Related Experiment Video
Updated: Feb 28, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Novel Hybrid Large Language Model Approach for Reporting Panoramic Radiographs and Performance Comparison with
Yunus Balel1, Kaan Sağtaş2, Fatih Teke2,3
1Department of Oral and Maxillofacial Surgery, Faculty of Dentistry, Sivas Cumhuriyet University, 58140, Merkez, Sivas, Turkey. yunusbalel@hotmail.com.
This study introduces a hybrid AI framework for dental radiology, combining deep learning image analysis with large language models (LLMs) to improve reporting accuracy and reduce errors in panoramic radiograph interpretation.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Dental Radiology
- Machine Learning Applications
Background:
- Current multimodal AI systems struggle with panoramic radiograph interpretation due to low diagnostic accuracy and high hallucination rates.
- Large language models (LLMs) show promise for clinical reporting but require integration with robust image analysis for reliability.
Purpose of the Study:
- To develop and evaluate a hybrid AI framework integrating deep learning image analysis with LLM-driven reporting for enhanced reliability in dental radiology.
- To assess the performance of this framework in terms of reporting accuracy, structural validity, consistency, and hallucination rates.
Main Methods:
- A YOLOv12 model was trained on 30,954 panoramic radiographs for tooth detection and 14-category segmentation.
- Detection and segmentation outputs were converted to structured JSON data for processing by locally hosted LLMs (DeepSeek R1, Mistral, Llama 3.2, Gemma 3, Qwen3, SmolLM3).
- Reporting accuracy was evaluated on 50 unseen radiographs, comparing AI outputs against expert assessments.
Main Results:
- The segmentation model achieved an overall F1-score of 0.708, with high performance for ectopic/supernumerary teeth (0.994), impacted teeth (0.990), and implants (0.984).
- All hybrid LLMs generated structurally valid JSON outputs (100%). DeepSeek R1 demonstrated the highest reporting accuracy (466 true findings) and lowest hallucinations (30).
- Commercial LLMs exhibited hallucinations in 100% of reports, whereas the hybrid framework significantly minimized these errors.
Conclusions:
- Integrating structured image-derived findings with LLM reasoning markedly improves reporting accuracy and minimizes hallucinations in dental radiology.
- The developed hybrid framework offers a promising and reliable solution for AI-assisted interpretation of dental radiographs, outperforming current commercial LLMs.
More Related Videos
09:10Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
10:23Author Spotlight: Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023