Related Experiment Video
Updated: Jan 28, 2026

Semi-Automated Planimetric Quantification of Dental Plaque Using an Intraoral Fluorescence Camera
Published on: January 27, 2023
Comparative Evaluation of Vision-Language Models for Detecting and Localizing Dental Lesions from Intraoral Images
Maria Jahan1, Al Ibne Siam1, Lamim Zakir Pronay2
1Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh.
Abstract:
To assess the efficiency of vision-language models in detecting and classifying carious and non-carious lesions from intraoral photo imaging. A dataset of 172 annotated images were classified for microcavitation, cavitated lesions, staining, calculus, and non-carious lesions. Florence-2, PaLI-Gemma, and YOLOv8 models were trained on the dataset and model performance. The dataset was divided into 80:10:10 split, and the model performance was evaluated using mean average precision (mAP), mAP50-95, class-specific precision and recall. YOLOv8 outperformed the vision-language models, achieving a mean average precision (mAP) of 37% with a precision of 42.3% (with 100% for cavitation detection) and 31.3% recall. PaLI-Gemma produced a recall of 13% and 21%. Florence-2 yielded a mean average precision of 10% with a precision and recall was 51% and 35%. YOLOv8 achieved the strongest overall performance. Florence-2 and PaLI-Gemma models underperformed relative to YOLOv8 despite the potential for multimodal contextual understanding, highlighting the need for larger, more diverse datasets and hybrid architectures to achieve improved performance.
Related Concept Videos
Vision
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Color Vision
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition

