Related Experiment Video
Updated: Jan 28, 2026

09:34
Semi-Automated Planimetric Quantification of Dental Plaque Using an Intraoral Fluorescence Camera
Published on: January 27, 2023
2.5K
Comparative Evaluation of Vision-Language Models for Detecting and Localizing Dental Lesions from Intraoral Images
Maria Jahan1, Al Ibne Siam1, Lamim Zakir Pronay2
1Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh.
Journal of Imaging
|January 27, 2026
Summary
YOLOv8 demonstrated superior performance in detecting dental caries from intraoral images compared to vision-language models. Further research with larger datasets is needed to enhance the capabilities of Florence-2 and PaLI-Gemma for dental lesion classification.
Area of Science:
- Artificial Intelligence in Dentistry
- Medical Imaging Analysis
- Computer Vision
Background:
- Intraoral photo imaging is crucial for diagnosing dental caries.
- Vision-language models offer potential for automated analysis of dental images.
- Current models require evaluation for their efficiency in detecting various dental lesions.
Purpose of the Study:
- To evaluate the performance of Florence-2, PaLI-Gemma, and YOLOv8 in detecting and classifying carious and non-carious dental lesions.
- To compare the efficacy of these models using a standardized dataset of intraoral photographs.
Main Methods:
- A dataset of 172 annotated intraoral images was used, classifying lesions such as microcavitation, cavitated lesions, staining, calculus, and non-carious lesions.
- Models were trained on an 80:10:10 data split.
- Performance was assessed using mean average precision (mAP), mAP50-95, and class-specific precision and recall metrics.
Main Results:
- YOLOv8 achieved the highest overall performance with a mean average precision (mAP) of 37%, 42.3% precision, and 31.3% recall, including 100% precision for cavitation detection.
- PaLI-Gemma achieved a recall of 13% and 21% (precision not specified).
- Florence-2 yielded a mean average precision (mAP) of 10%, with 51% precision and 35% recall.
Conclusions:
- YOLOv8 significantly outperformed Florence-2 and PaLI-Gemma in detecting and classifying dental lesions from intraoral images.
- Despite multimodal capabilities, Florence-2 and PaLI-Gemma underperformed, suggesting a need for larger, more diverse datasets.
- Hybrid architectures may be necessary to improve the performance of vision-language models in dental imaging applications.
Related Concept Videos
Vision
60.0K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
60.0K
Language
904
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
904
Color Vision
1.5K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.5K
Components of Language
809
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
809
Language Development
892
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
892
Language and Cognition
763
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
763

