Related Experiment Video
Updated: Aug 3, 2026

Digital Hybrid Model Preparation for Virtual Planning of Reconstructive Dentoalveolar Surgical Procedures
Published on: August 5, 2021
From inconsistent annotations to ground truth: Aggregation strategies for annotations of proximal carious lesions in
Vanessa Klein1, Martha Büttner2, Gerd Göstemeyer3
1Department of Oral Diagnostics, Digital Health and Health Services Research, Charité - Universitätsmedizin Berlin, Berlin, Germany; Conservative Dentistry and Periodontology, LMU University Hospital, LMU Munich, Munich, Germany.
Objectives:
Annotating carious lesions on images is challenging. For artificial intelligence (AI) applications, the aggregation of heterogeneous multi-examiner annotations into one single annotation (e.g. via majority voting, MV) is usually needed. We assessed different aggregation strategies for multi-examiner annotations of primary proximal carious lesions on orthoradial radiographs and Near-Infrared Light Transillumination (NILT) images.
Methods:
A total of 1007 proximal surfaces from 522 extracted posterior teeth were assessed by five dentists. Histological analysis provided the gold standard. Surfaces were classified as (1) sound, (2) enamel lesion or (3) dentin lesion. Four label aggregation strategies - MV, Weighted Majority Voting (WMV), Dawid-Skene (DS), and multi-annotator competence estimation (MACE) - were applied to unimodal (radiographs, NILT) and multimodal (combined) datasets. The area under the receiver operating characteristic curve (AUROC) was the primary outcome metric.
Results:
According to the gold standard, 637 (63 %) surfaces were sound, 280 (28 %) showed carious lesions limited to the enamel, and 90 (9 %) showed lesions extending into the dentin. For radiographs, aggregation using MACE outperformed MV, WMV and DS significantly across all lesion depths (p < 0.002). For NILT, MACE significantly outperformed MV across all lesion depths (p < 0.001) and DS for enamel and dentin lesions (p ≤ 0.002). In the multimodal dataset, DS outperformed the other label aggregation strategies across all lesion depths significantly (p < 0.05).
Conclusions:
The commonly applied MV may be suboptimal. There is a need for informed application of specific aggregation strategies, depending on the dataset characteristics.
Clinical Significance:
Most AI applications for dental image analysis are trained on a single annotation, usually resulting from aggregated multi-examiner annotations of each image. However, since these annotations are usually aggregated in an in vivo setting where no definitive ground truth is available, the choice of aggregation strategy plays a crucial role.
Related Concept Videos
X-ray Imaging
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Tooth Anatomy
The Crown, Neck, and Root
The visible part of the tooth is referred to as the crown. It's covered by enamel, the hardest substance in the human body. The crown is uniquely shaped for each type of tooth, allowing for different functions such as cutting, tearing, or grinding food.
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT

