Related Experiment Video
Updated: Aug 6, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
COLLECT: A Counterfactual Oral Lesion Library for Trustworthy AI Evaluation
K T Tafti1, Z D'Souza1, A Ajam1
1Faculty of Dental Medicine and Oral Health Sciences, McGill University, Montreal, Quebec, Canada.
Journal of Dental Research
|July 25, 2026
Summary
This study introduces COLLECT, a new dataset for evaluating AI in diagnosing oral lesions. It helps make AI models more transparent and reliable for clinical use in dentistry.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Oral Pathology
Background:
- Intraoral lesions present diagnostic challenges due to varied appearances.
- Deep learning models excel at oral lesion classification but lack clinical interpretability.
- A dedicated counterfactual dataset for intraoral images is needed to enhance model transparency.
Purpose of the Study:
- To introduce the Counterfactual Oral Lesion Library for Explainable Concept Testing (COLLECT).
- To provide a framework for assessing the interpretability of AI models for intraoral lesion diagnosis.
- To create a counterfactual dataset for evaluating AI sensitivity to clinical features.
Main Methods:
- Developed COLLECT, a dataset of 600 edited intraoral lesion images across 4 types (aphthous ulcer, geographic tongue, hairy tongue, oral squamous cell carcinoma).
- Modified images based on clinical concepts: size, color, and opacity.
- Evaluated common convolutional neural networks using metrics like true class probability, entropy, TCAV, and attribution scores.
Main Results:
- Counterfactual edits decreased model accuracy and increased uncertainty, with opacity and color changes having the most significant impact.
- Network layers showed varying responsiveness to concept changes, with later layers relying more on baseline features.
- Attribution scores indicated opacity, color, and size changes most affected model confidence, with some models showing bias towards background anatomy.
Conclusions:
- COLLECT enables systematic evaluation of AI models under clinically relevant counterfactual conditions.
- The dataset and framework improve transparency in automated oral lesion diagnosis.
- Findings highlight the need for robust AI evaluation to ensure reliance on lesion-specific features.