Related Experiment Videos
CODE (comprehensive oral mucosa database with explanations): a smartphone-based normative intraoral image dataset for
Madan Kumar Parangimalai Diwakar1, Ranganathan Kannan2, Lavanya Chandra2
1Department of Public Health Dentistry, Ragas Dental College and Hospital, Chennai, Tamil Nadu, 600119, India.
Abstract:
The CODE (Comprehensive Oral Mucosa Database with Explanations) dataset is a publicly available smartphone-based intraoral image repository developed to facilitate artificial intelligence (AI) research in oral healthcare. The dataset was generated as part of the Indian Council of Medical Research (ICMR)-funded SMITA project using a standardized hub-and-spoke oral health screening framework implemented across Tamil Nadu, India. A total of 5,632 intraoral images were acquired from 704 participants using standardized smartphone imaging protocols, of which 5,566 images satisfied predefined quality assessment criteria and were included in the final dataset. The repository comprises 5,386 normative images and 180 abnormal images representing variations from normal oral mucosa, oral potentially malignant disorders (OPMDs), and oral cancer. Images were captured from predefined intraoral anatomical sites under standardized acquisition procedures and linked to anonymized participant-level demographic and clinical metadata. Expert-reviewed polygon annotations were generated using the VGG Image Annotator (VIA) platform and are provided in JSON format together with annotated and unannotated JPEG images and structured metadata files. Technical validation was performed using an unsupervised convolutional autoencoder to demonstrate the suitability of the dataset for representation learning and anomaly detection rather than diagnostic performance evaluation. The predominance of normative images reflects the intended purpose of the dataset as a resource for baseline anatomical feature learning, self-supervised learning, and anomaly detection. The CODE dataset provides a reusable benchmark resource for image segmentation, explainable AI, image quality assessment, tele-dentistry, multimodal learning, and future AI-assisted oral disease screening research.