Related Experiment Videos
A Deep Learning system for Automated Quality Control of Medical Questionnaire data in Large-Scale Cohort Studies
Minmin Cao1,2, Yanan Duan2,3, Dong Lang2,4,5
1Chengdu University of Traditional Chinese Medicine, Chengdu 610000, China.
Abstract:
Recent advances in natural language processing combined with deep learning have transformed disease prediction, diagnosis, and treatment, yet their application in automated quality control of medical questionnaire data remains largely unexplored. Here, we present an end-to-end system tailored for large-scale cohort interview data. A fine-tuned Whisper model was developed for automatic audio-to-text transcription and optimized for recognizing domain-specific medical terminology and regional dialects, achieving an accuracy of over 91%. Using the DeepSeek framework, we jointly analyzed transcribed texts, interviewer records, and questionnaire content from 522 medical interviews to identify quality issues, achieving the area under the receiver operating characteristic curve (AUC) values of 0.893, 0.953, and 0.973 for omitted question stems, response-record inconsistencies, and insufficient follow-up, respectively. The system achieved high agreement with the senior specialist reference (F1 = 0.843, Kappa = 0.835) and significantly outperformed the junior specialist group (P < 0.01). Evaluation on an independent validation dataset (n = 96), which was collected using a structurally distinct brain health questionnaire, further demonstrated robust performance and generalizability (F1 = 0.831, Kappa = 0.821). Collectively, this framework enhances the efficiency and reliability of medical questionnaire data processing and offers a practical artificial intelligence-driven solution for quality control in global health data, especially in low-resource dialect settings.