Related Experiment Videos
Robustness of Healthcare ML Under Data Quality Degradation: A Dimension-Wise Analysis on MIMIC-IV
Nour Idris Pacha1,2,3, Saber Aloui1, Asma Rabaoui2
1Angers University Hospital, Angers, France.
Abstract:
Using the MIMIC-IV database (1,500-7,000 patients), we assess the robustness of healthcare machine learning models under controlled data quality (DQ) degradations applied to training or test data across five dimensions. Model performance declined with increasing degradation, with the largest losses observed when degradations occurred at inference time. Completeness, coherence, and validity were the most detrimental dimensions, whereas precision and uniqueness had more limited impact. Models adapted to training-time imperfections but remained sensitive to previously unseen degradations at deployment. Clinically, these findings highlight the importance of preserving data completeness and coherence at inference to support reliable healthcare AI systems.
Insights
Healthcare machine learning models degrade with data quality issues, especially during inference. Maintaining data completeness and coherence is crucial for reliable AI in clinical settings.
Area of Science:
- Healthcare Machine Learning
- Data Quality in AI
- Clinical Informatics
Background:
- Machine learning models are increasingly used in healthcare.
- The performance of these models is sensitive to data quality.
- Understanding data quality's impact is vital for reliable AI deployment.
Purpose of the Study:
- To assess the robustness of healthcare machine learning models against controlled data quality degradations.
- To identify which data quality dimensions most impact model performance.
- To evaluate model performance when degradations occur during training versus inference.
Main Methods:
- Utilized the MIMIC-IV database with 1,500-7,000 patients.
- Applied controlled data quality degradations across five dimensions (completeness, coherence, validity, precision, uniqueness).
- Evaluated model performance on training and test data subjected to these degradations.
Main Results:
- Model performance decreased as data quality degradation increased.
- The most significant performance drops occurred when degradations were applied at inference time.
- Completeness, coherence, and validity were the most detrimental data quality dimensions.
- Models showed some adaptation to training-time imperfections but remained vulnerable to novel inference-time degradations.
Conclusions:
- Healthcare AI model robustness is significantly challenged by data quality issues, particularly at inference.
- Ensuring data completeness and coherence during inference is critical for trustworthy clinical AI systems.
- Model adaptation to training data imperfections does not guarantee resilience to deployment-time data quality variations.
Related Concept Videos
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
Improving Translational Accuracy
Improving Translational Accuracy
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe and...
Healthcare Agencies II
Parish nursing is a growing specialty nursing profession that focuses on holistic healthcare, health promotion, and illness prevention. It blends professional nursing practice with a health ministry, focusing on health and healing within the context of a Christian community. Parish nurses serve as health educators, referral sources, and lay...
Bioequivalence Data: Statistical Interpretation