Related Experiment Video
Updated: Aug 9, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
The impact of inconsistent human annotations on AI driven clinical decision making
Aneeta Sylolypavan1, Derek Sleeman2, Honghan Wu3,4
1Institute of Health Informatics, University College London, London, United Kingdom.
Annotation inconsistencies among clinical experts in supervised learning are common. Our study shows that simple consensus methods yield suboptimal models, but focusing on
Area of Science:
- Machine Learning
- Clinical Informatics
- Medical AI
Background:
- Supervised learning in healthcare relies on expert annotations, but inconsistencies arise from expert bias and judgment variations.
- The impact of these annotation inconsistencies on real-world clinical model performance is understudied.
Purpose of the Study:
- To investigate the implications of expert annotation inconsistencies in supervised learning for Intensive Care Unit (ICU) datasets.
- To evaluate current best practices for establishing gold-standard models and consensus in clinical AI development.
Main Methods:
- Developed individual supervised learning models using a common ICU dataset annotated independently by 11 clinical experts.
- Performed internal validation using Fleiss' kappa and external validation on a HiRID dataset using Cohen's kappa.
- Analyzed model agreement for discharge decisions versus mortality prediction.
Main Results:
- Internal validation showed fair agreement (Fleiss' κ = 0.383) among expert models.
- External validation revealed low pairwise agreement (average Cohen's κ = 0.255) between models.
- Models showed less agreement on discharge decisions (Fleiss' κ = 0.174) than mortality prediction (Fleiss' κ = 0.267).
- Standard consensus methods like majority voting produced suboptimal models.
Conclusions:
- No single 'super expert' reliably emerged, and simple consensus approaches are insufficient for optimal clinical AI.
- Assessing annotation learnability and using 'learnable' subsets for consensus building can lead to optimal models.
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Errors occurring during blood pressure monitoring
Several factors...
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
The Availability Heuristic
Methods of Documentation III: PIE

