Related Experiment Video
Updated: Aug 3, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Hurdles to Artificial Intelligence Deployment: Noise in Schemas and "Gold" Labels
Mohamed Abdalla1, Benjamin Fine1
1Institute for Better Health, Trillium Health Partners, Mississauga, Ontario, Canada (M.A., B.F.); and Centre for Information Technology, Department of Computer Science (M.A.), and Department of Medical Imaging (B.F.), University of Toronto, 40 St George St, Room 4283, Toronto, ON, Canada M5S 2E4.
Noise in medical imaging datasets, including labeling schema variations and inconsistent annotations, challenges the reliability of artificial intelligence (AI). Addressing these dataset creation issues is crucial for safe clinical AI deployment.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Data Science
Background:
- Clinicians express concerns regarding the safety and robustness of artificial intelligence (AI) in medical imaging, despite AI performance reports.
- Underreported sources of noise in imaging AI include variations in labeling schema definitions and noise within the labeling process itself.
Purpose of the Study:
- To investigate the impact of labeling schema variations and labeling process noise on medical imaging AI.
- To quantify the extent of label inconsistency and noise in publicly available datasets and radiologist annotations.
Main Methods:
- Compared schema overlap between two public datasets and a third-party vendor.
- Analyzed individual radiologist annotations from the CheXpert test set to quantify labeling noise.
- Evaluated label agreement across different classes to identify sources of unreliability.
Main Results:
- Low agreement (<50%) was found between labeling schemas of different datasets and vendors.
- Label noise varied by class; high agreement (>90%) for pneumothorax and medical devices, but low agreement for pneumonia and consolidation.
- The reliability of 'ground truth' labels for low-agreement classes was questionable, indicating dependence on the annotating radiologist group.
Conclusions:
- Noise in labeling schemas and gold standard annotations is prevalent in medical imaging classification.
- This noise negatively impacts the downstream clinical deployment and trustworthiness of AI tools.
- Potential solutions involving task design, annotation methods, and model training are discussed to enhance clinical AI trust.
More Related Videos
07:31Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
Published on: February 8, 2019
07:26Executing Complexity-Increasing Queries in Relational MySQL and NoSQL MongoDB and EXist Size-Growing ISO/EN 13606 Standardized EHR Databases
Published on: March 19, 2018
Related Concept Videos
Schemas
Natural and Artificial Concepts
Heuristics
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
Stereotype Content Model
Hypothesis: Accept or Fail to Reject?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as: