Related Experiment Video
Updated: Sep 15, 2025

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Is a score enough? Pitfalls and solutions for AI severity scores
Michael H Bernstein1, Marly van Assen2, Michael A Bruno3
1Department of Diagnostic Imaging, Brown Radiology Human Factors Lab, Rhode Island Hospital, Warren Alpert School of Medicine of Brown University, Providence, RI, USA. Michael_Bernstein@brown.edu.
Artificial intelligence (AI) scores in radiology have six human factor limitations impacting their usefulness. Providing false discovery and omission rates could mitigate these AI score limitations.
Area of Science:
- Radiology
- Artificial Intelligence (AI)
- Psychological Science
- Statistics
Background:
- Artificial intelligence (AI) tools in radiology commonly provide severity scores, indicating pathology likelihood.
- The utility and transparency of these AI-generated scores remain under-examined.
- Existing research has not sufficiently addressed the radiologist-AI interaction.
Purpose of the Study:
- To elucidate six human factors limitations of AI scores in radiology.
- To propose a hypothesis for mitigating these limitations.
- To discuss empirical testing of the proposed hypothesis.
Main Methods:
- Analysis of AI score utility drawing on principles from psychological science and statistics.
- Identification of six key human factors limitations: inter-AI variability, intra-AI variability, inter-radiologist variability, intra-radiologist variability, unknown score distribution, and perceptual challenges.
- Formulation of a hypothesis involving false discovery rate (FDR) and false omission rate (FOR).
Main Results:
- Six human factors limitations were identified that undermine the utility of AI severity scores.
- These limitations include variability across and within AI systems and radiologists, unknown score distributions, and perceptual challenges.
- A hypothesis is proposed: FDR and FOR thresholds can mitigate these limitations.
Conclusions:
- The utility of AI severity scores in radiology is significantly constrained by human factors.
- Addressing variability and perceptual challenges is crucial for effective radiologist-AI interaction.
- Incorporating FDR and FOR as thresholds offers a potential strategy to enhance the reliability and utility of AI scores.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Non-equilibrium in the Cell
Survival Tree
Building a Survival Tree
Constructing a...
Machines: Problem Solving II
Machines: Problem Solving I
The toggle clamp system is a machine structure consisting of movable, pin-connected multi-force members that form a stabilized system to transmit forces. The...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
Mean Absolute Deviation
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...