Related Experiment Video
Updated: Sep 14, 2025

A Neonatal Imaging Model of Gram-Negative Bacterial Sepsis
Published on: August 12, 2020
Human performance evaluation of a pediatric artificial intelligence sepsis model
Swaminathan Kandaswamy1, Naveen Muthu1,2, Nikolay Braykov2
1Pediatrics, Emory University School of Medicine, Atlanta, GA, 30307, United States.
Insights
An artificial intelligence (AI) sepsis alert improved clinician situation awareness in the emergency department (ED), but also showed some automation bias and increased workload. This study highlights the feasibility of mixed-methods evaluation for AI in clinical practice.
Area of Science:
- Clinical Informatics
- Artificial Intelligence in Healthcare
- Patient Safety
Background:
- Sepsis remains a critical challenge in pediatric emergency care.
- Predictive artificial intelligence (AI) models offer potential for early sepsis detection.
- Evaluating the impact of AI on human performance is crucial for successful implementation.
Purpose of the Study:
- To assess the influence of an implemented AI model for predicting pediatric sepsis on human performance measures in the emergency department (ED).
- To evaluate clinician situation awareness, trust, workload, and automation bias related to the AI sepsis alert.
Main Methods:
- A mixed-methods approach combining qualitative interviews with quantitative electronic health record (EHR) data analysis.
- Interviews with 40 ED providers and nurses within 72 hours of a patient being flagged by the AI sepsis model.
- Assessment of human performance metrics including situation awareness, explainability, human-computer agreement, workload, trust, automation bias, and staff-patient relationships.
Main Results:
- The AI sepsis alert improved clinician situation awareness, influencing patient care management, resource allocation, and monitoring.
- Clinicians reported an average trust of 3.8/5 in the AI alert; however, 28% of sepsis huddles occurred without an alert, indicating potential automation bias.
- Antibiotic treatment rates for sepsis cases were similar pre- and post-intervention without a huddle, but doubled with the intervention; NASA Task Load Index increased from 43 to 57.
Conclusions:
- The AI sepsis prediction model generally showed positive impacts on human performance, enhancing situation awareness and satisfaction with alert-driven sepsis huddles.
- Evidence of automation bias and a slight increase in workload were observed.
- The study demonstrates the feasibility of a mixed-methods approach for evaluating AI in clinical practice and suggests future research should focus on reducing measurement burden and correlating human performance with clinical outcomes.
Objective:
To assess the influence of an implemented artificial intelligence model predicting pediatric sepsis (defined by IPSO-Improving Pediatric Sepsis Outcomes collaborative) in the emergency department (ED) on human performance measures.
Materials And Methods:
Two ED sites within a large pediatric health system in the Southeastern United States between January 1, 2021 and April 1, 2024. We interviewed ED providers and nurses within 72 hours of caring for a patient identified as potentially having sepsis by the predictive model. Thematic analysis of qualitative data was combined with electronic health record queries to assess measures of human performance, including situation awareness, explainability, human-computer agreement, workload, trust, automation bias, and relationship between staff and patients.
Results:
We interviewed 40 clinicians. Participants found that the sepsis alert improved situation awareness, leading to changes in patient care management, resource allocation, and/or monitoring. Participants reported an average trust in the model-based alert of 3.8/5. Only 28% (555/1977) of sepsis huddles were done without alert firing, suggesting some automation bias. Treatment with antibiotics for IPSO sepsis cases was similar pre- and post-intervention without a huddle (9.3% vs 10.5%), though treatment doubled with huddle intervention (22.7%). NASA Task Load Index increased from 43 to 57 post-intervention. There was no report of adverse relationships with patients post-intervention.
Discussion:
Human performance appeared to be generally positive with improved situation awareness and satisfaction with the alert-driven huddle. However, there was some evidence of automation bias and a slight increase in workload with the intervention.
Conclusion:
This study demonstrates the feasibility of evaluating multiple dimensions of human performance using a mixed methods approach for an AI model implemented in clinical practice. Future studies should aim to reduce the measurement burden of human performance metrics associated with AI implementation in acute care settings and assess the correlation between human performance measures and clinical outcomes.

