Related Experiment Video
Updated: Jun 12, 2026

10:38
Observational Study Protocol for Repeated Clinical Examination and Critical Care Ultrasonography Within the Simple Intensive Care Studies
Published on: January 16, 2019
20.7K
Evaluating AI-based comprehensive clinical decision support for sepsis and ARDS: protocol for a Clinician Turing Test
Antonia Angeli Gazola1, Nicholas S Bishop1, Benjamin E Schmid1
1Palliative and Advanced Illness Research (PAIR) Center, University of Pennsylvania Perelman School of Medicine, Philadelphia, Pennsylvania, USA.
BMJ Open
|December 25, 2025
Summary
This study introduces a Clinician Turing Test to evaluate an AI ventilator assistant (AVA) for sepsis and ARDS patients. If clinicians cannot distinguish AVA’s recommendations from human ones, it signals preclinical safety.
Area of Science:
- Critical Care Medicine
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
Background:
- Evaluating AI clinical decision support systems (CDSSs) in practice is challenging due to limited early-stage data.
- High-stakes decisions in the Intensive Care Unit (ICU) for conditions like sepsis and Acute Respiratory Distress Syndrome (ARDS) necessitate robust AI validation.
- The AI Ventilator Assistant (AVA) was developed for sepsis ARDS patients on mechanical ventilation, but predictive performance alone is insufficient for safety assessment.
Purpose of the Study:
- To introduce and validate a novel Clinician Turing Test for assessing the safety and appropriateness of the AI Ventilator Assistant (AVA).
- To determine if critical care clinicians can differentiate between AI-generated and human-generated treatment recommendations for sepsis ARDS patients.
- To provide a preclinical signal of AVA's safety and appropriateness before large-scale clinical deployment.
Main Methods:
- A multisite, randomized, electronic, vignette-based Phase 1b study employing a Clinician Turing Test design.
- Recruitment of 350 critical care clinicians (physicians, advanced practice providers) across six US hospitals.
- Participants reviewed clinical vignettes with AI-generated or human-enacted treatment plans, randomly assigned (1:1), to identify the source. Primary endpoint: accuracy in source identification via mixed-effects logistic regression.
Main Results:
- The study is designed to assess the primary endpoint of clinician accuracy in distinguishing AI-generated from human-generated treatment recommendations.
- Secondary endpoints include evaluating clinicians' perceptions of safety, appropriateness, confidence, and interest in AI CDSSs.
- The study aims to provide preliminary data on AI CDSS clinical appropriateness without the risks of actual deployment.
Conclusions:
- The Clinician Turing Test offers a novel approach to evaluating AI CDSS safety and appropriateness in a preclinical setting.
- Successful 'passing' of the Turing test by AVA would indicate a strong preclinical signal of its safety and appropriateness.
- This validation method informs decisions regarding future clinical implementation and evaluation of AI CDSSs in real-world environments.
