Related Experiment Videos
Theoretical Exploration of Error Thresholds for Clinical AI Decision Support in Nursing: Exploratory Simulation Study
1Faculty of Nursing, Shumei University, 1-1 Daigaku-cho, Yachiyo, Chiba, 276-0003, Japan, 81 47-488-2111.
Background:
Clinical AI decision support is being introduced into nursing practice; however, existing large language models (LLMs) demonstrate only moderate accuracy on complex clinical tasks, raising questions about the level of accuracy required for safe clinical use across varying levels of clinician experience and task complexity.
Objective:
The aim of the study is to develop an empirically calibrated simulation model of human-AI reliance and error in nursing decision-making and estimate the AI accuracy required to achieve specified error rate targets.
Methods:
A linear reliance model with coefficients for AI accuracy (A), clinician experience (E), and task complexity (C) was calibrated using weighted least squares against 9 empirical data points from 3 independent randomized experiments on AI-assisted decision-making (N=3502). Predicted error was computed as reliance×(1-A) across a 27-cell factorial design. Study-level bootstrap (2000 iterations) quantified calibration uncertainty. To contextualize the simulation's operating range, the accuracy of contemporary general-purpose LLMs on complex clinical tasks was drawn from published benchmarks.
Results:
Calibration placed βA at 0.201 (bootstrap 95% CI 0.023-0.234; P(βA>0)>.99). For the novice×high-complexity combination, the minimum AI accuracy values required to achieve error rates <10% and<20% were 0.89 and 0.78, respectively (bootstrap 95% CIs 0.88-0.90 and 0.75-0.79). At the moderate accuracy levels currently reported for general-purpose LLMs on complex clinical tasks (approximately 0.5-0.7), the model predicts error rates of roughly 26% to 41% in this high-risk condition.
Conclusions:
In this model, keeping predicted error rates below a stringent target (<10%) for high-complexity nursing decision support by novice clinicians requires AI accuracy of at least approximately 0.89, a level that current general-purpose LLMs may not reliably reach on complex clinical tasks. Because the model is calibrated on nonnursing reliance data, these thresholds are illustrative model outputs, not nursing-derived empirical standards. The Athreshold framework provides a decision-theoretic tool for evaluating the minimum AI accuracy requirement by user-and-task profile. Behavioral validation in nursing contexts remains an essential next step. Because the framework is independent of any specific model, it remains applicable as AI systems improve.
Related Concept Videos
Current Trends in Nursing II
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters assessment...