Related Experiment Videos
Clinical predictive artificial intelligence evaluation: A narrative review of trial designs and practical
Maxime Fosset1,2,3,4, Joris Pensier3,4,5, Boris Jung3,4,6,7
1Medical Intensive Care Unit, Lapeyronie Montpellier University Hospital, University of Montpellier, Montpellier, France.
None:
Artificial intelligence (AI) predictive models demonstrate potential for transforming clinical decision-making across medicine. However, conventional randomized controlled trials (RCTs), the gold standard for evaluating medical interventions, are ill-suited for clinical AI tools due to their static design, lengthy timelines, and inability to accommodate algorithms that evolve and adapt to changing clinical contexts. In this narrative review, we outline the limitations of traditional evaluation frameworks and propose a paradigm shift toward adaptive, iterative, and context-specific assessment methodologies. Evaluating clinical AI in practice requires three interdependent but epistemologically distinct activities: performance monitoring, which tracks the technical characteristics of the deployed model (calibration, discrimination, data drift, alert burden, fairness, workflow fidelity); clinical impact monitoring, which observationally and prospectively tracks whether the initial clinical benefit appears sustained over time; and scientific evidence generation, which produces causal estimates of deployment effects on patient outcomes through pragmatic, adaptive trial designs and causal inference techniques. We propose a predictive-AI-specific framework that links performance monitoring, clinical impact monitoring, evidence generation, causal estimands, and governance of model updates into one coherent decision pathway for clinicians and trialists. We present a governance-driven escalation protocol specifying when monitoring signals should trigger formal evidence generation, a decision pathway mapping signal types (performance or clinical impact) to trial design classifications, and a guide to causal inference methods for clinical AI trials. Drawing from adaptive platform and pragmatic trial designs, we recommend continuous monitoring approaches that prioritize patient-centered outcomes, health equity, and workflow integration over narrow performance metrics, and provide actionable steps to design a clinical AI trial. Successful implementation requires clinician engagement, transparency, and ongoing education regarding AI capabilities and limitations. Within this new evaluation paradigm, predictive AI can progress from a promising technology to reliable clinical tools that improve patient outcomes, support clinical decision-making, and uphold ethical standards in routine practice.