Related Experiment Videos
Process-Oriented, Behaviorally Anchored Assessment of Clinical Reasoning in Large Language Models and the Effect of
Adnan Agha1, Muhammad Jalil2, Eram Anwar1
1Department of Internal Medicine, College of Medicine and Health Sciences, United Arab Emirates University, AlAin, Abu Dhabi, United Arab Emirates.
Background:
Most clinical reasoning evaluations in large language models (LLMs) score only the final answer, usually multiple-choice accuracy, which explains little about model reasoning. Two developments stress this gap: reasoning-optimized models are now common, and several expose an explicit extended thinking control. Whether that deliberation improves the reasoning process remains untested with a validated instrument.
Objective:
This protocol introduces and aims to validate the Rapid Evaluation Assessment of Clinical Reasoning Tool (REACT)-AI, a Behaviorally Anchored Rating Scale (BARS) with 13 subdomains for process-oriented assessment of AI clinical reasoning, that is, the quality of the externalized reasoning rather than its faithfulness to the model's internal computation. It also builds a conflict-of-interest-controlled LLM-as-judge pipeline and tests whether extended thinking improves reasoning quality.
Methods:
This 2-phase, prospective, comparative study has a within-model thinking-mode factor. Six flagship models are run in standard and extended thinking or reasoning modes, giving 12 conditions: 4 providers are toggled within the same model, and 2 pair a standard model with a dedicated reasoning model. In phase 1, 5 standardized urgent care vignettes are answered under all 12 conditions, 3 runs per condition (180 AI outputs), and by an independent expert clinician panel (n=5) and senior and junior medical students. Every output is scored by dual-blind human raters and, in parallel, by LLM judges, with no model judging its own family. Acceptance criteria for deploying the judge at scale are a weighted κ of at least 0.60 and an intraclass correlation coefficient above 0.75. In phase 2, the validated pipeline scores 3600 AI outputs. A self-correction turn and a paired Gulf English condition run alongside and are reported separately, with the AI metacognition assessment rubric and disinformation generation rate as secondary instruments.
Results:
As of June 2026, ethics approval was obtained (ERSC_2025_6124), and the study is registered on the Open Science Framework under embargo. Responses to the 5 phase 1 vignettes were collected from senior and junior medical students between January and June 2026; 173 scripts were received and remain sealed, unopened, and unscored. No AI outputs have been generated, and the expert panel remains unrecruited. Following the June 1, 2026, model-version lock, AI generation and scoring will begin in September 2026, with phase 1 calibration through December 2026 and phase 2 from January to April 2027; results are expected in winter 2027. The study will report the validated instrument, judge calibration against human experts, self-preference bias by model family, and the effect of extended thinking on reflection and metacognition, including null findings.
Conclusions:
REACT-AI and its judge pipeline are built to be reusable. They close the gap between accuracy benchmarks and genuine reasoning assessment and provide a template for studying reasoning mode effects.
Trial Registration:
OSF Registries osf.io/jep5t; https://osf.io/jep5t.
International Registered Report Identifier (Irrid):
DERR1-10.2196/103220.