Related Experiment Video
Updated: Oct 1, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Prospective multicentre expert evaluation of an artificial intelligence-based oncology clinical decision support
Monika Shekhawat1, Ashok Kumar Nagar2, Arpita A Gupta3
1Department of Surgical Oncology, King George's Medical College, Lucknow, India.
Introduction:
Cancer is a major public health challenge in India, with an estimated 1.4 million new cases each year. Despite this burden, the ratio of oncologists to new cancer patients is critically low, about one specialist per 1600 new diagnoses per year. The AI-driven CDSS could help offset the oncologist shortage in LMICs. OneRx.AI is an RAG platform grounded in international and national oncology guidelines. We evaluated its Expert-rated accuracy, clinical safety, Guideline concordance and decision impact, confidence, efficiency, and inter-rater reliability (IRR) in Indian practice.
Methodology:
A prospective, multicentre, expert-validation study conducted at three tertiary oncology centres in India. Three Oncologist-clinicians independently evaluated de-identified queries across five cancer types. Individuals and common-prompt evaluation phases were analysed separately. Co-primary endpoints were assessed using one-sample z-tests with 95% Wilson confidence intervals, and inter-rater reliability was measured using the pre-specified Gwet's AC1 statistic.
Results:
Of 124 submitted clinical queries, 110 (88.7%) met the eligibility criteria. Expert-rated fully correct clinical accuracy was 82.7%, expert-rated clinical safety was 100%, and fully concordant guideline recommendations were 94.5%. Composite rates for Expert-rated clinical accuracy and guideline concordance were 100%. Expert-rated Clinical accuracy showed Gwet's AC1 = 0.707. Expert-rated Safety and guideline concordance received uniform ratings across evaluators. Citation accuracy was 86.4%, decision-influence rate 92.7%, physician-confidence increased by a mean of 1.69 (p < 0.0001), and median time to clinical recommendation decreased from 389.2 s to 47.2 s, representing an 87.9% reduction (p < 0.0001).
Conclusion:
OneRx.AI demonstrated high expert-rated clinical performance in this exploratory, ceiling-affected evaluation, supporting progression to a larger Stage II study.