Related Experiment Video
Updated: May 26, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are They Objectively up to Medical Standards? A Real-Life Assessment
Natalia Castaño-Villegas1, María Camila Villa1, Katherine Monsalve Barrientos1
1Arkangel AI, Bogotá, Colombia.
Objective:
To compare the performance of 4 large language model chatbots in response time and quality of clinical answers, evaluated by specialists using predefined validity criteria.
Participants And Methods:
Between June 1 and September 20, 2025, four clinical vignettes (orthopedics, pediatrics, gynecology, and psychiatry) were developed by independent experts and answered by 4 conversational agents: Arkangel AI, OpenEvidence, ChatGPT, and Medisearch. Each vignette included 4 questions (diagnosis, clinical management, research, and general knowledge). Responses were independently evaluated by external clinicians using an 8-criterion Likert scale assessing correctness, consensus agreement, absence of bias, adherence to standards of care, timeliness, patient safety, authenticity of cited references, and contextual appropriateness. Response times were summarized using medians and interquartile ranges.
Results:
A total of 128 question-answer pairs (1024 evaluations) were analyzed. Overall satisfaction ranged from 71.1% (727) to 93% (952) across agents, with statistically significant differences (Kruskal-Wallis P<.001). Dissatisfaction was most frequent for reference authenticity in some ChatGPT modes (75% [n 91/128]-97% [n 119/128]), whereas Arkangel AI-Deep, ChatGPT-Deep, and OpenEvidence showed 100% satisfaction for this criterion. Satisfaction was high for correctness and consensus, with greater variability for bias and patient safety. Satisfaction differed by specialty, with higher scores in gynecology and lower scores in pediatrics (P<.001). Median response times ranged from 18 seconds to 12.8 minutes, with significant differences across agents and modes (Wilcoxon P<.05).
Conclusion:
Large language model chatbots showed substantial variability across validity dimensions when assessed using expert clinical judgment, supporting the need for standardized, multidimensional evaluation frameworks.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Current Trends in Nursing II
Healthcare Agencies II
Parish nursing is a growing specialty nursing profession that focuses on holistic healthcare, health promotion, and illness prevention. It blends professional nursing practice with a health ministry, focusing on health and healing within the context of a Christian community. Parish nurses serve as health educators, referral sources, and lay...
Healthcare Agencies I
Health Information Technology and Healthcare Information System
Health Information Technology, commonly called HIT, integrates advanced information systems and technology in healthcare settings. Its primary functions include:
