Related Experiment Video
Updated: Aug 29, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Comparative evaluation of artificial intelligence-assisted literature search tools for identifying clinically
Riccardo Di Febo1,2,3, Maximiliano Jeanneret Medina4, Alexandre Renaud5
1Department of Cardiology, Hôpital Saint Philibert, Groupement des Hôpitaux de L'Institut Catholique de Lille, Université Catholique de Lille, 115 Rue du Grand But, 59160 Lille, France.
Aims:
The rapid expansion of biomedical literature challenges clinicians' and researchers' ability to identify clinically meaningful evidence. We systematically compared five literature search tools, four artificial intelligence (AI)-assisted and one conventional, across clinically relevant cardiology research scenarios, using a blinded expert-validated gold standard to assess their ability to retrieve relevant and key references.
Methods And Results:
We evaluated ChatGPT-5, Elicit, Consensus, Scite, and PubMed across four cardiology topics defined by maturity and specificity, with multiple standardized prompts. Three electrophysiology experts independently and blindly rated all retrieved references, defining two gold standards: expert-rated relevance and expert-selected key references. ChatGPT-5 achieved the highest proportion of relevant articles (90% [88-100], P < 0.001) and the highest key-reference overlap (60% [43-68], P < 0.001), whereas Scite performed lowest (20% and 10%, respectively). The tool was the primary determinant of performance (partial R 2 = 0.50), whereas prompt formulation had no significant effect. In a pre-specified subanalysis restricted to clinical studies, ChatGPT-5 and human-conducted systematic reviews overlapped by 42% (96% of shared articles highly relevant), with 58% distinct references, indicating complementary AI and human retrieval; ChatGPT-5 produced hallucinated citations when long reference lists were requested for emerging topics, underscoring the need for human verification.
Conclusion:
AI-assisted tools showed heterogeneous performance, ChatGPT-5 performing best in this cardiology setting. These preliminary, context-specific findings support hybrid human-AI strategies in which AI complements rather than replaces transparent database searches such as PubMed; larger-scale, multi-domain studies are needed to confirm and generalize them.