Related Experiment Video
Updated: Aug 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Efficacy of a Large Language Model Data Extraction System in Evidence Reviews for Emerging Infectious Diseases: A
Masahiro Ishikane1, Yuki Kataoka2,3,4,5,6,7, Yasushi Tsujimoto4,8,9,10
1Disease Control and Prevention Center, National Centre for Global Health and Medicine, Japan Institute for Health Security, Shinjuku, Tokyo, Japan.
Background:
Rapid evidence synthesis during emerging infectious and re-emerging disease outbreaks is critical, yet traditional systematic reviews rarely meet urgent timelines. Large language models (LLMs) may accelerate evidence synthesis by extracting data from publications. We compared an LLM-assisted data extraction system with manual extraction.
Methods:
We conducted a 1:1, open-label, 2-period, randomized crossover trial at the National Center for Global Health and Medicine, a national reference center for emerging infectious diseases in Japan (2025). Five experienced reviewers extracted predefined items from mpox-related articles under 2 conditions: (i) LLM-assisted extraction using OpenAI's o3 model to generate structured summaries and (ii) manual review of PDF files. The primary outcome was task completion time; secondary outcomes were extraction accuracy and adverse events. Mixed-effects models included condition as a fixed effect and participant and paper IDs as random effects. The protocol, source code, and data are available at https://github.com/SRWS-PSG/emerging_infection_24K13518_open.
Results:
Five evaluators (4 physicians and 1 pharmacist; 6-10 years postgraduation) completed 20 task-level evaluations (LLM, n = 9; no LLM, n = 11). Mean completion time was 27.5 minutes with LLM assistance versus 34.5 minutes without. The LLM-assisted condition was 7.9 minutes faster on average (95% CI -1.5 to 17.3; P = .099). Extraction accuracy was 100% in both conditions, and no adverse events were reported.
Conclusions:
LLM assistance might reduce data extraction time by ∼23% (7.9 minutes per article; 95% CI -1.5 to 17.3 minutes) with no observed loss of accuracy. Although statistical uncertainty remains, LLM integration may offer practical value for rapid evidence synthesis during public health emergencies as tools and prompting strategies mature.
