Related Experiment Video
Updated: Jun 29, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for automated PRISMA 2020 adherence checking
Yuki Kataoka1, Ryuhei So2, Masahiro Banno3
1Center for Postgraduate Clinical Training and Career Development, Nagoya University Hospital, Nagoya, Aichi, Japan; Center for Medical Education, Graduate School of Medicine, Nagoya University, Nagoya, Aichi, Japan; Scientific Research Works Peer Support Group (SRWS-PSG), Osaka, Japan; Department of Internal Medicine, Kyoto Min-iren Asukai Hospital, Kyoto, Japan; Department of Healthcare Epidemiology, Kyoto University Graduate School of Medicine / School of Public Health, Kyoto, Japan; Department of International and Community Oral Health, Tohoku University Graduate School of Dentistry, Sendai, Miyagi, Japan.
Background:
Evaluating adherence to PRISMA 2020 guideline remains a burden in the peer review process. However, there is a lack of shareable benchmarks for evaluating large language model (LLM) performance in this task.
Methods:
We constructed a copyright-aware benchmark of 108 Creative Commons-licensed systematic reviews. We first conducted parameter optimization using five SRs from the Suda dataset, then compared five checklist input formats (Markdown, JSON, XML, plain text, and manuscript-only control) using ten development-phase LLMs on ten further SRs from the Suda dataset, and finally validated the locked Markdown pipeline using nineteen LLMs on ten SRs from the Tsuge dataset as additional frontier models became available during the study period.
Results:
Supplying structured PRISMA 2020 checklists yielded 78.7-79.7% accuracy versus 45.2% for manuscript-only input, with paired aggregate analyses showing that structured formats outperformed manuscript-only input while structured formats did not differ significantly from one another. In the validation sample, accuracy ranged from 68.5% to 86.0% with distinct sensitivity-specificity trade-offs. Using Qwen3-Max on the full dataset (n = 120), we achieved 95.1% sensitivity and 49.3% specificity.
Discussion:
Structured checklist provision substantially improves LLM-based PRISMA assessment. However, given the observed proportion of false positives, human expert verification remains essential before editorial decisions.