Related Experiment Video
Updated: Sep 10, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Use of artificial intelligence to support the assessment of the methodological quality of systematic reviews
Manuel Marques-Cruz1, Filipe Pinto2, Rafael José Vieira3
1Faculty of Medicine, MEDCIDS - Department of Community Medicine, Information and Health Decision Sciences, University of Porto, Porto, Portugal; CINTESIS@RISE - Health Research Network, University of Porto, Porto, Portugal; Public Health Unit Marão e Douro Norte, Local Health Unit Trás-os-Montes e Alto Douro, Vila Real, Portugal.
Objectives:
Published systematic reviews display a heterogeneous methodological quality, which can impact decision-making. Large language models (LLMs) can support and make the assessment of the methodological quality of systematic reviews more efficient, aiding in the incorporation of their evidence in guideline recommendations. We aimed to develop an LLM-based tool for supporting the assessment of the methodological quality of systematic reviews.
Methods:
We assessed the performance of 8 LLMs in evaluating the methodological quality of systematic reviews. In particular, we provided 100 systematic reviews for eight LLMs (five base models and three fine-tuned models) to evaluate their methodological quality based on a 27-item validated tool (Reported Methodological Quality (ReMarQ)). The fine-tuned models had been trained with a different sample of 300 manually assessed systematic reviews. We compared the answers provided by LLMs with those independently provided by human reviewers, computing the accuracy, kappa coefficient and F1-score for this comparison.
Results:
The best performing LLM was a fine-tuned GPT-3.5 model (mean accuracy = 96.5% [95% CI = 89.9%-100%]; mean kappa coefficient = 0.90 [95% CI = 0.71-1.00]; mean F1-score = 0.91 [95% CI = 0.83-1.00]). This model displayed an accuracy >80% and a kappa coefficient >0.60 for all individual items. When we made this LLM assess 60 times the same set of systematic reviews, answers to 18 of 27 items were always consistent (ie, were always the same) and only 11% of assessed systematic reviews showed inconsistency.
Conclusion:
Overall, LLMs have the potential to accurately support the assessment of the methodological quality of systematic reviews based on a validated tool comprising dichotomous items.
Related Concept Videos
Quality Assurance
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Data Validation
Key parameters for method validation include:
Qualitative Analysis
There are two main approaches to qualitative analysis:...
The Scientific Method in Nursing Process
When using research findings to change practice, one must understand the process used to guide a study. The scientific method is a systematic, step-by-step process that supports the data's validity, reliability, and generalizability. As a result, findings can be...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...

