Related Experiment Video
Updated: May 8, 2026

06:51
Measuring the Complete-arch Distortion of an Optical Dental Impression
Published on: May 30, 2019
8.1K
Performance of large language models conducting systematic review tasks in prosthodontics
Rata Rokhshad1, Parisa Motie2, Mohammadjavad Shirani3
1Resident, Department of Pediatric Dentistry, School of Dentistry, Loma Linda University, Loma Linda, Calif.
The Journal of Prosthetic Dentistry
|March 13, 2026
Summary
Large language models (LLMs) show promise for improving systematic reviews (SRs), particularly in data extraction. However, their variable performance across tasks necessitates continued human oversight for accurate and reliable results.
Area of Science:
- Bibliometrics
- Artificial Intelligence in Research
- Systematic Review Methodology
Background:
- Systematic reviews (SRs) are crucial for evidence synthesis but are often time-consuming and resource-intensive.
- The potential of large language models (LLMs) to enhance the efficiency of SR processes remains largely unexplored.
Purpose of the Study:
- To evaluate the accuracy and reliability of four LLMs (GPT-4, Gemini, Claude, Elicit) in performing key SR tasks.
- Assess LLM performance in full-text screening, data extraction, and risk of bias assessment over time.
Main Methods:
- A systematic search identified 59 articles for screening and 31 for data extraction.
- A 3-pronged prompting strategy (persona-based, few-shot learning, PICO criteria) was employed.
- LLM performance was measured using accuracy, precision, F1-score, sensitivity, specificity, data extraction quality scores, and Cohen kappa for risk of bias agreement.
Main Results:
- Claude demonstrated the highest sensitivity (97%) in full-text screening; Claude and Elicit achieved 86% accuracy and 87% F1-scores.
- GPT-4 excelled in data extraction with median scores of 5.0; Claude and Gemini showed comparable performance.
- Risk of bias assessment agreement with experts ranged from 55% to 90% across different criteria.
Conclusions:
- LLMs offer potential for increasing the efficiency of systematic reviews, especially in data extraction.
- Variable performance across different LLMs and SR tasks underscores the need for human oversight.

