Related Experiment Video
Updated: Sep 18, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Risk of Bias Assessment of Diagnostic Accuracy Studies Using QUADAS 2 by Large Language Models.
Daniel-Corneliu Leucuța1, Andrada Elena Urda-Cîmpean1, Dan Istrate1
1Department of Medical Informatics and Biostatistics, Iuliu Hațieganu University of Medicine and Pharmacy, 400349 Cluj-Napoca, Romania.
Large language models (LLMs) show moderate accuracy in assessing risk of bias in diagnostic accuracy studies using QUADAS 2. While not replacing human experts, LLMs can aid systematic reviews with supervision.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Biostatistics
Background:
- Diagnostic accuracy studies are crucial for evaluating medical test performance.
- Risk of bias (RoB) assessment in these studies commonly utilizes the Quality Assessment of Diagnostic Accuracy Studies (QUADAS) tool.
- Evaluating the capabilities of Large Language Models (LLMs) in RoB assessment is a novel area of research.
Purpose of the Study:
- To assess the accuracy of LLMs in evaluating RoB in diagnostic accuracy studies using QUADAS 2.
- To compare LLM performance against human expert assessments.
- To identify specific domains where LLMs excel or falter in RoB evaluation.
Main Methods:
- Four LLMs (ChatGPT 4o, Grok 3, Gemini 2.0 Flash, DeepSeek V3) were employed.
- Ten diagnostic accuracy studies were selected for assessment.
- Human experts and LLMs independently applied the QUADAS 2 tool to each study.
Main Results:
- LLMs achieved a mean accuracy of 72.95% across 110 signaling questions.
- Grok 3 (74.45%) and ChatGPT 4o (73.15%) demonstrated higher accuracy than DeepSeek V3 (70.00%) and Gemini 2.0 Flash (67.27%).
- Highest accuracy was observed in 'flow and timing', followed by 'index test', 'patient selection', and 'reference standard' domains, with documented reasoning errors.
Conclusions:
- LLMs demonstrate moderate capability in RoB assessment for diagnostic accuracy studies.
- Current LLM performance is insufficient to replace expert clinical and methodological judgment.
- LLMs show potential as supplementary tools in systematic reviews, contingent on mandatory human oversight.
Related Concept Videos
Detection of Gross Error: The Q Test
Bias in Epidemiological Studies
Improving Translational Accuracy
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Receiver Operating Characteristic Plot

