Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the methodologic quality of systematic reviews using generative large language models.
Bowen Yao1,2, Onuralp Ergun1,2, Maylynn Ding2
1Minneapolis VA Healthcare System, Minneapolis, MN, United States.
Generative large language models (LLMs) show potential for assessing systematic review (SR) quality. With specific instructions, GPT achieved 93% accuracy in quality assessment, indicating efficient and reliable evaluation capabilities.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Urology Research
Background:
- Assessing the methodological quality of systematic reviews (SRs) is crucial for evidence-based medicine.
- Generative large language models (LLMs) offer potential for automating complex analytical tasks.
Purpose of the Study:
- To evaluate the accuracy of generative LLMs in assessing the methodological quality of urological SRs.
- To compare LLM-based quality assessment with human expert penilaian.
Main Methods:
- 114 urological SRs were assessed by human experts and a customized GPT model.
- GPT underwent three zero-shot assessment iterations and an enhanced trial using chain-of-thought prompting.
- Performance metrics included accuracy, sensitivity, specificity, and F1 score against human judgments.
Main Results:
- GPT achieved 75% overall congruence with human reviewers, with 77% for critical criteria.
- The average F1 score was 0.66, and internal validity was high at 85%.
- Enhanced prompting improved critical criteria congruence to 91% and overall accuracy to 93%.
Conclusions:
- Generative LLMs demonstrate a promising capacity for efficient and accurate quality assessment of SRs in urology.
- LLM-based tools can potentially streamline the review process and support evidence synthesis.
More Related Videos
Related Concept Videos
Qualitative Analysis
There are two main approaches to qualitative analysis:...
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Improving Translational Accuracy
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

