Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Not Ready for Prime Time: Limitations of a Retrieval-Augmented Generation Large Language Model in Assessing Risk of
Samuel A Beber1, Katherine D Groff1, Tyler R Mange1
1Division of Pediatric Orthopaedic Surgery, Hospital for Special Surgery, New York, NY, USA.
Retrieval-augmented generation large language models (RAG-LLMs) showed low accuracy in assessing pediatric orthopaedic studies. Caution is advised when using AI tools for research until further validation is achieved.
Area of Science:
- Artificial Intelligence in Medical Research
- Systematic Review Augmentation
- Pediatric Orthopaedics Literature Analysis
Background:
- Large language models (LLMs) are increasingly used to enhance systematic reviews.
- LLMs can "hallucinate"; retrieval-augmented generation (RAG) limits this by using provided sources.
- This study assessed a RAG-LLM's accuracy in quality appraisal of pediatric orthopaedic observational studies.
Purpose of the Study:
- To evaluate the accuracy and reliability of a RAG-LLM (NotebookLM) for quality assessment of observational studies in pediatric orthopaedics.
- To compare the RAG-LLM's performance against manual review (ground truth).
Main Methods:
- Two existing systematic reviews of pediatric orthopaedic observational studies were included.
- NotebookLM assessed study quality using the Newcastle-Ottawa Scale (NOS).
- Intraclass correlation coefficients (ICC) measured agreement between RAG-LLM scores and manual review scores.
Main Results:
- Overall agreement between RAG-LLM instances and manual review was moderate (ICC(2,k) = 0.69).
- Agreement between individual RAG-LLM scores and manual review was poor (ICC(2,1) ranging from 0.081 to 0.27).
- Percent agreement ranged from 14.8% to 29.6%.
Conclusions:
- The RAG-LLM (NotebookLM) demonstrated low reliability and accuracy in quality assessment.
- Researchers should exercise caution when using AI tools like RAG-LLMs in research.
- Independent critical appraisal by researchers remains essential for new studies.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Bias in Epidemiological Studies
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Naturalistic Observations
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Confounding in Epidemiological Studies