Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy of LLMs to retrieve numeric data for meta-analysis in dentistry
Vito Carlo Alberto Caponio1, Alejandro I Lorenzo-Pouso2, Marco Magalhaes3
1Department of Life Sciences, Health and Health Professions, Link Campus University, Via del Casale Di San Pio V 44, 00165, Rome, Italy; ORALMED Research Group, Department of Dental Clinical Specialties, School of Dentistry, Complutense University, 28040 Madrid, Spain.
Large language models (LLMs) show high accuracy in extracting single numeric outcomes for dental systematic reviews and meta-analyses (SRMA). However, omission errors increase at higher data aggregation levels, limiting their independent use.
Area of Science:
- Dental research
- Evidence synthesis
- Artificial intelligence in healthcare
Background:
- Systematic reviews and meta-analyses (SRMA) are crucial for evidence-based dentistry.
- Data extraction in SRMA is time-consuming and prone to errors.
- Large language models (LLMs) offer potential for automating SRMA data extraction.
Purpose of the Study:
- To evaluate the accuracy of four LLMs (DeepSeek v3 R1, Claude 3.5 Sonnet, ChatGPT-4o, Gemini 2.0-flash) in extracting primary numeric outcomes from dental research.
- To compare the performance of different LLMs in data extraction for SRMA.
Main Methods:
- LLMs were queried via APIs using default settings and a SMART-format prompt.
- Data extraction accuracy was assessed at sub-outcome, outcome, and study levels.
- Errors were categorized as hallucinations, missed data, or omissions.
Main Results:
- Overall extraction accuracy was high at the sub-outcome level.
- Gemini 2.0-flash performed significantly worse than other models (p < 0.01).
- Claude 3.5 Sonnet and DeepSeek-v3 R1 demonstrated superior accuracy and lower omission rates in full-text extraction.
Conclusions:
- LLMs show significant potential for data extraction in dental SRMA but have limitations.
- Accuracy varies between models, and cost does not correlate with performance.
- Standardized outcome reporting and accurate, lower-cost LLMs can improve evidence synthesis efficiency.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Regression Toward the Mean
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Teeth
In the bud stage, the tooth germ (an aggregation of cells) starts to form in the developing jawbone. During the cap stage, the tooth germ differentiates into enamel organ, dental papilla, and dental sac, which will later develop into the tooth's enamel, dentin...

