Related Experiment Video
Updated: Jul 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can large language models approximate the results of meta-analyses in critical care? A meta-research study
Michael Pratte1, Shawn Thirukumar2, Caseng Zhang3
1University of Toronto, Interdivisional Department of Critical Care, Canada.
Background:
Large language models (LLMs) are capable of processing extensive textual data and synthesizing evidence to answer complex clinical questions. The labor-intensive nature of systematic reviews with meta-analyses (SRMAs) present a unique opportunity to evaluate the utility of LLMs as a novel method for evidence synthesis.
Objective:
This study assessed the ability of OpenAI's o3 DeepResearch model to approximate the direction of effect, magnitude of effect and certainty of evidence for clinical questions addressed by published meta-analyses in top critical care medicine journals.
Methods:
We constructed standardized prompts based on the PICO (Population, Intervention, Comparator, Outcome) from a convenience sample of 23 systematic reviews with meta-analyses published in high-impact critical care journals. The LLM's estimates of effect size and certainty of evidence ratings were compared to those reported in the original SRMAs.
Results:
The LLM demonstrated a concordance rate of 83 % (19 of 23 studies) for the magnitude of effect size and 91 % (21 of 23 studies) for the direction of effect. Concordance for certainty of evidence was also 91 %. Discrepancies were due to differences in study selection between the LLM and SRMAs, rather than model hallucination or misinterpretation.
Conclusions:
LLMs show promise as a new tool for rapid evidence synthesis in critical care, with outputs comparable to traditional meta-analyses in many cases. While not a replacement for systematic reviews, LLMs may enhance clinical decision-making, perform rapid evidence synthesis, and streamline future research workflows.
