Related Experiment Video
Updated: Jun 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for full-text methods assessment: a case study on mediation analysis
Wenqing Zhang1, Trang Nguyen1,2, Elizabeth A Stuart1,2
1Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD 21205, United States.
Objective:
Systematic reviews remain labor-intensive, particularly when extracting methodological details from full texts. Using mediation analysis as a case study, we evaluated whether large language models (LLMs) can match human-expert-level full-text methodological review on key causal assumptions (eg, no unmeasured confounding, temporal ordering) and best practices (eg, sensitivity analyses, interaction assessments, covariate adjustment) for psychiatry and psychology studies.
Materials And Methods:
We evaluated 6 LLMs from 3 major families (ChatGPT-4o-mini/4o/o3/5, Claude Sonnet 4, Gemini 2.5 Flash) on 180 full-text mediation analysis articles from 2013 to 2018 previously reviewed by expert methodologists. LLMs assessed 14 binary methodological criteria ranging from straightforward checks (eg, whether the exposure was randomized) to nuanced assessments (eg, whether the temporal ordering between mediator and outcome was established). Performance was benchmarked against expert consensus labels and individual reviewers using accuracy, precision, recall, F1, AUC, and PR-AUC.
Results:
LLM performance strongly correlated with human reviewers across methodological criteria (accuracy correlation 0.71; F1 correlation 0.95), indicating tasks difficult for humans were likewise challenging for models. Advanced LLMs achieved near-human accuracy on explicit methodological features but lagged behind top reviewers by up to 15% on inference-intensive tasks. Longer documents reduced model accuracy. Common model errors include overinterpreting on linguistic cues and colloquial use of technical terms.
Discussion And Conclusion:
Our findings support a criterion-specific human-AI collaboration strategy for full-text methodological assessment and provide a reproducible framework for future testing in other evidence-synthesis settings.
Related Concept Videos
Methods of Medium Optimization
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Two-Way ANOVA
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...