Related Experiment Video
Updated: Jun 19, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models for full-text methods assessment: a case study on mediation analysis
Wenqing Zhang1, Trang Nguyen1,2, Elizabeth A Stuart1,2
1Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD 21205, United States.
Summary
Large language models (LLMs) show promise in assisting with systematic reviews by matching human expert performance on methodological assessments. However, advanced LLMs still lag on complex inference tasks, suggesting a collaborative human-AI approach.
Area of Science:
- Psychiatry and Psychology Research
- Artificial Intelligence in Scientific Literature Review
Background:
- Systematic reviews are crucial for evidence synthesis but are labor-intensive, especially for extracting detailed methodological information from full-text articles.
- Assessing causal assumptions and methodological best practices in psychiatry and psychology studies requires expert-level review.
Purpose of the Study:
- To evaluate the performance of large language models (LLMs) in conducting full-text methodological reviews of mediation analysis studies.
- To compare LLM capabilities against human expert-level review on key causal assumptions and best practices.
Main Methods:
- Six LLMs (ChatGPT, Claude, Gemini) were tested on 180 full-text mediation analysis articles previously reviewed by methodologists.
- LLMs assessed 14 binary methodological criteria, with performance measured against expert consensus using accuracy, precision, recall, F1, AUC, and PR-AUC.
Main Results:
- LLM performance strongly correlated with human reviewers (accuracy correlation 0.71, F1 correlation 0.95).
- Advanced LLMs achieved near-human accuracy on explicit features but were up to 15% less accurate on inference-intensive tasks.
- Model accuracy decreased with longer documents, and common errors involved overinterpretation and misinterpretation of technical terms.
Conclusions:
- Findings support a criterion-specific human-AI collaboration strategy for full-text methodological assessment.
- A reproducible framework is provided for future LLM testing in evidence synthesis settings.
Related Concept Videos
Methods of Medium Optimization
Optimizing growth media enhances microbial proliferation and maximizes product yield. Statistical experimental design methodologies provide structured and reproducible approaches, offering progressively higher levels of robustness and efficiency.The One-Factor-at-a-Time (OFAT) MethodThe One-Factor-at-a-Time (OFAT) method involves adjusting a single variable while keeping all others constant. However, it cannot detect interactions between variables, often leading to suboptimal outcomes when...
Regression Analysis
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Two-Way ANOVA
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...