Related Experiment Video
Updated: Jun 30, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational
Medrxiv : the Preprint Server for Health Sciences
|June 29, 2026
Summary
Automated LLM pipelines, LitScreen and VarEx, streamline systematic reviews for chronic disease epidemiology by reducing screening and extraction time by up to 90%. These tools enhance causal inference by improving the comparability of exposures, outcomes, and covariates across studies.
Area of Science:
- Epidemiology
- Bioinformatics
- Artificial Intelligence
Background:
- Systematic reviews of observational studies are crucial for causal inference in chronic disease epidemiology.
- Challenges include the large volume of literature and inconsistent confounder control.
- Need for transparent, open methods to reduce screening burden and standardize variable extraction.
Purpose of the Study:
- Develop and evaluate modular LLM-based pipelines (LitScreen and VarEx) for automating study screening and variable extraction.
- Focus on observational systematic reviews for use cases like hypertension-to-Alzheimer's Disease and Related Dementias (ADRD) and PTSD-to-self-harm outcomes.
Main Methods:
- An end-to-end workflow using MEDLINE queries, LitScreen for three-phase screening, and VarEx for retrieval-augmented variable extraction.
- LitScreen combines abstract screening, criterion-wise adjudication, and full-text verification.
- VarEx extracts exposures, outcomes, and covariates, classifying them into predefined categories.
Main Results:
- VarEx achieved high performance (e.g., 0.80 precision, 0.79 recall for covariates in hypertension-ADRD).
- LitScreen maintained high recall while significantly reducing screening time (80-90% reduction compared to manual review).
- Performance was validated on multiple datasets, including expert-annotated corpora.
Conclusions:
- A retrieval-augmented LLM framework can automate key aspects of systematic reviews for observational studies.
- These tools generate structured covariate inventories, improving efficiency and reproducibility of evidence synthesis.
- The LLM framework acts as an assistant, augmenting rather than replacing human reviewers.