Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluation of Large Language Model Performance in Assessing Health Economic Study Quality
Chen Dun1,2, Cody J Couperus3, Seohu Lee1
1Biomedical Informatics and Data Science Johns Hopkins University School of Medicine, Baltimore, Maryland.
Introduction:
Economic evaluations are essential for informed healthcare decision-making but often face challenges due to inconsistent reporting and methodological complexity. Large Language Models (LLMs) offer a scalable alternative for evaluating adherence to such standards. Building on Hileas, a previously developed tool, this study assesses the accuracy of LLM-generated evaluations compared with human reviewers, aiming to quantify reliability, identify limitations, and advance automated, but assistive quality assessment methods in health economic research.
Methods:
In all, 110 peer-reviewed economic evaluation papers were evaluated using the CHEERS checklist through structured LLM prompts and scored by 2 human reviewers on a 0-4 ordinal scale. Interrater agreement and LLM performance were measured using Cohen's kappa, sensitivity, specificity, and area under the curve. LLM outputs were compared against human consensus ratings, and usability of the review platform was assessed with the System Usability Scale.
Results:
Among 2860 item-level evaluations, 25.3% showed disagreement between human reviewers, with generally low interrater reliability (kappa=-0.07 to 0.43). Compared with human consensus, the LLM achieved 72.3% to 94.7% agreement, with areas under the curve up to 0.96 but variable performance across checklist items. At the paper level, LLM-assigned CHEERS scores (median, 17) were consistently lower than human-reviewed scores (median, 18-21).
Conclusion:
This study demonstrated an exploratory proof-of-concept application of LLMs to research quality evaluation. Our results suggests that the LLM was generally able to provide well-reasoned evaluations that closely aligned with human assessments, although with some limitations in fully supporting its judgments.
More Related Videos
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Bias in Epidemiological Studies
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...
Analysis of Population Pharmacokinetic Data

