Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A statistical framework for evaluating the repeatability and reproducibility of large language models
Cathy Shyr1,2,3, Boyu Ren4, Chih-Yuan Hsu2
1Department of Biomedical Informatics, Vanderbilt University Medical Center, 2525 West End Avenue, Nashville, 37203, TN, USA.
Abstract:
A major concern in applying large language models (LLMs) to medicine is their reliability. Because LLMs generate text by sampling the next token (or word) from a probability distribution, the stochastic nature of this process can lead to different outputs even when the input prompt, model architecture, and parameters remain the same. Variation in model output has important implications for reliability in medical applications, yet it remains underexplored and lacks standardized metrics. To address this gap, we propose a statistical framework that systematically quantifies LLM variability using two metrics: repeatability, the consistency of LLM responses across repeated runs under identical conditions, and reproducibility, the consistency across runs under different conditions. Within these metrics, we evaluate two complementary dimensions: semantic consistency, which measures the similarity in meaning across responses, and internal stability, which measures the stability of the model's underlying token-generating process. We applied this framework to medical reasoning as a use case, evaluating LLM repeatability and reproducibility on standardized United States Medical Licensing Examination (USMLE) questions and real-world rare disease cases from the Undiagnosed Diseases Network (UDN) using validated medical reasoning prompts. LLM responses were less variable for UDN cases than for USMLE questions, suggesting that the complexity and ambiguity of real-world patient presentations may constrain the model's output space and yield more stable reasoning. Repeatability and reproducibility did not correlate with diagnostic accuracy, underscoring that an LLM producing a correct answer is not equivalent to producing it consistently. By providing a systematic approach to quantifying LLM repeatability and reproducibility, our framework supports more reliable use of LLMs in medicine and biomedical research.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Reliability and Validity
Random and Systematic Errors
Language and Cognition