Related Experiment Video
Updated: Apr 11, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Reproducibility and Robustness of Large Language Models for Mobility Functional Status Extraction
Xingyi Liu1, Muskan Garg1, Eunji Jeon1
1Department of AI and Informatics, Mayo Clinic, Rochester, USA.
Medrxiv : the Preprint Server for Health Sciences
|April 10, 2026
Summary
Evaluating large language models (LLMs) for clinical information extraction (IE) is crucial. This study found that prompt paraphrasing and model choice significantly impact LLM stability, but self-consistency can improve reliability.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Artificial Intelligence in Healthcare
Background:
- Clinical narrative text is rich in patient data but challenging for information extraction (IE) due to variability.
- Large language models (LLMs) show promise for clinical IE, but their reproducibility and robustness are critical for deployment.
- Quantifying LLM stability is essential for reliable clinical applications.
Purpose of the Study:
- To evaluate the reproducibility and robustness of three open-weight large language models (LLMs) for clinical information extraction.
- To assess the impact of prompt variations and model architecture on LLM performance and stability.
- To provide recommendations for improving LLM reliability in clinical settings.
Main Methods:
- Evaluated three distinct LLMs (Llama 3.3, Llama 4, MedGemma) on binary clinical IE tasks related to the ICF mobility framework.
- Quantified intra-prompt reproducibility (repeated sampling) and inter-prompt robustness (paraphrased prompts).
- Measured predictive performance (F1-score) and stability (Fleiss' Kappa), analyzing factor effects with ANOVA.
Main Results:
- Increasing temperature generally decreased LLM agreement, with model-dependent effects.
- Prompt paraphrasing significantly reduced LLM stability, especially for Mixture-of-Experts (MoE) models.
- Self-consistency via majority voting improved stability (Kappa) and often maintained or improved performance (F1-score).
Conclusions:
- LLM reliability in clinical IE is sensitive to prompt design and model architecture.
- Self-consistency offers a practical method to enhance the stability and performance of LLMs for clinical tasks.
- A reproducible framework is presented for evaluating and improving LLM reliability in healthcare.
Related Concept Videos
Improving Translational Accuracy
3.8K
3.8K
Improving Translational Accuracy
15.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.6K

