Related Experiment Video
Updated: Mar 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations.
Benlu Wang1, Iris Xia1, Yifan Zhang2,3
1Department of Computer Science, Yale University, CT, USA.
Large language models (LLMs) show potential in medicine, but their calculation abilities need better evaluation. This study introduces a new method to assess LLM medical calculations, improving trustworthiness for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in medical tasks but their calculation accuracy is poorly evaluated.
- Existing benchmarks overlook systematic reasoning failures, risking clinical misjudgments.
- Clinical trustworthiness requires a more rigorous assessment of medical calculations.
Purpose of the Study:
- To develop a clinically faithful methodology for evaluating LLM medical calculations.
- To identify and analyze reasoning failures in LLMs performing medical calculations.
- To propose an improved LLM framework for reliable medical computation.
Main Methods:
- Restructured MedCalc-Bench dataset and proposed a step-by-step evaluation pipeline (formula selection, entity extraction, computation).
- Developed an automatic error analysis framework for scalable, explainable diagnostics.
- Introduced MedRaC, a modular agentic pipeline combining retrieval-augmented generation and Python code execution.
Main Results:
- GPT-4o accuracy dropped from 62.7% to 43.6% under granular evaluation, revealing masked errors.
- Automatic error analysis aligned with human expert judgment.
- MedRaC improved LLM accuracy significantly, from 16.35% to 53.19% without fine-tuning.
Conclusions:
- Current medical calculation benchmarks for LLMs are inadequate and may lead to serious clinical errors.
- A granular, step-by-step evaluation and error analysis framework enhances clinical trustworthiness.
- The MedRaC pipeline offers a promising approach for reliable LLM-based medical calculations, advancing safe clinical adoption.
More Related Videos
Related Concept Videos
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Health Literacy
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Methods of Documentation III: PIE

