Related Experiment Video
Updated: May 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language model agents can use tools to perform clinical calculations
Alex J Goodell1, Simon N Chu2, Dara Rouholiman3
1Department of Anesthesiology, Pain, and Perioperative Medicine, Stanford University School of Medicine, Stanford, CA, USA. agoodell@stanford.edu.
Large language models (LLMs) struggle with medical calculations. Integrating task-specific tools significantly improved accuracy, showing potential for safer clinical AI applications.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Computational Medicine
Background:
- Large language models (LLMs) demonstrate expertise in medical question answering.
- LLMs exhibit significant limitations, including hallucinations and arithmetic errors.
- Current LLM performance in clinical calculations hinders integration into healthcare workflows.
Purpose of the Study:
- To evaluate the accuracy of ChatGPT in performing medical calculations.
- To assess the impact of agentic augmentation strategies on LLM performance in medical calculations.
- To determine the efficacy of task-specific tools in improving LLM reliability for clinical computations.
Main Methods:
- ChatGPT's performance was evaluated on 48 distinct medical calculation tasks.
- Three augmentation methods were tested: retrieval-augmented generation, a code interpreter, and task-specific tools (OpenMedCalc).
- Over 10,000 trials were conducted to assess model performance with and without augmentation.
Main Results:
- ChatGPT produced incorrect responses in one-third of initial trials.
- Models augmented with task-specific tools showed substantial improvements.
- LLaMa and GPT-based models with tools reduced incorrect responses by 5.5-fold and 13-fold, respectively.
Conclusions:
- LLMs face challenges with medical arithmetic, impacting clinical utility.
- Task-specific, machine-readable tools are effective in mitigating LLM calculation errors.
- Integrating specialized tools can enhance the reliability of LLMs for clinical applications.
More Related Videos
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Statistical Software for Data Analysis and Clinical Trials
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...

