Related Experiment Video
Updated: Jun 21, 2026

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Using LLMs to Interpret Arterial Blood Gases: Comparison of a Novel Math Scratchpad with Different Prompting Methods
Praveen Meka1,2, Christine Tsien Silvers2,3, Bharath Gunapati3
1Dana-Farber Cancer Institute, Boston, MA.
A novel math scratchpad significantly improved large language models' (LLMs) accuracy in interpreting Arterial Blood Gases (ABGs) for clinical decision support systems. This method enhances diagnostic capabilities by overcoming LLMs' calculation limitations.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) show promise in healthcare but struggle with precise calculations, limiting their use in clinical decision support systems (CDSS).
- Accurate interpretation of Arterial Blood Gases (ABGs) requires complex calculations, posing a challenge for current LLM capabilities.
- Existing LLM approaches for medical data analysis often lack the computational precision needed for critical diagnostic tasks.
Purpose of the Study:
- To evaluate a novel method integrating a custom math scratchpad with LLMs for enhanced Arterial Blood Gas (ABG) interpretation.
- To compare the diagnostic accuracy of different LLM prompting strategies, including zero-shot, Retrieval-Augmented Generation (RAG), and a combined approach with a math scratchpad.
- To address the limitations of LLMs in performing domain-specific calculations within a CDSS context.
Main Methods:
- Developed and implemented a novel math scratchpad integrated into an LLM-based CDSS.
- Compared three interpretation methods: zero-shot prompting, prompt engineering with RAG, and a hybrid approach combining the math scratchpad, RAG, and prompt engineering.
- Evaluated the methods on a dataset of 50 Arterial Blood Gas (ABG) results.
Main Results:
- The LLM-integrated CDSS using the novel math scratchpad (Method 3) achieved an accuracy of 86% (43/50), with a 95% confidence interval of 74%-93%.
- Method 2 (RAG and prompt engineering) resulted in 78% accuracy (39/50) [CI 65%-87%].
- Method 1 (zero-shot prompting) demonstrated significantly lower accuracy at 48% (24/50) [CI 35%-61%].
Conclusions:
- The integration of a math scratchpad effectively overcomes the computational limitations of LLMs in ABG interpretation.
- The novel hybrid method significantly enhances the accuracy of LLM-based CDSS for complex clinical data analysis.
- Further validation with real-world patient ABG data is recommended to confirm the clinical utility of this approach.
More Related Videos
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
06:13Author Spotlight: Exploring Olfactory Influences on Corticospinal Excitability - Insights and Innovations in Neurological Research
Published on: January 19, 2024
Related Concept Videos
Blood Studies I: ABG and VBG
Arterial Blood Gas (ABG)
Arterial Blood Gas (ABG) studies are crucial for assessing the lungs' ability to supply oxygen and remove carbon dioxide, reflecting the patient's ventilation status. They also help understand the kidneys' capacity to reabsorb or...
Applications of Integration to Find Blood Flow