Related Experiment Video
Updated: Sep 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Reasoning Capabilities of Large Language Models for Medical Coding and Hospital Readmission Risk
Parvati Naliyatthaliyazchayil1, Raajitha Muthyala1, Judy Wawira Gichoya2
1Department of Biomedical Engineering and Informatics, Luddy School of Informatics, Computing and Engineering, Indiana University Indianapolis, 535 W Michigan Street, Indianapolis, IN, 46202, United States, 1 317 274 0439.
Large language models show moderate success in zero-shot clinical diagnosis and risk prediction but struggle with medical coding. Task-specific fine-tuning and human oversight are crucial for reliable healthcare applications.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
- Natural Language Processing in Medicine
Background:
- Large language models (LLMs) demonstrate potential in healthcare for clinical reasoning and decision support.
- However, their reliability for critical tasks like diagnosis, coding, and risk prediction without specific training is uncertain.
Purpose of the Study:
- To evaluate and compare the zero-shot performance of reasoning and non-reasoning LLMs on clinical diagnosis, ICD-9 code prediction, and hospital readmission risk stratification.
- To assess LLMs' potential as general-purpose clinical decision support tools.
Main Methods:
- Utilized the MIMIC-IV dataset, analyzing 300 hospital discharge summaries.
- Standardized, zero-shot prompts were used across reasoning and non-reasoning LLMs, incorporating rationale elicitation for transparency.
- Performance was measured using F1-scores and correctness percentages, with statistical analysis.
Main Results:
- LLMs showed moderate success in zero-shot diagnosis and risk prediction but significantly underperformed in ICD-9 code prediction.
- OpenAI-O3 demonstrated superior performance in diagnosis and ICD-9 coding among tested models.
- Reasoning models offered marginal performance gains and improved interpretability but generated verbose outputs.
Conclusions:
- Current LLMs require task-specific fine-tuning and human-in-the-loop validation for clinical use.
- Medical coding remains a significant challenge for LLMs in zero-shot settings.
- Further research is needed to improve LLM stability, reliability, and performance on diverse clinical data.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Improving Translational Accuracy
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Patient-centered Care