Related Experiment Video
Updated: Sep 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Open-Source Large Language Models to Extract Symptoms from Clinical Notes
Yunbing Bai1, Wanting Cui1, Joseph Finkelstein1
1Department of Biomedical Informatics, School of Medicine, University of Utah, Salt Lake City, Utah.
Large language models (LLMs) show promise in extracting symptoms, signs, and ICD-10 codes from clinical notes. Llama 3.3-70B demonstrated superior performance in this automated clinical data extraction task.
Area of Science:
- Natural Language Processing (NLP) in Healthcare
- Clinical Informatics
- Artificial Intelligence (AI) in Medicine
Background:
- Automated extraction of clinical information from electronic health records (EHRs) is crucial for improving healthcare efficiency.
- Large language models (LLMs) offer potential for analyzing unstructured clinical text.
- Evaluating LLM performance on specific clinical tasks like symptom extraction and coding is essential.
Purpose of the Study:
- To assess the capability of open-source foundational large language models (LLMs) in extracting symptoms and signs (S&S) and their corresponding ICD-10 codes from clinical notes.
- To compare the performance of different Llama model versions (Llama 3.1-13B, Llama 3.3-70B, Me-Llama-13B) on S&S extraction and ICD-10 code generation.
Main Methods:
- Utilized the public MTSamples dataset, focusing on genitourinary conditions.
- Manually annotated a subset of the dataset for ground truth comparison.
- Evaluated three Llama model versions on S&S extraction and ICD-10 code generation tasks, measuring consistency, runtime, and performance metrics (recall, precision).
Main Results:
- Llama 3.3-70B exhibited the best overall performance among the tested models.
- This model achieved a fast runtime and high consistency.
- For S&S extraction, Llama 3.3-70B reported an average recall of 0.87 and precision of 0.71.
- For ICD-10 code generation, it achieved an average recall of 0.71 and precision of 0.54.
Conclusions:
- Open-source LLMs, particularly Llama 3.3-70B, are effective tools for automated extraction of symptoms, signs, and ICD-10 codes from clinical notes.
- The findings suggest significant potential for LLMs to streamline clinical data processing and coding in healthcare settings.
- Further research can explore optimizing LLM performance for diverse clinical datasets and complex coding requirements.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Statistical Software for Data Analysis and Clinical Trials
Improving Translational Accuracy