Related Experiment Video
Updated: Sep 18, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing large language models for acute heart failure classification and information extraction from French
Adrien Bazoge1, Matthieu Wargny2, Pacôme Constant Dit Beaufils3
1Nantes Université, CHU Nantes, Pôle Hospitalo-Universitaire 11: Santé Publique, Clinique des données, INSERM, CIC 1413, F-44000, Nantes, France; Nantes Université, École Centrale Nantes, CNRS, LS2N, UMR 6004, F-44000, Nantes, France.
Large language models (LLMs) can identify acute heart failure (AHF) hospitalizations from clinical notes. A specialized French biomedical model outperformed a general LLM, though the general model excelled at extracting specific quantitative data.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing for Healthcare
- Clinical Data Extraction
Background:
- Acute heart failure (AHF) diagnosis is challenging due to reliance on unstructured electronic health record (EHR) data.
- Extracting accurate clinical information from free-text notes is crucial for understanding AHF.
Purpose of the Study:
- To evaluate large language models (LLMs) for automated identification of AHF hospitalizations and extraction of clinical details.
- To compare the performance of a general-purpose LLM (Qwen2-7B) against a French biomedical LLM (DrLongformer).
Main Methods:
- Utilized clinical notes from Nantes University Hospital for model training and evaluation.
- Employed supervised fine-tuning and in-context learning (few-shot, chain-of-thought prompting).
- Conducted an ablation study to assess the impact of data volume and annotation characteristics.
Main Results:
- DrLongformer achieved higher performance in AHF classification (F1=0.878) and most information extraction tasks compared to Qwen2-7B (F1=0.80).
- Qwen2-7B demonstrated superior performance in extracting quantitative outcomes (e.g., weight, BMI) after fine-tuning.
- Model performance significantly correlated with training data volume, with diminishing returns after 250 documents; longer annotations were detrimental.
Conclusions:
- Biomedical-specific LLMs show strong potential for accurate AHF identification and data extraction from EHRs.
- General-purpose LLMs can be effective, particularly for specific quantitative data extraction, with fine-tuning.
- On-premise deployable small language models offer a viable solution for improving clinical data collection and AHF symptom identification within hospital systems.
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Heart Failure III: Clinical Manifestations
Heart Failure V: Medical Management
Heart Failure Drugs: β-Blockers
Heart Failure Drugs: Diuretics
Heart Failure I: Introduction

