Related Experiment Video
Updated: Jun 10, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
508
Automated Transformation of Unstructured Cardiovascular Diagnostic Reports into Structured Datasets Using
Sumukh Vasisht Shankar1, Lovedeep S Dhingra1, Arya Aminorroaya1
1Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, New Haven, CT, USA.
Medrxiv : the Preprint Server for Health Sciences
|October 17, 2024
Summary
This study introduces HeartDx-LM, a novel large language model (LLM) system for automatically extracting valuable data from unstructured echocardiogram reports. This innovation enhances cardiovascular research and patient care by making complex data accessible.
Area of Science:
- Cardiovascular medicine
- Artificial intelligence
- Medical informatics
Background:
- Cardiovascular diagnostic testing generates rich data often trapped in unstructured reports.
- Manual data abstraction limits real-time patient care and research applications.
- Automating data extraction from echocardiogram reports is crucial for clinical utility.
Purpose of the Study:
- To develop and validate a novel method for automated data extraction from unstructured transthoracic echocardiogram (TTE) reports.
- To leverage large language models (LLMs) for transforming free-text clinical narratives into structured, computable data.
- To improve the accessibility of cardiovascular data for patient care and research.
Main Methods:
- A two-step process using generative (Llama2 70b) and interpretative (Llama2 13b) LLMs was employed.
- Llama2 70b generated varied TTE report formats from real-world data.
- Llama2 13b was fine-tuned to extract 18 key echocardiographic fields from free-text narratives, creating the HeartDx-LM model.
Main Results:
- The HeartDx-LM model achieved high accuracy in extracting data from contemporary (98.7%) and older (87.1%) echocardiogram reports.
- External validation on MIMIC-III (87.9%) and MIMIC-IV (91.3%) datasets demonstrated robust performance.
- Stable performance was achieved with as few as 500 annotated reports for fine-tuning.
Conclusions:
- A novel method using paired LLMs automates the extraction of unstructured echocardiographic reports into tabular datasets.
- This scalable strategy transforms unstructured reports into computable elements.
- The approach promises to enhance cardiovascular care quality and facilitate research.

