Related Experiment Video
Updated: Jun 5, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using large language models to enhance clinically-driven missing data recovery algorithms in electronic health
Sarah C Lotspeich1, Abbey N Collins2, Brian J Wells3
1Department of Statistical Sciences, Wake Forest University, Winston-Salem, NC 27109, United States.
JAMIA Open
|June 4, 2026
Summary
Algorithms using large language models (LLMs) can recover missing electronic health record (EHR) data, mimicking expert chart reviews. These clinically-driven tools offer a scalable solution for improving EHR data quality.
Area of Science:
- Health Informatics
- Clinical Data Management
- Artificial Intelligence in Healthcare
Background:
- Electronic health record (EHR) data frequently contain missing values and errors, impacting data quality and clinical research.
- Traditional chart reviews are effective but resource-intensive, limiting their application to large patient cohorts.
- Previous work introduced a roadmap protocol using auxiliary diagnoses to impute missing EHR data.
Purpose of the Study:
- To evaluate the accuracy and scalability of a roadmap-driven algorithm for recovering missing EHR data.
- To compare the performance of LLM-enhanced roadmaps against traditional chart reviews.
- To assess the feasibility of applying these algorithms to large-scale EHR datasets.
Main Methods:
- Developed and refined roadmap algorithms using International Classification of Diseases, 10th revision (ICD-10) codes.
- Iteratively enhanced roadmaps with large language models (LLMs) and clinical expertise to identify auxiliary diagnoses.
- Validated algorithm performance against expert chart reviews for 100 patients and tested scalability on 1000 patients from an extensive EHR.
Main Results:
- Expert chart reviews recovered 12% (49/413) of missing EHR values in 100 patients.
- LLM-enhanced roadmap algorithms recovered 20%-22% (83-89/413) of missing values.
- A clinician-approved LLM-enhanced algorithm recovered 18% (73/413) of missing values, balancing expansion and clinical relevance.
- Application to 1000 patients increased the median non-missing EHR values per patient from 6 to 7.
Conclusions:
- Clinically-driven algorithms, augmented by LLMs, can accurately recover missing EHR data, comparable to manual chart reviews.
- These algorithms demonstrate scalability for application to large EHR datasets.
- Future work could extend these methods to address other data quality issues in EHRs.