Related Experiment Video
Updated: Jun 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Enhancing early detection of cognitive decline in the elderly: a comparative study utilizing large language models in
Xinsong Du1, John Novoa-Laurentiev2, Joseph M Plasek1
1Division of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, MA, 02115, USA; Department of Medicine, Harvard Medical School, Boston, MA, 02115, USA.
An ensemble model combining large language models (LLMs) and traditional machine learning significantly improved the detection of cognitive decline signs in electronic health records (EHRs). This approach enhances diagnostic accuracy by leveraging complementary error profiles.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Informatics
- Natural Language Processing
Background:
- Large language models (LLMs) show promise in healthcare but require evaluation for detecting specific clinical conditions in electronic health records (EHRs).
- This study assesses LLMs for identifying cognitive decline indicators in clinical notes, comparing their performance and error patterns against traditional models.
- Understanding these differences informs strategies for improving LLM performance in clinical settings.
Purpose of the Study:
- To evaluate the effectiveness of large language models (LLMs) in detecting cognitive decline from electronic health record (EHR) clinical notes.
- To compare the error profiles of LLMs (GPT-4, Llama 2) with traditional machine learning models.
- To develop an optimal LLM-based method for identifying cognitive decline and assess the performance of an ensemble model.
Main Methods:
- Analysis of clinical notes from Mass General Brigham (2019 diagnosis of mild cognitive impairment).
- Development and comparison of prompts for GPT-4 and Llama 2 using various approaches (hard prompting, retrieval augmented generation).
- Creation of an ensemble model combining LLMs and traditional models (hierarchical attention network, XGBoost) using a majority vote, evaluated with confusion-matrix-based scores.
Main Results:
- GPT-4 showed superior accuracy and efficiency over Llama 2 but did not surpass traditional models individually.
- The ensemble model significantly outperformed individual models across all metrics (p < 0.01), achieving 90.2% precision, 94.2% recall, and 92.1% F1-score.
- The ensemble model substantially improved precision (from 70%-79% to >90%), with only 3.2% mutual errors across all models, indicating diverse error profiles.
Conclusions:
- LLMs and traditional models trained on local EHR data have distinct error profiles.
- An ensemble approach combining these models is complementary and enhances diagnostic performance for cognitive decline detection.
- Future research should explore integrating LLMs with localized models and domain-specific knowledge for task-specific performance enhancement.
More Related Videos
Related Concept Videos
Language and Cognition
Cognitive Development During Adulthood
Dementia
The progression of dementia is generally gradual....
Cognitive Enhancers: Cholinesterase Inhibitors and NMDA Receptor Antagonists
Alzheimer's Disease: Overview
The clinical diagnosis of AD hinges on the presence of memory and other cognitive impairments. Biomarkers, such as changes in Aβ...

