Related Experiment Video
Updated: Jul 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Generating Alzheimer's narratives using large language models
Paula Andrea Perez-Toro1,2, Mahmoud Almizel3, Elmar Nöth3
1Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany. paula.andrea.perez@fau.de.
Background:
Analyzing semi-spontaneous speech is a promising direction for supporting Alzheimer's disease (AD) assessment, yet progress is limited by the scarcity of annotated clinical data. Large Language Models (LLMs) offer new opportunities to generate synthetic narratives that may resemble speech patterns of both patients with AD and healthy controls during cognitive evaluation tasks such as the Cookie Theft Picture description.
Methods:
This study evaluates whether models including GPT, T5/Flan-T5, LLaMA, Mistral, and Qwen can generate clinically plausible picture-description narratives under two configurations: Human-to-Bot, where an LLM responds directly to real interviewer prompts, and Bot-to-Bot, where two LLMs simulate both interviewer and participant roles. Models were fine-tuned on transcripts from the DementiaBank Pitt Corpus and assessed using lexical and semantic metrics, as well as human expert ratings. Generated narratives were further used to augment training data for an AD vs. healthy control classifier based on BERT embeddings and an MLP architecture.
Results:
LLMs differed substantially in their ability to reproduce clinically meaningful and semantically coherent narratives of patient-interviewer interactions. Mistral, LLaMA, and Qwen achieved the strongest automatic evaluation metrics, e.g., BERTScores above 0.90 in the Human-to-Bot condition-and produced narratives rated by human experts as fluent, plausible, and diagnostically informative. When combining real and synthetic narratives for classifier training, the highest F1-score reached 0.84, outperforming models trained on real data alone (F1 = 0.74). Synthetic data generated in Human-to-Bot settings contributed most to diagnostic improvements, whereas Bot-to-Bot interactions exhibited greater variability and reduced clinical realism.
Conclusion:
LLMs can generate high-quality synthetic narratives that enhance downstream AD classification and show promising clinical plausibility in cognitive assessment contexts. Incorporating LLM-generated data provides a scalable strategy for mitigating data scarcity in dementia research. Future work should focus on improving fully synthetic dialogue quality, expanding multilingual capabilities, and refining evaluation frameworks to better capture clinically relevant linguistic features.
Related Concept Videos
Alzheimer Disease l: Introduction
Alzheimer's Disease: Overview
The clinical diagnosis of AD hinges on the presence of memory and other cognitive impairments. Biomarkers, such as changes in Aβ and tau...
Alzheimer Disease ll: Pathophysiology
Alzheimer's Disease: Treatment
Dementia l: Introduction
Language and Cognition
