Related Experiment Video
Updated: May 27, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Automating collateral histories in dementia: Development and proof‑of‑concept evaluation of the LUMEN conversational
Judith R Harrison1, Alexander Robertson2, Song Ling Tang3
1Translational and Clinical Research Institute, Newcastle University, Newcastle, UK.
Background:
Collateral histories from carers are central to dementia diagnosis but are often collected inconsistently and variably documented. With rising demand on memory services and the emergence of disease-modifying therapies requiring timely diagnosis, there is increasing need for structured and efficient assessment approaches. Conversational AI powered by large language models (LLMs) may support standardised collateral history acquisition while maintaining clinician oversight. We developed LUMEN, a stakeholder-informed prototype designed to generate structured collateral summaries for clinical review.
Methods:
A five-stage patient, public and professional involvement programme (approximately 232 participants) co-designed the question set, interface and outputs. Seven open-source LLMs were benchmarked; Qwen3-30B-A3B was selected to generate structured summaries from interview transcripts. Six clinician-authored vignettes representing Alzheimer's disease, dementia with Lewy bodies, vascular dementia, frontotemporal dementia, mild cognitive impairment and normal cognition were used to generate 54 synthetic dialogues (27 clinician role-played, 27 GPT-4 generated). Diagnostic categories were assigned using a deterministic rule-based rubric applied to structured summaries. Two clinicians independently rated each dialogue. Outcomes included exploratory evaluation of alignment with diagnostic categories measured by area under the receiver operating characteristic curve (AUROC) and Cohen's κ, and System Usability Scale (SUS) scores.
Results:
In this small synthetic vignette-based dataset, macro-average AUROC was 0.95; these values reflect performance under closed-loop proof-of-concept conditions rather than real-world diagnostic accuracy. Discrimination was highest for Alzheimer's disease and vascular dementia (AUROC = 1.00 in this synthetic dataset) and lowest for mild cognitive impairment (AUROC = 0.77). Agreement between categories assigned by the rule-based rubric and averaged clinician ratings was κ = 0.88 (95% CI 0.83-0.93). Mean SUS score was 78.1/100.
Conclusions:
In a small, closed-loop synthetic proof-of-concept dataset, this LLM-assisted, rubric-based pipeline showed that structured summaries could be processed reproducibly by the rubric and separated diagnostic categories under controlled conditions. These findings do not show real-world diagnostic performance. Further evaluation is required to determine clinical usefulness, robustness and workflow impact.
Related Concept Videos
Dementia l: Introduction
Alzheimer Disease l: Introduction
Automatic Processing and Automatic Social Behavior