Related Experiment Video
Updated: Jul 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the utility of large language models for detecting and simulating language dysfunction
Yan Cong1,2, Jiyeon Lee3, Nalin Rajput4
1School of Languages and Cultures, Purdue University, West Lafayette, IN, United States.
Abstract:
This study investigates the utility of large language models (LLMs) in simulating and classifying language dysfunction, with a focus on aphasia. We conducted three studies to evaluate the extent to which LLMs can approximate surface-level linguistic features of individuals with aphasia and whether synthetic data can support classification tasks when real data is limited. In Study 1, we prompted an LLM to generate synthetic utterances and used human rating surveys to evaluate whether the resulting outputs capture surface-level features associated with agrammatic language production. Results showed that approximately half of raters were unable to reliably distinguish AI-generated from human-produced agrammatic speech. This pattern suggests that the synthetic data captured some surface-level features associated with agrammatic utterances; however, low inter-rater agreement indicates substantial rater uncertainty and task ambiguity-likely reflecting the difficulty of making reliable judgments from brief, text-only excerpts-limiting conclusions about clinical fidelity. In Study 2, we assessed the performance of (fine-tuned) LLMs on two binary classification tasks: (1) detecting agrammatic features and (2) classifying utterances based on an LLM-derived surprisal index (negative log-likelihood of a token given its context). We examined models' successes and failures by comparing performance across training conditions and against a classical machine learning baseline (logistic regression) model. We found preliminary and suggestive evidence that the classical machine learning model remained superior for binary agrammatic utterances detection, however, fine-tuned LLMs showed advantages in approximating the LLM-derived surprisal index compared to their pre-trained counterparts and the baseline logistic regression model. We additionally conducted an exploratory analysis in Study 3, where we prompted an LLM for end-to-end aphasia severity prediction. Our overall results suggest that, as a proof-of-concept, transformer-based models, particularly when fine-tuned on curated synthetic data, can learn task-relevant patterns of language dysfunction. Our findings demonstrate the potential utility of LLMs in modeling not only language function (language produced by the general population) but also patterns of language dysfunction (such as those observed in aphasia), offering insights into the linguistic features of disordered language. We cautiously conclude that, with further development, LLMs may serve as useful tools for data evaluation and generation in the study of language dysfunction.
Related Concept Videos
Language and Cognition
Introduction to Language of Pathophysiology ll
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...