Related Experiment Video
Updated: Sep 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using Large Language Model Embeddings to Analyze Neural Activity During Language Processing
Pengli Shao1, Griffin A Shimamoto1, Jing Cai2
1Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA, USA.
Abstract:
Human language emerges from the dynamic interplay of many interconnected elements, including meaning, context, and grammar. This richness makes it difficult to study the neural processes that respond to language. A major challenge is the absence of a precise, comprehensive numerical representation of language that can serve as a bridge between language and neural activity. Recent advances in large language models (LLMs) offer a potential solution by translating language into numerical representations that preserve meaning, grammar, and context. In this paper, we provide a general protocol for using LLM embeddings in conjunction with electrophysiological recordings during language tasks to reveal specific neural patterns in response to language activity. Our approach enables the identification of neural units selectively involved in language production and comprehension by correlating LLM embeddings with neural signals precisely aligned to word onset timestamps. The workflow consists of four stages: (i) extraction and preprocessing of numerical language features representing meaning and context obtained from pretrained language models, (ii) preprocessing and temporal alignment of neural recordings with word-onset timestamps, (iii) correlation and specificity analysis, and (iv) visualization of significantly responding neural units.

