Related Concept Videos
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
You might also read
Related Articles
Articles linked to this work by shared authors, journal, and citation graph.
Spontaneous, controlled acts of reference between friends and strangers.
Related Experiment Video
Updated: Jul 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can large language models help augment English psycholinguistic datasets?
1Department of Cognitive Science, UC San Diego, 9500 Gilman Dr., La Jolla, CA, 92093-0515, USA. sttrott@ucsd.edu.
Large language models (LLMs) like GPT-4 can generate reliable psycholinguistic norms, matching human agreement levels for tasks like word similarity and sensorimotor associations. This offers a scalable alternative to costly human data collection for language and cognition research.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Area of Science:
- Cognitive Science
- Computational Linguistics
- Psycholinguistics
Background:
- Psycholinguistic datasets, or norms, are crucial for language and cognition research.
- Collecting human judgments for these norms is time-consuming and expensive, especially for large-scale or multi-dimensional data.
- Large language models (LLMs) present a potential solution to overcome these data collection challenges.
Purpose of the Study:
- To investigate the feasibility of using LLMs, specifically GPT-4, to generate psycholinguistic norms for English.
- To compare LLM-generated semantic judgments against human-generated data (the gold standard).
- To assess the utility of LLM-generated norms in statistical modeling and identify potential limitations.
Main Methods:
- GPT-4 was employed to collect various semantic judgments (e.g., word similarity, contextualized sensorimotor associations, iconicity) for English words.
- LLM-generated judgments were systematically compared against established human-generated norms.
- Substitution analyses were conducted, replacing human norms with LLM norms in statistical models.
Main Results:
- GPT-4 judgments showed positive correlations with human judgments across different datasets.
- In some cases, LLM agreement levels met or exceeded average human inter-annotator agreement.
- Substitution analyses indicated that LLM norms generally preserve the direction of parameter estimates in statistical models, though magnitudes may change.
Conclusions:
- LLMs can effectively augment the creation of large-scale psycholinguistic datasets, offering a cost-effective and efficient alternative to human data collection.
- Systematic differences exist between LLM and human norms, requiring careful consideration of factors like data contamination, LLM choice, and validity.
- The study provides valuable LLM-generated norm data and discusses critical considerations for their future use in research.