Related Experiment Video
Updated: Sep 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Human Expertise and Large Language Model Embeddings in the Content Validity Assessment of Personality Tests
Nicola Milano1, Michela Ponticorvo1, Davide Marocco1
1Department of Humanistic Studies, Natural and Artificial Cognition Laboratory "Orazio Miglino," University of Naples Federico II, Naples, Italy.
Large Language Models (LLMs) show promise in evaluating psychometric instruments like the Big Five Questionnaire (BFQ) and Big Five Inventory (BFI). Hybrid systems combining human expertise and AI offer scalable, objective test development.
Area of Science:
- Psychometrics and Psychological Assessment
- Artificial Intelligence in Behavioral Science
- Natural Language Processing for Measurement
Background:
- Content validity is crucial for ensuring psychological measures accurately reflect intended constructs.
- Traditional methods rely on human expert judgment, which can be time-consuming and subjective.
- The increasing complexity of psychological constructs necessitates innovative validation approaches.
Purpose of the Study:
- To investigate the application of Large Language Models (LLMs) in assessing the content validity of psychometric instruments.
- To compare the performance of LLMs against human expert evaluations in semantic item-construct alignment.
- To explore the impact of LLM training strategies on content validity assessment accuracy.
Main Methods:
- Human experts (graduate psychology students) used the Content Validity Ratio to evaluate items for the Big Five Questionnaire (BFQ) and Big Five Inventory (BFI).
- State-of-the-art LLMs, including multilingual and fine-tuned models, analyzed item embeddings to predict construct mappings.
- Comparison of semantic item-construct alignment accuracy between human and AI-driven approaches.
Main Results:
- Human validators demonstrated higher accuracy with the behaviorally rich BFQ items.
- LLMs showed superior performance with the linguistically concise BFI items.
- LLM performance was significantly influenced by training strategies, with lexically focused models outperforming general ones.
Conclusions:
- LLMs offer a scalable and objective approach to content validity assessment in psychometric instrument development.
- Hybrid systems integrating human expertise with AI precision present a promising direction for robust test validation.
- LLMs have the potential to transform psychological assessment methodologies, enhancing efficiency and objectivity.
More Related Videos
Related Concept Videos
Introduction to Personality Psychology
Early Theories of Personality
The study of...
Stereotype Content Model
Self-Report Tests of Personality
Five-Factor Theory of Personality
Openness reflects creativity, curiosity, and openness to new experiences. Individuals scoring high in openness are imaginative, have a wide range of interests, and are independent...
Personality Theory by Eysenck and Eysenck
In the extroversion/introversion dimension, highly extroverted people are sociable, outgoing, and easily connect with others. In contrast,...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...

