Related Experiment Video
Updated: Apr 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
The AI interviewer: multi-faceted evaluation of adaptive questioning by large language models.
Anastasia Panfilova1,2, Vadim Bolshev1,2, Mikhail Mozikov3
1Institute of Psychology of the Russian Academy of Sciences, Moscow, Russia.
Scientific Reports
|April 4, 2026
Summary
This study evaluates six large language models (LLMs) as adaptive interviewers in psychological research. Gemini 2.5 Pro excels in empathy, while GPT-5 Chat prioritizes speed, showcasing systematic trade-offs in LLM interviewing capabilities.
Area of Science:
- Artificial Intelligence
- Psychology
- Human-Computer Interaction
Background:
- Large language models (LLMs) are increasingly used as interviewers in research and human-computer interaction.
- Systematic evaluations of LLM interviewing behavior are limited, hindering their effective deployment.
Purpose of the Study:
- To introduce a modular LLM agent for semi-structured psychological interviews.
- To present a controlled, multi-faceted evaluation protocol for assessing interviewer quality across six state-of-the-art LLMs.
Main Methods:
- Developed an LLM agent for adaptive interviews with 54 main questions.
- Standardized interview context using human transcripts and a single LLM interviewee for fair comparison.
- Expert psycholinguists evaluated interviewer behavior on benevolence, necessity, context-awareness, openness, and justified skip, complemented by efficiency and linguistic metrics.
Main Results:
- Gemini 2.5 Pro demonstrated the most empathic tone; GPT-5 Chat optimized for speed and precision.
- Grok 4 achieved exhaustive coverage but with higher latency; Claude Sonnet 4 offered balanced versatility.
- Linguistic markers correlated with human judgments, indicating alignment between stylistic choices and perceived interview quality.
Conclusions:
- Systematic trade-offs exist in LLM interviewing capabilities, with distinct models excelling in different aspects.
- The developed toolkit enables principled deployment and auditing of LLM interviewers in psychological research.
- Researchers can now better match LLM capabilities to study goals for enhanced empathy, appropriateness, and effectiveness.
Related Concept Videos
Language Development
1.1K
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
1.1K
Language and Cognition
978
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
978
Self-Evaluation Maintenance Model
396
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...
396
