Related Experiment Video
Updated: Mar 19, 2026

07:14
Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
522
Automated scoring of the Ambiguous Intentions Hostility Questionnaire with fine-tuned large language models
Yizhou Lyu1, Dennis Combs2, Dawn Neumann3
1Department of Psychology, University of California, Los Angeles.
Psychological Assessment
|March 17, 2026
Summary
Large language models can now automate scoring for the Ambiguous Intentions Hostility Questionnaire (AIHQ), a tool measuring hostile attribution bias. This AI-powered scoring aligns with human ratings, speeding up psychological assessments.
Area of Science:
- Psychology
- Artificial Intelligence
- Computational Linguistics
Background:
- Hostile attribution bias, the tendency to perceive social interactions as hostile, is measured by the Ambiguous Intentions Hostility Questionnaire (AIHQ).
- AIHQ open-ended responses offer deep insights but require laborious human scoring.
- Automating AIHQ scoring could significantly enhance research and clinical efficiency.
Purpose of the Study:
- To evaluate the efficacy of large language models (LLMs) in automating the scoring of AIHQ open-ended responses.
- To assess the alignment of LLM-generated scores with human expert ratings.
- To explore the potential of LLMs in facilitating psychological assessments.
Main Methods:
- Utilized a dataset of AIHQ responses from individuals with and without traumatic brain injury (TBI).
- Fine-tuned two LLMs on human-generated ratings for half of the dataset.
- Tested the fine-tuned models on the remaining AIHQ responses and an independent dataset.
Main Results:
- LLM-generated ratings demonstrated strong alignment with human ratings for hostility and aggression attributions.
- Fine-tuned models exhibited higher alignment compared to baseline models.
- The models successfully replicated known group differences in attributional styles between TBI and non-TBI groups.
Conclusions:
- LLMs can effectively and accurately automate the scoring of AIHQ responses, streamlining psychological assessments.
- This automated approach holds promise for both research and clinical applications, particularly for populations like those with TBI.
- Accessible scoring interfaces were developed to encourage broader adoption of LLM-based AIHQ scoring.
