Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Data Validation01:03

Data Validation

Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Self-Report Tests of Personality01:22

Self-Report Tests of Personality

Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Behavioral signs of trauma on the Rorschach: Development of the Trauma Experience Index.

Psychological trauma : theory, research, practice and policy·2025
Same author

Developmental cascades from early childhood attachment security to adolescent level of personality functioning among high-risk youth.

Development and psychopathology·2024
Same author

Comparing the Validity of the Rorschach Performance Assessment System and Exner's Comprehensive System to Differentiate Patients and Nonpatients.

Assessment·2023
Same author

Cross-cultural investigation from nine countries on the associations of antisocial traits and the WHO's containment measures for the COVID-19 pandemic.

Scandinavian journal of psychology·2022
Same author

Legal Admissibility of the Rorschach and R-PAS: A Review of Research, Practice, and Case Law.

Journal of personality assessment·2022
Same author

Rorschach Performance Assessment System (R-PAS) Interrater Reliability in a Brazilian Adolescent Sample and Comparisons With Three Other Studies.

Assessment·2020

Related Experiment Video

Updated: Jun 20, 2026

Advancing Dyslexia Assessment in Children Through Computerized Testing
09:00

Advancing Dyslexia Assessment in Children Through Computerized Testing

Published on: August 16, 2024

Evaluating LLM-Based Coders in Psychological Assessment: A Validation Framework With Application to the Rorschach

Ruam P F A Pimentel1, Gregory J Meyer1,2

  • 1Rorschach Performance Assessment System, OH, USA.

Assessment
|June 19, 2026
PubMed
Summary

A new framework validates large language model (LLM) scoring in psychological assessment. LLM coders demonstrated reliable and valid scoring, comparable to human experts, for Rorschach Morbid Content.

Keywords:
Rorschachautomated scoringmorbid contentpsychological assessmentpsychometricsvalidation framework

More Related Videos

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
12:55

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties

Published on: September 27, 2020

Related Experiment Videos

Last Updated: Jun 20, 2026

Advancing Dyslexia Assessment in Children Through Computerized Testing
09:00

Advancing Dyslexia Assessment in Children Through Computerized Testing

Published on: August 16, 2024

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
12:55

Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties

Published on: September 27, 2020

Area of Science:

  • Psychological assessment
  • Artificial intelligence in mental health
  • Psychometrics

Background:

  • Large language models (LLMs) are emerging tools in psychological assessment.
  • Current standards for evaluating LLM scoring accuracy are insufficient.
  • Reliable and valid scoring is crucial for clinical utility.

Purpose of the Study:

  • To introduce a reproducible framework for evaluating LLM-based scoring systems.
  • To assess the reliability and validity of LLM coders in psychological testing.
  • To provide practical guidance for implementing automated scoring.

Main Methods:

  • Developed a validation framework with pre-validation and standardized phases.
  • Applied a two-agent LLM workflow for Morbid Content (MOR) scoring in the Rorschach task.
  • Evaluated LLM performance on an independent dataset (n=84) using agreement and validity metrics.

Main Results:

  • LLM coder achieved good response-level agreement (kappa = .72-.74) and excellent protocol-level agreement (ICC = 0.94-0.95) with human assessors.
  • Demonstrated high internal consistency (ICC = 0.97-0.99) and replicated external validity (r = .59-.71).
  • LLM validity metrics were comparable to those of human coders (r = .54-.65).

Conclusions:

  • The proposed framework offers a clear and reproducible method for LLM scoring validation.
  • LLM-based scoring systems can achieve reliable and valid results in psychological assessment.
  • This work guides the ethical and practical implementation of automated coders in clinical settings.