Related Experiment Videos
Psychometric applications of generative artificial intelligence: Lifecycle, risks, and research agenda
David Villarreal-Zegarra1, Yscenia Paredes-Gonzales2, Jackeline García-Serna2,3
1Escuela de Posgrado, Universidad Continental, Lima, Peru.
Abstract:
Generative artificial intelligence (GenAI), including large language models and large multimodal models, is increasingly used to support psychological assessment and digital mental health measurement. This review proposes a lifecycle framework for evaluating these uses without weakening psychometric standards. The framework covers construct definition; item, prompt, indicator, or signal generation; evaluation of content, response processes, usability, and data quality; piloting and calibration; validation; fairness and measurement invariance; scoring and interpretation; documentation; and post-deployment monitoring. We argue that digital traces from smartphones, chatbots, ecological momentary assessment, wearables, social media, clinical notes, and multimodal systems should not be treated as psychometric measures by default. They should remain raw data, candidate indicators, or algorithmic scores until their construct interpretation and intended use are theoretically specified, technically verified, empirically calibrated, and validated with human data. GenAI may help generate candidate items, refine wording, classify open-text responses, extract structured information, support multimodal integration, and assist scoring under explicit rules. These uses may improve efficiency and scale, but they do not establish validity, objectivity, fairness, or clinical meaning. We distinguish two complementary facets: GenAI for psychometrics, in which models support measurement development and scoring, and psychometrics for GenAI, in which psychometric methods evaluate model behavior when LLMs are part of measurement workflows. Key risks include construct drift, face validity without structural validity, compressed variability in synthetic respondents, algorithmic bias, lack of invariance, prompt sensitivity, model drift, automation bias, data-security and privacy failures, and loss of subjectivity and disagreement. Responsible use requires human oversight, transparent reporting, validation with human data, fairness evaluation, secure data governance, and continuous monitoring. GenAI should augment, not replace, psychometric science.