Related Experiment Videos
Human vs AI in NCLEX-style item review: Reliability, agreement, and efficiency using the NEIA rubric
Rachel Cox Simms1, Desirée Hensel2, Robert Cavanaugh3
1MGH Institute of Health Professions, School of Nursing, 36 First Avenue, Charlestown Navy Yard, Boston, MA 02129, USA.
Aim:
To establish initial inter-rater reliability estimates for the NEIA rubric across human nursing faculty and AI model configurations and to compare reliability and agreement patterns between rater types.
Background:
High-quality NCLEX-style exam items are essential for valid nursing assessment, yet faculty-developed questions often contain flaws that compromise test validity. The NEIA rubric provides structured criteria for evaluating item quality, but consistent application requires reliable scoring. Generative artificial intelligence offers potential for scalable, consistent rubric application, though its reliability compared with human raters remains unexplored.
Design:
Psychometric study establishing inter-rater reliability estimates for the NEIA rubric.
Methods:
Thirty-three nursing faculty evaluated 15 NCLEX-style items (5 low-quality, 5 moderate-quality, 5 high-quality) using a planned incomplete design (6 items per rater). Seven AI configurations (ChatGPT custom GPT and Claude Projects) evaluated all items in a complete crossed design. Inter-rater reliability was assessed using intraclass correlation coefficients and Krippendorff's alpha. Agreement with expert classifications was measured using percentage exact agreement.
Results:
Individual human raters demonstrated poor reliability (ICC = 0.317; α = 0.147) and 47.3% exact agreement with expert classifications. AI models showed good reliability (ICC = 0.709; α = 0.77) and 69.5% average agreement, with best models achieving 80% agreement. Under the study conditions, AI exhibited fewer extreme misclassifications and more consistent enforcement of stop criteria.
Conclusions:
AI models applied item-quality criteria more consistently than individual faculty and achieved higher agreement with developer-defined expert classifications. Findings support a hybrid model where AI conducts initial screening while faculty retain final decision-making authority.
Related Concept Videos
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...
Nursing Evaluation
Section...