Related Experiment Videos
Human-Governed Validation of Artificial Intelligence-Generated Medical Assessment Artifacts: A Technical Report
Vinicius C Destefani1, Matheus Feliciano C Ferreira1, Afranio C Destefani2
1Medical Education, Centro Universitário Integrado, Campo Mourão, BRA.
Abstract:
Generative artificial intelligence (AI) models can create fluent medical assessment materials, including multiple-choice questions, distractors, clinical vignettes, answer explanations, and blueprint tags. However, linguistic plausibility does not establish measurement validity. Using AI-generated text as ready-made assessment material may introduce risks related to clinical accuracy, item-writing quality, blueprint alignment, distractor functioning, fairness, and score interpretation. This technical report presents a human-governed validation workflow that treats AI-generated medical assessment artifacts as candidates requiring staged review rather than immediate deployment. The workflow is organized into seven gates: artifact taxonomy, AI provenance and human accountability, clinical and content review, item-writing and language review, blueprint alignment, psychometric-readiness decision, and final human approval. The psychometric-readiness gate explicitly asks whether item analysis, distractor analysis, differential item functioning (DIF) review, or other empirical evaluation is required before an artifact is used in a scored assessment. By separating expert clinical review from psychometric evidence, the workflow preserves human accountability while acknowledging that expert approval alone cannot support all score-based interpretations. We recommend introducing AI-generated assessment materials into medical education pipelines as preliminary candidates subject to documented governance, human oversight, and empirical evaluation when used for scored or consequential assessment.