Related Experiment Videos
Evaluating psychiatric conference posters: Benchmarking a custom generative pre-trained transformer against human
Anusa Arunachalam Mohandoss1, Rooban Thavarajah2,3, Raj Kiran Donthu4
1Department of Psychiatry, Shri Sathya Sai Medical College and Research Institute, Affiliated to Sri Balaji Vidyapeeth, Deemed to be University, SBV Nagar, Chennai Campus, Ammapettai, Chengalpet Taluk, Kancheepuram, Nellikuppam, Tamil Nadu, India.
Background:
Scientific poster assessment lacks standardized and discipline-neutral rubrics. Assessment by human reviewers (HRs) is subject to inter-rater variability.
Aim:
To assess PA²IRS (Poster Assessment via AI-Integrated Rubric System) framework for AI-assisted psychiatric poster evaluation, and conducted a reliability study comparing AI and HR agreement.
Methods:
PA²IRS was developed through an AI-assisted iterative criterion refinement process modelled on Delphi principles. Sixty posters (30-case reports/series [CR], 20 original research [OR], 10-systematic review-Meta-analysis [SRMA]) were randomly sampled. Three qualified mental health professionals served as independent reviewers. A custom-GPT (GPT-5.2, GO-subscription) provided AI assessments across three domains: Domain-A (content quality, poster-type specific), Domain-B (visual), and Domain-C (impact). PA²IRS is a 100-point instrument combining an AI-assessable poster component and an in-person interview component. This study concerns only the poster component. Intraclass correlation coefficients (ICCs), Passing-Bablok regression, Bland-Altman analysis, and variance component analysis were performed using appropriate statistical tools.
Results:
AI-human single-measure ICCs [Overall (0.62), Domain-A (0.63), Domain-B (0.44), Domain-C (0.55)] met or exceeded human-human ICCs (0.42, 0.40, 0.27, 0.48) across all domains. Four-rater ICC (with AI) reached 0.75. Variance ratios (AI-human vs inter-human spread) were ≤1.0 across all domains for all posters combined. The SRMA subgroup showed variance ratios of 0.09-0.13 for Domains A and B (Bartlett P ≤ 0.002). Overall score bias was 0.11 percentage points (pp); Domain-A showed a consistent maximum positive bias of 5.5 pp across subgroups.
Conclusion:
AI-human agreement was within or exceeded the inter-human reliability range across three domains. The domain-dependent agreement pattern is consistent with dual-process cognitive theory. PA²IRS supports use as a scalable, standardized, and cross-disciplinarily competent formative biomedical poster assessment tool.