Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a

Shahid Akhtar Akhund1, Sadia Qazi2, Muhammad Atif Mazhar2

  • 1Department of Medical Education and Department of Anatomy, College of Medicine, Alfaisal University, Riyadh, Saudi Arabia.

Abstract

Related Concept Videos