Leveraging computerized adaptive testing for cost-effective evaluation of large language models in medical

Tianpeng Zheng1,2, Zhehan Jiang3,4, Jiayi Liu5

  • 1Institute of Medical Education, Health Science Center, Peking University, Beijing, China.

Summary

A new computerized adaptive testing (CAT) framework significantly reduces evaluation time and cost for large language models (LLMs) in healthcare. This psychometrically rigorous method enables efficient, scalable assessment of medical knowledge.

Related Concept Videos