How does AI compare to the experts in a Delphi setting: simulating medical consensus with large language models

Young Suk Park1,2, Dongjae Jeon1, Songchang Shi2,3

  • 1Department of Surgery, Seoul National University Bundang Hospital, Seongnam-si, Seoul National University College of Medicine, Seoul, Republic of Korea.

Summary

Large language models (LLMs) achieved higher consensus than human experts in a modified Delphi study. Their collective intelligence closely matched human expert decisions, showing promise for medical consensus-building.