Related Experiment Video
Updated: Jun 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
How does AI compare to the experts in a Delphi setting: simulating medical consensus with large language models
Young Suk Park1,2, Dongjae Jeon1, Songchang Shi2,3
1Department of Surgery, Seoul National University Bundang Hospital, Seongnam-si, Seoul National University College of Medicine, Seoul, Republic of Korea.
Background:
Several attempts have been made to enhance decision-making capabilities of large language models (LLMs) through debate and collaboration, simulating human-like deliberative processes. However, limited research exists on whether the collective intelligence of LLMs can reproduce consensus decisions of human experts. We investigated consensus-building processes and outcomes among LLMs using a modified Delphi method, comparing results to a human expert Delphi study.
Methods:
We conducted a three-round Delphi study involving eight LLMs, evaluating 135 medical statements from the International Federation for the Surgery of Obesity and Metabolic Disorders 2024 Delphi study. LLMs independently assessed statements in Round 1, refined their opinions based on feedback integration in Round 2, and engaged in pairwise debate in Round 3. Consensus was defined as ≥70% agreement. Concordance was defined as identical outcomes between LLMs and human experts, either both reaching or both failing to reach consensus.
Results:
LLMs achieved a higher overall consensus rate than human experts (93.3% vs. 81.5%, P = 0.002). Initial independent evaluations yielded consensus on 117 statements (86.7%), with five additional statements reaching consensus after feedback integration and four more following structured debates. Concordance between LLM and human expert consensus outcomes was observed in 78.5% of statements overall, and in 91.8% of statements where human experts had achieved consensus. The consensus rates between LLMs and human experts demonstrated a strong positive correlation (Spearman's rho = 0.73, P < 0.001). Substantial variation was observed among individual LLMs in their likelihood of changing decisions in response to peer feedback during Round 2 (0-44.4%). Similarly, considerable differences existed between LLMs in their ability to persuade others (0-63.6%) or their susceptibility to persuasion (0-80.0%) during Round 3.
Conclusion:
LLM-based Delphi methods demonstrated high clinical consensus closely aligned with human expert decisions. LLMs effectively simulated structured human-like deliberative reasoning, though they tended to adopt more guideline-driven and conservative positions. However, the use of commercial LLM platforms limits control over model parameters that may affect reproducibility. While the current study suggests that LLMs hold promise as complementary tools in medical consensus-building processes, further research addressing parameter optimization is warranted.

