Related Experiment Video
Updated: Jun 26, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
How does AI compare to the experts in a Delphi setting: simulating medical consensus with large language models
Young Suk Park1,2, Dongjae Jeon1, Songchang Shi2,3
1Department of Surgery, Seoul National University Bundang Hospital, Seongnam-si, Seoul National University College of Medicine, Seoul, Republic of Korea.
International Journal of Surgery (London, England)
|October 15, 2025
Summary
Large language models (LLMs) achieved higher consensus than human experts in a modified Delphi study. Their collective intelligence closely matched human expert decisions, showing promise for medical consensus-building.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Collective Intelligence
Background:
- Large language models (LLMs) show potential for enhancing decision-making through simulated debate and collaboration.
- Limited research exists on whether LLM collective intelligence can replicate human expert consensus.
Purpose of the Study:
- To investigate consensus-building processes and outcomes among LLMs using a modified Delphi method.
- To compare LLM consensus results with those from a human expert Delphi study.
Main Methods:
- A three-round Delphi study involving eight LLMs evaluated 135 medical statements.
- LLMs independently assessed statements, refined opinions with feedback, and engaged in pairwise debates.
- Consensus was defined as ≥70% agreement; concordance measured identical outcomes between LLMs and human experts.
Main Results:
- LLMs achieved a higher overall consensus rate (93.3%) than human experts (81.5%).
- Concordance between LLM and human expert consensus was high (78.5% overall, 91.8% for human-agreed statements).
- Significant variation existed in individual LLMs' decision-changing likelihood and persuasion abilities.
Conclusions:
- LLM-based Delphi methods demonstrated high clinical consensus, closely aligning with human expert decisions.
- LLMs simulated deliberative reasoning but adopted more conservative, guideline-driven positions.
- Further research on parameter optimization is needed for LLM reproducibility in medical consensus-building.

