:,

Mitul Gupta1, John Virostko2, Christopher Kaufmann1

  • 1The University of Texas at Austin, Dell Medical School, Department of Diagnostic Medicine, Austin, TX, USA.

PubMed
概括

这项研究评估了放射学中的大型语言模型 (LLM),发现GPT-4最初最准确,但随着时间的推移而下降. 持续的基准测试对于评估医学应用中的LLM可靠性至关重要.

相关概念视频