商业和开源大型语言模型的比较,用于标记胸部X射线图报告
Felix J Dorfner1, Liv Jürgensen1, Leonhard Donle1
1From the Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School, 149 Thirteenth St, Charlestown, MA 02129 (F.J.D., T.R.B., M.C.C., A.E.K., C.P.B.); Department of Radiology, Charité-Universitätsmedizin Berlin, corporate member of Freie Universität Berlin and Humboldt Universität zu Berlin, Berlin, Germany (F.J.D., L.D., F.A.M., F.B., L.J.); Department of Pediatric Oncology, Dana-Farber Cancer Institute, Boston, Mass (L.J.); Department of Diagnostic and Interventional Radiology, Technical University of Munich, Munich, Germany (L.C.A.); Mass General Brigham Data Science Office, Boston, Mass (J.S., T.S., C.P.B.); Microsoft Health and Life Sciences (HLS), Redmond, Wash (J.M.); Klinikum rechts der Isar, Technical University of Munich, Munich, Germany (K.K.B.); Department of Radiology and Nuclear Medicine, German Heart Center Munich, Munich, Germany (K.K.B.); and Department of Cardiovascular Radiology and Nuclear Medicine, Technical University of Munich, School of Medicine and Health, German Heart Center, TUM University Hospital, Munich, Germany (K.K.B.).
在零射击胸部X射线报告标签方面,GPT-4略高于开源大型语言模型 (LLM). 然而,用例子提示少数人,缩小了绩效差距,显示了开源LLMs的可比结果.
科学领域:
- 医疗成像中的人工智能
- 放射学自然语言处理用于放射学.
- 医疗保健中的机器学习
背景情况:
- 大型语言模型 (LLM) 正在迅速发展,有许多商业和开源选项可供选择.
- 之前的研究集中在GPT-4用于放射学报告分析,但与领先的开源LLM缺乏现实世界的比较.
- 从胸部X射线图报告中准确地提取发现对于临床决策至关重要.
研究的目的:
- 将领先的开源LLM与GPT-4的性能进行比较,从胸部X射线图报告中提取相关发现.
- 在这个任务中评估零射击和少数射击提示策略的有效性.
主要方法:
- 对两个独立数据集的自由文本胸部X光学报告 (ImaGenome和马萨诸塞州总医院) 的回顾性分析.
- 商业型号 (GPT-3.5 Turbo,GPT-4) 与开源型号 (Mistral-7B,Mixtral-8×7B,Llama 2-13B,Llama 2-70B,Qwen1.5-72B) 和CheXbert/CheXpert-labeler.com) 的比较,这些商业型号的使用情况是如下:
- 使用零射击和少数射击提示进行评估,通过F1得分测量性能,并使用McNemar测试进行比较.
主要成果:
- 在ImaGenome数据集上,Llama 2-70B获得了0.97 (零射击) 和0.97 (少数射击) 的微F1得分,与GPT-4的0.98.8密切匹配.
- 在机构数据集上,一个集成开源模型实现了微F1得分0.96 (零射击) 和0.97 (少数射击),与GPT-4的0.98和0.97.97.相当.
- GPT-4在零射击标签方面表现出优越性,但少数射击提示显著改善了开源模型的性能,产生了可比的结果.
结论:
- 虽然GPT-4在零射击报告标签方面表现出色,但使用最小的示例进行少数射击提示使开源LLM能够达到接近GPT-4的性能水平.
- 在不同的数据集和LLM架构中,几次射击提示的有效性各不相同.
- 开源的LLM显示了在放射学报告分析中的临床应用的巨大潜力,特别是当它们与少数射击学习进行微调时.
更多相关视频
02:09Multi-modal Pulmonary Imaging: Using Complementary Information from CT and Hyperpolarized 129Xe MRI to Evaluate Lung Structure-Function
Published on: April 12, 2024
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
相关概念视频
Molecular Models
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
