来自生物医学文献的定性指标,用于评估临床决策中的大型语言模型:叙述性审查
Cindy N Ho1, Tiffany Tian1, Alessandra T Ayers1
1Diabetes Technology Society, Burlingame, CA, USA.
BMC medical informatics and decision making
|November 27, 2024
概括
研究人员评估了临床决策中的大型语言模型 (LLM),发现使用了准确性和完整性等共同标准. 然而,报告方法有很大的差异,突出了评估LLM绩效的标准化需求.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 临床决策支持 临床决策支持
背景情况:
- 大型语言模型 (LLM),以ChatGPT为例,自2022年底以来,在临床决策支持方面获得了吸引力.
- 医学界在评估临床环境中的LLM绩效方面缺乏共识.
研究的目的:
- 审查和综合用于评估临床应用中的大型语言模型 (LLM) 的方法和标准.
- 确定共同的评估指标,并突出显示在医疗保健LLM绩效报告标准的差异.
主要方法:
- 对2022年12月1日至2024年4月1日期间发表的文章进行了PubMed的文献综述.
- 审查的重点是评估LLM产生的诊断或治疗计划的研究,选择了108篇相关出版物.
主要成果:
- GPT-3.5,GPT-4,Bard,LLaMa/Alpaca和Bing Chat是最常被评估的法学士课程.
- 关键的评估标准包括准确性,完整性,适当性,洞察力和一致性.
- 在研究报告发现和评估LLM绩效的方式中观察到显著的差异.
结论:
- 在过去的1.5年里,高质量的LLMs的一致标准已经出现.
- 需要对定性评估指标进行标准化报告,以推进医疗保健中LLM的研究.
- 制定标准化指标将促进对LLM实用性和安全性的更强大和可比的研究.
更多相关视频
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.4K
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
7.5K
相关概念视频
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
121
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
121
Sensitivity, Specificity, and Predicted Value
192
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
192
Methods of Documentation VI: Case Management Model
559
The case management model is a multidisciplinary approach that involves healthcare professionals from diverse disciplines, such as physicians, nurses, therapists, social workers, and pharmacists, working collaboratively to address the various needs of patients. Each healthcare professional brings unique expertise and perspectives, contributing to a more comprehensive understanding of the patient's condition and tailoring treatment plans accordingly.
For example, a patient with a chronic...
For example, a patient with a chronic...
559
Improving Translational Accuracy
2.5K
2.5K
Clinical Trials
6.6K
Clinical trials are prospective experimental studies conducted on humans to determine the safety and efficacy of treatments, drugs, diet methods, and medical devices. Using statistics in clinical trials enables researchers to derive reasonable and accurate conclusions from the collected data, allowing them to make wise decisions in uncertain situations. In medical research, statistical methods are crucial for preventing errors and bias.
There are four phases in a clinical trial. A phase one...
There are four phases in a clinical trial. A phase one...
6.6K
Receiver Operating Characteristic Plot
82
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
82
