聊天GPT是否准备好在器官特异性药物毒性研究中公开使用?
Skylar Connor1, Leihong Wu1, Ruth A Roberts2
1National Center for Toxicological Research, US Food and Drug Administration, Jefferson, AR 72079, USA.
Drug discovery today
|January 22, 2025
概括
大型语言模型 (LLM) 在评估公共卫生应用中药物毒性的准确性中等. 为了提高可靠性,可能需要像Retrieval Augmented Generation这样的高级框架.
科学领域:
- 人工智能在医学中的应用
- 药物监督 药物监督 药物监督
- 计算毒理学计算毒理学
背景情况:
- 像ChatGPT这样的大型语言模型 (LLM) 越来越多地被使用,这引发了人们对其在公共卫生中的可靠性的担忧.
- 准确的药物毒性评估对于患者安全和监管监督至关重要.
研究的目的:
- 评估GPT-4在评估肝脏,心脏和功能药物毒性的准确性.
- 为了比较GPT-4药物毒性评估的一般和专家提示的性能.
主要方法:
- 使用两种不同的方法提示GPT-4:"一般提示"和"专家提示".
- 评估与来自美国食品和药物管理局 (FDA) 药物标签文件的专家评估进行了基准对比.
- 对肝脏,心脏和脏终点的药物毒性进行了评估.
主要成果:
- 与"一般提示" (48-72%) 相比",专家提示"的准确性更高 (64-75%).
- 在药物毒性评估中,GPT-4的整体性能中度,这表明直接用于公共卫生的限制.
- 在不同的器官系统中,GPT-4的准确性各不相同.
结论:
- GPT-4显示出潜力,但由于准确度中等,在公共卫生中需要谨慎应用.
- 改进,如检索增强生成 (RAG),可能是必要的,以提高药物毒性评估的LLMs的可靠性.
- 需要进一步的研究来优化LLM框架,以应对敏感的公共卫生任务.
相关概念视频
Preclinical Development: Overview
4.8K
Preclinical development consists of a series of tests that ensure the safety and efficacy of a new therapeutic compound before it is tested in humans. There are four main phases to this process. First, safety pharmacology tests are conducted to ensure the drug does not produce any acutely harmful effects. These tests examine parameters such as bronchoconstriction, cardiac dysrhythmias, blood pressure changes, and ataxia. Next, preliminary toxicological testing is performed to determine the...
4.8K
Toxicity Testing in Animals
200
Toxicity tests in animals are grounded on two main assumptions: first, the effects observed in laboratory animals can be extrapolated to humans, especially when adjusted for body surface area; second, high-dose exposure in animals is essential to identify potential human hazards from lower doses. This is based on the quantal dose-response concept, which faces the challenge of extrapolating results from relatively few test animals to much larger human populations. For example, a 0.01% incidence...
200


