PIEE循环:在临床决策中为Red Teaming大型语言模型提供结构化框架
Maissa Trabilsy1, Srinivasagam Prabha1, Cesar A Gomez-Cabello1
1Division of Plastic Surgery, Mayo Clinic, 4500 San Pablo Rd, Jacksonville, FL 32224, USA.
一个新的框架,PIEE (规划和准备,信息收集和快速生成,执行和评估),有助于评估医疗保健中的人工智能 (AI). 它对临床环境中的安全性和可靠性进行压力测试的大型语言模型 (LLM).
科学领域:
- 医疗信息学 医疗信息学
- 人工智能安全问题 人工智能安全问题
- 临床决策支持 临床决策支持
背景情况:
- 大型语言模型 (LLM) 提供医疗保健的好处,但对患者的安全,准确性和道德构成风险.
- 目前还没有一种标准化的方法来评估临床决策中的LLM安全性.
研究的目的:
- 引入PIEE周期,这是一个结构化的红色团队框架,用于评估医疗保健决策中的AI安全性.
- 在临床场景中为压力测试LLM提供系统方法.
主要方法:
- PIEE周期包括规划和准备,信息收集和快速生成,执行和评估.
- 敌对提示 (越狱,社会工程,分散注意力的攻击) 用于压力测试法学士.
- 绩效使用危害检测率,幻觉率 (TruthfulQA),安全性,可靠性,偏见 (BBQ) 和道德评分等指标进行评估.
主要成果:
- PIEE框架可以模拟现实世界的临床场景,以测试LLM的稳定性.
- 评估指标提供了对LLM绩效的定量和定性评估.
- 该框架可以适应各个医学专业,用整形外科的例子来说明.
结论:
- PIEE周期为评估医学LLM临床可靠性和伦理完整性提供了实际基础.
- 虽然是概念性的,但该框架是为所有医疗提供者提供,以确保安全的AI集成.
- 持续的验证是必要的,以确认框架的有效性.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Methods of Documentation III: PIE
Critical Thinking II
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Patient-centered Care
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
