在大型语言模型中评估反LGBTQIA+医学偏见
Crystal T Chang1, Neha Srivathsa2, Charbel Bou-Khalil3
1Department of Dermatology, Stanford University, Stanford, California, United States of America.
PLOS digital health
|September 8, 2025
概括
这项研究发现,大型语言模型 (LLM) 表现出反LGBTQIA+偏见和错误信息,不适当的反应经常发生. 需要进一步开发以提高准确性并减少对LGBTQIA+患者的偏见.
科学领域:
- 医疗保健中的人工智能
- 医疗信息学 医疗信息学
- 健康 公平 卫生 公平
背景情况:
- 大型语言模型 (LLM) 在临床环境中越来越多地用于患者沟通和决策支持.
- 现有的研究强调了LLM的基于种族和性别的偏见,但反LGBTQIA+偏见仍未得到充分研究.
- 医疗保健差异不成比例地影响LGBTQIA+个体,强调需要评估这一群体中的LLM偏见.
研究的目的:
- 评估四个领先的大型语言模型 (LLM) 传播反LGBTQIA+医学偏见和错误信息的潜力.
- 评估LLM对涉及LGBTQIA+身份的提示的响应的适当性和临床实用性.
- 在敏感的临床环境中建立一个基准数据集,用于评估未来的LLM绩效.
主要方法:
- 四名法学士 (Gemini 1.5 Flash,Claude 3 Haiku,GPT-4o,斯坦福医学安全GPT) 得到了38对问题和合成临床笔记的提示.
- 提示程序旨在探索具有或没有明确的LGBTQIA+身份术语的临床情况,在相关和不相关的临床环境中.
- 经过医学培训的审查员和LGBTQIA+健康专家评估了安全性,隐私,准确性,偏见和临床实用性的反应.
主要成果:
- 这四个LLM都对带有和没有LGBTQIA+身份术语的提示产生了不适当的响应.
- 不适当反应的比例在LGBTQIA+特定提示的43-62%,在一般提示的47-65%之间.
- 幻觉/准确性是不适当分类的最常见原因,其次是偏见和安全问题,LGBTQIA+提示引发了更严重的偏见.
结论:
- 法律专业医生显示出传播反LGBTQIA+医学偏见和错误信息的巨大潜力.
- 与适当的相比,不适当的反应的临床实用性得分较低.
- 未来的研究必须专注于提高LLM准确性,减少偏见,并为LGBTQIA+患者护理量身定制输出.
相关概念视频
Improving Translational Accuracy
3.6K
3.6K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Bias in Epidemiological Studies
1.3K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
1.3K
Stereotype Content Model
15.4K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.4K
Stereotypes, Prejudice, and Discrimination
95.0K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
95.0K
Classification of Illness
8.6K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
8.6K


