在大型语言模型中评估道德能力的路线图
Julia Haas1, Sophie Bridgers2, Arianna Manzini2
1Google DeepMind, London, UK. juliahaas@google.com.
Nature
|February 18, 2026
概括
评估大型语言模型 (LLM) 的道德能力对于它们的安全部署至关重要. 需要新的评估方法来应对AI伦理学中的模仿和复杂性等挑战.
科学领域:
- 人工智能伦理学 人工智能伦理学
- 计算道德 计算机道德
- 人工智能 安全 AI 安全
背景情况:
- 大型语言模型 (LLM) 越来越多地用于敏感的角色,需要了解他们的道德能力.
- 目前的评估重点是道德表现 (输出适当性),而不是道德能力 (基于道德考虑的推理).
- 评估道德能力对于预测AI行为,建立信任和证明道德归因至关重要.
研究的目的:
- 超越评估道德表现的范围,评估法学士的道德能力.
- 识别和解决评估LLM道德能力的基本挑战.
- 提出一个科学基础的AI道德能力评估路线图.
主要方法:
- 确定关键挑战:副本问题 (模仿与理解),道德多维性 (情境敏感因素) 和道德多元化 (全球人工智能标准).
- 倡导一系列对抗性和确认性评估.
- 开发一个框架来评估AI的道德能力.
主要成果:
- 由于模型架构和道德复杂性,LLM道德能力评估面临重大障碍.
- 副本问题,道德多维性和道德多元主义被认为是关键的挑战.
- 建议采用结构化的评估方法来科学评估LLM道德能力.
结论:
- 对LLM道德能力的强有力的科学理解需要解决已确定的挑战.
- 负责任地将道德能力归咎于LLM,需要严格,有科学依据的评估.
- 未来的AI开发和部署必须优先考虑对AI道德能力的伦理评估.
相关概念视频
Kohlberg's Theory of Moral Development
1.1K
Kohlberg's theory of moral development uses the Heinz dilemma — a thought experiment in which a man, Heinz, must decide whether to steal an unaffordable drug to save his dying wife — to illustrate the evolution of moral reasoning. This framework, divided into three levels with two stages, highlights how individuals' understanding of right and wrong becomes increasingly complex.
Pre-Conventional Level
At the pre-conventional level, morality is primarily driven by personal...
Pre-Conventional Level
At the pre-conventional level, morality is primarily driven by personal...
1.1K
Improving Translational Accuracy
3.7K
3.7K
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Stereotype Content Model
15.5K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.5K
Self-Evaluation Maintenance Model
343
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...
343
Language Development
973
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
973

