大規模な言語モデルにおける道徳的能力の評価のためのロードマップ
Julia Haas1, Sophie Bridgers2, Arianna Manzini2
1Google DeepMind, London, UK. juliahaas@google.com.
Nature
|February 18, 2026
まとめ
大型言語モデル (LLM) の道徳的能力の評価は,その安全な展開に不可欠です. AI倫理における模倣や複雑さなどの課題に対処するために,新しい評価方法が必要です.
科学分野:
- 人工知能 倫理 人工知能 倫理
- コンピューティング・モラル コンピューティング・モラル
- AI 安全性 AI 安全性
背景:
- 大型言語モデル (LLM) は,敏感な役割においてますます使用され,その道徳的な能力の理解が求められています.
- 現在の評価は,道徳的能力 (道徳的考慮に基づいた推論) よりも,道徳的パフォーマンス (出力の適切性) に焦点を当てています.
- 道徳的能力の評価は,AIの行動を予測し,信頼を築き,道徳的属性を正当化するために不可欠です.
研究 の 目的:
- 道徳的業績の評価を超えて,LLMの道徳的能力を評価すること.
- LLMの道徳的能力の評価における根本的な課題を特定し,対処する.
- 人工知能の道徳的能力の科学的根拠に基づく評価のためのロードマップを提案する.
主な方法:
- 主要な課題の特定:ファクシミレ問題 (模倣対理解),道徳的多次元性 (文脈に敏感な要因),道徳的多元主義 (グローバルなAI標準).
- 敵対的および確認的評価のセットのための提唱.
- 人工知能の道徳的能力を評価するための枠組みの開発.
主要な成果:
- LLMの道徳的な能力の評価は,モデルアーキテクチャと道徳的な複雑さにより,大きな障害に直面しています.
- ファクシミレ問題,道徳的多次元性,道徳的多元主義は,重要な課題として認識されています.
- LLMの道徳的能力を科学的に評価するために,構造化された評価アプローチが提案されています.
結論:
- LLMの道徳的能力に関する強力な科学的理解は,特定された課題に取り組むことを要求します.
- LLMに道徳的能力の責任ある付与は,厳格で科学的に根拠のある評価を必要とします.
- 将来のAIの開発と導入は,AIの道徳的な能力の倫理的評価を優先しなければなりません.
関連する概念動画
Kohlberg's Theory of Moral Development
1.1K
Kohlberg's theory of moral development uses the Heinz dilemma — a thought experiment in which a man, Heinz, must decide whether to steal an unaffordable drug to save his dying wife — to illustrate the evolution of moral reasoning. This framework, divided into three levels with two stages, highlights how individuals' understanding of right and wrong becomes increasingly complex.
Pre-Conventional Level
At the pre-conventional level, morality is primarily driven by personal...
Pre-Conventional Level
At the pre-conventional level, morality is primarily driven by personal...
1.1K
Improving Translational Accuracy
3.7K
3.7K
Improving Translational Accuracy
15.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.2K
Stereotype Content Model
15.5K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.5K
Self-Evaluation Maintenance Model
343
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...
343
Language Development
973
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
973

