チェス,ショギ,自己プレイをマスターする一般的な補強学習アルゴリズム
David Silver1,2, Thomas Hubert3, Julian Schrittwieser3
1DeepMind, 6 Pancras Square, London N1C 4AG, UK. davidsilver@google.com dhcontact@google.com.
まとめ
新しい人工知能アルゴリズム AlphaZeroは 強化学習を使って チェスやゴーのようなゲームで 超人的なパフォーマンスを発揮します ランダムなゲームでトップのプログラムを倒すのです ランダムなゲームを倒すのです
科学分野:
- 人工知能
- 機械学習
- ゲーム理論
背景:
- チェスは人工知能 (AI) 研究の長年の基準となっています.
- 現在の人工知能のチェスプログラムは 人間の専門知識と 手作りによる評価機能に依存しています
- 強化学習 (RL) の最近の進歩は,ゴーのような複雑なゲームに希望を示しています.
研究 の 目的:
- 複数の挑戦的なゲームで超人的なパフォーマンスを達成できる 一般的なAIアルゴリズムを開発します
- 領域特有の知識なしに自己学習による補強学習の有効性を実証する.
主な方法:
- この研究は,AlphaGo Zeroの方法論を一般化する統一されたアプローチであるAlphaZeroアルゴリズムを導入しています.
- AlphaZeroは深層ニューラルネットワークとモンテカルロツリー検索を利用し,自己学習による補強学習で訓練されます.
- アルゴリズムは,ランダムなゲームから始まるゲームのルールをのみ提供します.
主要な成果:
- アルファゼロは,チェス,ショギ,ゴーで超人的なパフォーマンスを達成しました.
- このアルゴリズムは 世界チャンピオンレベルのAIプログラムに 説得力のある勝利を収めました
- これは一般的なゲームAIの 重要な進歩を示しています
結論:
- 自己学習による強化学習は 複雑な戦略的領域で 超人的なAIを開発するための強力なパラダイムです
- アルファゼロは より一般的な人工知能システムへの 重要な一歩を踏み出しています
- このアプローチは,AIゲーム開発における広範なドメイン特有のエンジニアリングの必要性を排除します.
関連する概念動画
Master Transcription Regulators
7.8K
Master transcription regulators are regulatory proteins that are predominantly responsible for regulating the expression of multiple genes. Often these genes work in concert to drive a complex process. Activation of a master transcription regulator can lead to a cascade of transcriptional activation necessary for that outcome. These regulators can directly bind to the regulatory sequences of the various genes involved, or they can indirectly regulate transcription by binding to regulatory...
7.8K
Master Transcription Regulators
2.8K
2.8K
Reinforcement
919
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
919
Social Foundations of Self I: Play and Game
213
The development of self in children is deeply rooted in social interactions, mainly through stages of play and structured games. These stages, outlined by sociologist George Herbert Mead, illustrate how children progressively learn to understand and adopt social roles, forming a cohesive sense of self.The Play Stage: Imitation and Simple Role-TakingIn the early years of childhood, the play stage is characterized by imitative behavior, where children engage in role-playing based on familiar...
213
Corrosion of Reinforcement
577
The corrosion of steel reinforcement within concrete is a process influenced by the material's inherent properties and external factors. The high pH level of around 13, provided by calcium hydroxide present in concrete, initially protects the steel reinforcement by promoting the formation of a passive iron oxide layer on its surface.
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
577
Reinforcement Schedules
503
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
503


