人間が知らずに ゴルフをマスターする
David Silver1, Julian Schrittwieser1, Karen Simonyan1
1DeepMind, 5 New Street Square, London EC4A 3TW, UK.
Nature
|October 21, 2017
まとめ
新しい人工知能のアルゴリズムである アルファゴー・ゼロは 人間のデータなしで 補強学習で学習します 勝負データでトレーニングすることで 超人的なパフォーマンスを達成しました
科学分野:
- 人工知能
- 強化学習
- ゲーム理論
背景:
- 人工知能 (AI) の重要な目標は,複雑な領域で超人的な熟練度を達成できるアルゴリズムを作成することです.
- AlphaGoのような以前の AI システムは 人間の専門家データからの 監督学習と 自己学習による強化学習を利用していました
- しかし,これらのシステムは依然として人間の知識と指導に依存していました.
研究 の 目的:
- 人工知能のアルゴリズムを開発し 人工知能のアルゴリズムを開発し 人工知能のアルゴリズムを開発し 人工知能のアルゴリズムを開発し
- 自己学習を通じて 繰り返しパフォーマンスを向上させることができる
- Goのゲームで超人的なパフォーマンスを 達成するために 新しい自己学習のAIアプローチを使います
主な方法:
- この研究は,強化学習のみに基づいた アルゴリズムである AlphaGo Zero を導入します.
- ニューラルネットワークは アルファゴー・ゼロの 移動選択とゲーム結果を 予測するように訓練されています
- このニューラルネットワークはツリー検索アルゴリズムを強化し,次のイテレーションでは動作選択の改善と自己プレイの強化につながります.
主要な成果:
- アルファゴーゼロは ゼロから学びました 人間のデータや ガイドラインや ゲームルールを超えた領域の知識は ありませんでした
- アルゴリズムはゴーゲームで 超人的な性能を達成しました
- アルファゴー・ゼロは,以前のチャンピオンレベルのアルファゴーを100対0で倒した.
結論:
- 人工知能が難しい領域で 超人的なパフォーマンスを発揮するには 人工知能のデータなしで 補強学習だけでは十分です
- 自己学習と自己改善は AIの学習と開発の強力なメカニズムです
- アルファゴーゼロは AIの重要な進歩であり アルゴリズムトレーニングの新たなパラダイムを示しています
関連する概念動画
Observational Learning
1.1K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Understanding Deception
183
Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...
183
Purposive Learning
535
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
535
High-Level and Low-Level Awareness
799
Controlled processes in human consciousness represent high-alert mental states where individuals deliberately focus their attention on achieving specific goals. Controlled processes can be seen in situations like mastering new technology, where a person might become so absorbed that they ignore surrounding distractions. Such processes involve selective attention, requiring one to concentrate on particular elements of experience while disregarding others. These are governed by executive...
799
Avoidance Learning and Learned Helplessness
2.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Machines: Problem Solving II
689
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
689


