Dominar el juego de Go sin conocimiento humano

David Silver1, Julian Schrittwieser1, Karen Simonyan1

  • 1DeepMind, 5 New Street Square, London EC4A 3TW, UK.

Nature
|October 21, 2017
PubMed
Resumen

Un nuevo algoritmo de inteligencia artificial, AlphaGo Zero, aprende únicamente a través del aprendizaje por refuerzo sin datos humanos. Logró un rendimiento sobrehumano en el juego de Go entrenando con sus propios datos de juego.

Videos de Conceptos Relacionados

Observational Learning01:12

Observational Learning

Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.1K
Understanding Deception01:14

Understanding Deception

Deception is a pervasive aspect of human communication. Empirical studies have shown that most individuals engage in some form of deceit on a daily basis, with approximately 20% of social exchanges involving deceptive elements. Lying follows a developmental trajectory, peaking during adolescence and declining with age, possibly due to the maturation of cognitive control and social accountability.Cognitive and Social Factors in Deception DetectionDespite its prevalence, accurately detecting...
183
Purposive Learning01:22

Purposive Learning

E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
535
High-Level and Low-Level Awareness01:19

High-Level and Low-Level Awareness

Controlled processes in human consciousness represent high-alert mental states where individuals deliberately focus their attention on achieving specific goals. Controlled processes can be seen in situations like mastering new technology, where a person might become so absorbed that they ignore surrounding distractions. Such processes involve selective attention, requiring one to concentrate on particular elements of experience while disregarding others. These are governed by executive...
799
Avoidance Learning and Learned Helplessness01:14

Avoidance Learning and Learned Helplessness

Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.7K
Machines: Problem Solving II01:30

Machines: Problem Solving II

Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
689