Video Experimental Relacionado
Updated: Jan 14, 2026

11:53
The Modular Design and Production of an Intelligent Robot Based on a Closed-Loop Control Strategy
Published on: October 14, 2017
12.1K
Descubrimiento de algoritmos de aprendizaje por refuerzo de última generación
Junhyuk Oh1, Greg Farquhar2, Iurii Kemaev2
1Google DeepMind, London, UK. junhyuk@google.com.
Nature
|October 22, 2025
Resumen
Las máquinas ahora pueden descubrir reglas avanzadas de aprendizaje por refuerzo (RL), superando a las diseñadas por el hombre. Este avance en la inteligencia artificial se logró a través del metaaprendizaje de las experiencias de los agentes.
Área de la Ciencia:
- Inteligencia artificial
- Aprendizaje automático
- Aprendizaje por refuerzo
Sus antecedentes:
- Los sistemas biológicos utilizan mecanismos de aprendizaje por refuerzo evolucionados.
- Los agentes artificiales actuales se basan en reglas de aprendizaje diseñadas manualmente.
- El descubrimiento de algoritmos autónomos de RL ha sido un desafío de larga data.
Objetivo del estudio:
- Demostrar que las máquinas pueden descubrir de forma autónoma las reglas de aprendizaje por refuerzo de última generación.
- Desarrollar un método para descubrir las reglas de aprendizaje basado en la experiencia a través del metaaprendizaje.
Principales métodos:
- Meta-aprendizaje de las experiencias colectivas de una población de agentes.
- La formación de agentes en una amplia gama de entornos complejos.
- Descubrir la regla RL específica que rige las actualizaciones de la política y la predicción.
Principales resultados:
- La regla RL descubierta superó a todas las reglas existentes en el índice de referencia de Atari.
- La regla descubierta superó los algoritmos RL de última generación en puntos de referencia desafiantes inéditos.
- Esto representa un avance significativo en el descubrimiento de algoritmos de aprendizaje por refuerzo.
Conclusiones:
- Es posible que las máquinas descubran de forma autónoma poderosos algoritmos de aprendizaje por refuerzo.
- La inteligencia artificial futura puede depender de algoritmos RL descubiertos automáticamente.
- Este enfoque cambia del diseño manual al descubrimiento basado en la experiencia de las reglas de aprendizaje de la IA.
Videos de Conceptos Relacionados
Reinforcement
826
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
826
Reinforcement Schedules
453
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
453
Observational Learning
824
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
824
Associative Learning
1.2K
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
1.2K
Law of Effect
2.5K
B.F. Skinner, a prominent figure in behavioral psychology, introduced operant conditioning by emphasizing the role of consequences in shaping behavior. This theory builds upon the law of effect proposed by Edward Thorndike, which posits that behaviors followed by satisfying outcomes are likely to be repeated. In contrast, those followed by unsatisfying outcomes are less likely to recur.
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
2.5K
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K