在人工智能中进行协作狩猎,使用深度强化学习
Kazushi Tsutsui1,2, Ryoya Tanaka2,3, Kazuya Takeda1,4
1Graduate School of Informatics, Nagoya University, Nagoya, Japan.
eLife
|May 7, 2024
概括
合作狩猎不需要高级认知. 基于经验的简单决策可以解释捕食者群体中复杂的协调,挑战以前关于大脑大小和社会行为的假设.
科学领域:
- 行为生态学 行为生态学
- 计算神经科学是一种神经科学.
- 进化生物学 进化生物学
背景情况:
- 合作狩猎被认为需要高水平的认知和大脑.
- 最近对小脑脊椎动物的协作狩猎的观察挑战了这一观念.
研究的目的:
- 调查复杂的协作狩猎策略是否可以从简单的决策过程中产生.
- 探索协调掠食者行为的认知要求.
主要方法:
- 使用计算式多代理模拟.
- 使用深度强化学习技术.
- 基于内部表现和先前经验的模型捕食者决策.
主要成果:
- 证明了复杂的协调可以从简单的,基于经验的决策规则中产生.
- 表明捕食者协调是强大的对抗不可预测的猎物行为.
- 发现依赖于距离的内部代表是新兴合作的关键.
结论:
- 协作狩猎不一定取决于先进的认知能力.
- 简单的决策机制可以解释捕食者的复杂社会行为.
- 研究结果提供了关于自然界中社交和合作的演变的见解.
相关概念视频
Observational Learning
166
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
166
Reinforcement
202
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
202
Predator-Prey Interactions
16.2K
Predators consume prey for energy. Predators that acquire prey and prey that avoid predation both increase their chances of survival and reproduction (i.e., fitness). Routine predator-prey interactions elicit mutual adaptations that improve predator offenses, such as claws, teeth, and speed, as well as prey defenses, including crypsis, aposematism, and mimicry. Thus, predator-prey interactions resemble an evolutionary arms race.
16.2K
Purposive Learning
118
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
118
Masking and Demasking Agents
2.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.4K
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K


