用DRL驱动的板球运动员:通过在真实和假设场景中的深度强化学习来模拟板球比赛
Mohammadreza Javadiha1, Jia Long Ji1, Wenqi Zhou1
1Computer Science Department, ViRVIG - Universitat Politécnica de Catalunya - BarcelonaTech, Barcelona, Spain.
Journal of sports sciences
|June 19, 2025
概括
深度增强学习 (DRL) 能够让虚拟代理学习球,创造出新的体育模拟. 这些DRL驱动的玩家可以复制真实比赛或探索假设条件,推进体育研究.
科学领域:
- 运动科学 运动科学 运动科学
- 人工智能的人工智能
- 计算机模拟计算模拟
背景情况:
- 深度强化学习 (DRL) 的进步使复杂的任务可以在最小的数据中进行学习.
- DRL有助于创建超越现实世界复制的体育模拟.
- 有限的先验知识场景从DRL驱动的模拟中受益.
研究的目的:
- 调查使用DRL进行模拟数据生成,用于Padel比赛.
- 复制真实的板球环境或探索假设的场景,改变参数.
- 专注于高水平的球员行为,而不是低水平的技能.
主要方法:
- 实施了一种概念验证DRL系统,用于Padel代理.
- 训练了各种各样的虚拟代理人在不同的条件下玩球.
- 在模拟匹配过程中观察到代理行为.
主要成果:
- 代理人学会协调运动类似于专业球员在特定参数组合的特定参数组合.
- 证明了DRL代理人适应改变的游戏条件的能力.
- 为分析和探索生成模拟的板球数据.
结论:
- DRL驱动的玩家为体育研究和模拟提供了强大的工具.
- DRL补充了传统模型,特别是当数据收集具有挑战性时.
- 通过DRL应用程序实现更具吸引力,更安全和更具包容性的体育运动的潜力.
相关概念视频
Reinforcement
353
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
353
Modeling and Similitude
346
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
346
Observational Learning
321
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
321
Reinforcement Schedules
243
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
243
Modeling in Therapy
155
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
155
Purposive Learning
210
E. C. Tolman emphasized the purposiveness of behavior — the idea that much of our behavior is goal-directed. For instance, employees who aim for a promotion work diligently to meet their targets. Tolman argued that when classical conditioning and operant conditioning occur, the organism acquires certain expectations. In classical conditioning, a child might fear a dog because they expect it to bite. In operant conditioning, a person might consistently work overtime because they expect a...
210


