模拟限量订单交易与深度强化学习的持续行动政策深度强化学习的交易
Avraam Tsantekidis1, Nikolaos Passalis1, Anastasios Tefas1
1School of Informatics, Aristotle University of Thessaloniki, Thessaloniki, Greece.
概括
本研究介绍了一种深度强化学习 (DRL) 方法,用于使用限量订单的交易代理,克服当前机器学习 (ML) 模型的局限性. 该方法有效地模拟限制价格,并战略性地使用市场订单,增强交易策略.
科学领域:
- 量化金融 量化金融
- 计算金融是指计算金融.
- 机器学习 机器学习
背景情况:
- 在算法交易中,市场订单存在滑落风险,导致不利的执行成本.
- 限量订单提供价格保护,但存在执行风险和复杂性,限制其在机器学习 (ML) 交易系统中的使用.
- 目前的ML交易文献主要使用市场订单,忽视了限量订单的好处.
研究的目的:
- 开发一个深度强化学习 (DRL) 框架,使交易代理商能够有效地利用限量订单.
- 为了应对基于ML的交易策略中限制订单执行风险的挑战.
- 整合一种混合方法,允许基于风险评估的限制和市场订单.
主要方法:
- 提出了一种新的DRL方法,使用连续概率分布建模极限价格.
- 该框架包括一个机制,在不执行的风险超过滑动成本时,可以切换到市场订单.
- 对多种货币对进行了广泛的实验,每小时提供价格数据.
主要成果:
- 拟议的DRL方法在模拟交易代理内部的限量订单行为方面表现出有效性.
- 混合订单策略在管理执行风险和滑落方面被证明是有利的.
- 实验验证证了该方法在各种货币对中的可行性.
结论:
- 开发的DRL方法成功地将限额订单纳入自动交易系统.
- 这项研究为将DRL应用于金融市场使用限量订单策略开辟了新的途径.
- 这些发现表明,通过实现复杂的订单类型建模,基于ML的算法交易取得了重大进展.
相关概念视频
Observational Learning
222
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
222
Reinforcement Schedules
212
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
212
Steps in the Modeling Process
257
Albert Bandura's theory of observational learning identifies four critical processes: attention, retention, motor reproduction, and reinforcement or motivation.
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...
Attention is the first necessary component for observational learning. It involves focusing on what the model is doing and saying. For example, if you decide to take a drawing class to enhance your skills, you need to pay close attention to the instructor's words and hand movements. The characteristics of the model significantly...
257
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
84
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
84
Reinforcement
289
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
289
Fixed Action Patterns
16.1K
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
16.1K


