多目标奖励概括:改善深度强化学习的性能,用于单一资产交易中的应用
Federico Cornalba1,2, Constantin Disselkamp3, Davide Scassola4
1Institute of Science and Technology Austria (ISTA), Am Campus 1, 3400 Klosterneuburg, Austria.
Neural computing & applications
|January 8, 2024
概括
本研究探讨了用于交易股票和加密货币的多目标深度强化学习. 拟议的算法显示了更好的预测稳定性,并且与单一目标方法相比,在稀疏的奖励方面表现更好.
科学领域:
- 计算金融是指计算金融.
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 传统的单一资产交易策略经常与动态的市场条件和稀疏的奖励信号作斗争.
- 深度强化学习 (DRL) 提供了自动交易的潜力,但需要仔细的奖励函数和折扣因子规范.
- 多目标优化可以为金融中复杂的决策流程提供更强大的框架.
研究的目的:
- 调查多目标深度强化学习 (MODRL) 算法的有效性,用于在股票和加密货币市场的单一资产交易.
- 评估算法在学习过程中概括奖励函数和折扣因子的能力.
- 将MODRL方法的性能和稳定性与传统的单一目标 (SO) DRL策略进行比较.
主要方法:
- 开发和实施一种包含奖励概括和折扣因子学习的新型MODRL算法.
- 对各种金融资产进行实证测试,包括加密货币 (BTCUSD,ETHUSDT,XRPUSDT) 和股票 (AAPL,SPY,NIFTY50).
- 对MODRL与SO策略进行比较分析,重点关注在稀疏奖励条件下的预测稳定性和性能.
主要成果:
- 与SO方法相比,MODRL算法展示了优越的奖励泛化能力.
- 初步的统计证据表明,在各种资产中,MODRL战略的预测稳定性得到了增强.
- 在以稀疏奖励机制为特征的场景中,MODRL算法表现出了显著的性能优势.
结论:
- 拟议的MODRL框架通过固有的处理复杂的奖励结构,为自动化单一资产交易提供了有希望的进步.
- 算法的概括奖励函数和折扣因子的能力有助于提高预测稳定性和稳定性.
- 开源代码的可用性促进了MODRL在金融交易中的进一步研究和应用.
相关概念视频
Multi-input and Multi-variable systems
106
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
106
Generalization, Discrimination, and Extinction
563
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
563
Decision Making: P-value Method
5.4K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
5.4K
Reinforcement Schedules
148
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
148
Associative Learning
375
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
375
Improving Translational Accuracy
10.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.5K


