基于因素的深度强化学习用于资产配置:对静态和动态β奖励设计的比较分析
1Seoul AI School, aSSIST University, Seoul, Republic of Korea.
PloS one
|December 30, 2025
概括
本研究介绍了基于因素的资产配置 (FDRL) 深度强化学习框架. 它通过纳入因子风险增加来增强投资策略,改善各种资产类别的风险调整回报.
科学领域:
- 量化金融 量化金融
- 机器学习 机器学习
- 计算经济学的计算经济学
背景情况:
- 传统的资产配置与市场波动和结构性破裂作斗争.
- 深度强化学习 (DRL) 往往忽视了对风险和适应至关重要的因素暴露.
- 因素暴露 ([公式:见文本]) 影响风险调整回报和适应性投资反应.
研究的目的:
- 为资产配置 (FDRL) 制定基于因素的深度强化学习框架.
- 将因素敏感性纳入DRL状态表示和奖励设计,以改善资产配置.
- 评估不同基于因素的奖励结构在不同市场条件下的表现.
主要方法:
- 开发了一个FDRL框架,使用滚动回归来估计因子灵敏度 (势头,波动性,偏差,体积).
- 实现了PPO,SAC和TD3算法,具有五种奖励变体 (夏普,索尔蒂诺,静态-[公式:见文本],动态-[公式:见文本],动量-[公式:见文本]).
- 测试了跨股票,加密货币,宏观经济工具和混合投资组合的框架,并进行了广泛的稳定性检查.
主要成果:
- 基于因素的奖励在各类资产中产生了不同但可解释的结果.
- 在股票方面,Dynamic-[公式:见文本]将年化回报提高到23-24%和夏普比率提高到1.27.
- 加密货币表现出高回报率 (38-43%),但制度敏感性;宏观工具受益于静态稳定性;混合投资组合表现出色的势头.
结论:
- FDRL框架提供了一种新的资产配置方法,将因子风险整合到DRL中.
- 对因素敏感的奖励表明异质但有价值的结果,改善绩效和风险管理.
- 该框架在动态投资策略中将适应性响应与可解释性和风险纪律相协调.
更多相关视频
07:05Operant Protocols for Assessing the Cost-benefit Analysis During Reinforced Decision Making by Rodents
Published on: September 10, 2018
6.4K
08:24The Joint Effect of Social Comparison and Social Distance on Evaluation of Intertemporal Choice Outcomes in Event-related Potential Studies
Published on: August 25, 2023
1.1K
相关概念视频
Reinforcement Schedules
429
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
429
Equity Theory
228
Equity theory explains how our sense of fairness influences the dynamics of close relationships. Rooted in social psychology, the theory posits that individuals evaluate fairness by comparing the ratio of their contributions to the rewards they receive. Relationship satisfaction is highest when these ratios are perceived as balanced between partners, promoting mutual reciprocity and a sense of justice.Equity vs. Equality in RelationshipsEquity is distinct from equality. Fairness does not...
228
Decision Making: P-value Method
6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K
Dynamic Equilibrium
61.4K
A reversible chemical reaction represents a chemical process that proceeds in both forward (left to right) and reverse (right to left) directions. When the rates of the forward and reverse reactions are equal, the concentrations of the reactant and product species remain constant over time and the system is at equilibrium. A special double arrow is used to emphasize the reversible nature of the reaction. The relative concentrations of reactants and products in equilibrium systems vary greatly;...
61.4K
Reinforcement
781
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
781
Actuarial Approach
276
The actuarial approach, a statistical method originally developed for life insurance risk assessment, is widely used to calculate survival rates in clinical and population studies. This method accounts for participants lost to follow-up or those who die from causes unrelated to the study, ensuring a more accurate representation of survival probabilities.
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
276
