HARL-TRADE:第2レベルの高周波取引のための階層的な適応強化学習フレームワーク
Hao Shi1, Xinting Zhang2, Desheng Wu2
1School of Computer Science and Technology, University of the Chinese Academy of Sciences, Beijing, China.
Chaos (Woodbury, N.Y.)
|February 18, 2026
まとめ
この研究は,注意に基づくメタエージェントを使用して,高周波取引 (HFT) の適応的階層的枠組みを導入します. それは,不安定な市場での適応力を高め,有意義なリターンで既存の方法を上回ります.
科学分野:
- 定量金融 定量金融とは
- 人工知能 (AI) とは,人工知能 (AI) のことです.
- アルゴリズムによる取引 (Algorithmic Trading) とは
背景:
- 高周波取引 (HFT) は,不安定な市場条件に適応する戦略を必要とします.
- 既存の分散型サブエージェントフレームワークは,市場条件の厳格な配分により,適応性が限られている.
研究 の 目的:
- HFTにおけるダイナミックなサブエージェントの調整のための注意に基づくメタエージェントを持つ新しい階層的枠組みを提案する.
- 多様な市場体制をナビゲートする際の適応性とパフォーマンスを向上させる.
主な方法:
- 注意に基づくメタエージェントを含む階層的な枠組みを開発しました.
- 最適なサブエージェントの重量調整のための市場埋め込みと強化学習の活用.
- ダイナミックなサブエージェントの割り当てと複数頭の注意力メカニズムを実装しました.
主要な成果:
- 提案されたフレームワークは,42.15%の総リターンと,HFTの過去のデータで4.19のシャープ比率を達成しました.
- 最先端のベースラインと比較して優れたパフォーマンスを実証した.
- アブラーション研究では,ダイナミックな割り当てと注意力メカニズムの有効性が確認されました.
結論:
- 注意に基づく階層的枠組みは,高周波取引における優れた適応性とパフォーマンスを提供します.
- メタエージェントによるサブエージェントのダイナミックな調整は,市場の変動や移行を効果的に対処します.
さらに関連する動画
07:31A Computerized Functional Skills Assessment and Training Program Targeting Technology Based Everyday Functional Skills
Published on: February 13, 2020
6.8K
11:09RBDT: A Computerized Task System based in Transposition for the Continuous Analysis of Relational Behavior Dynamics in Humans
Published on: July 17, 2021
2.5K
関連する概念動画
Real-World Application of Classical Conditioning
1.6K
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
1.6K
Reinforcement Schedules
541
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
541
Observational Learning
1.0K
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
1.0K
