Related Experiment Video
Updated: Dec 13, 2025

05:41
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
9.8K
Action-specialized expert ensemble trading system with extended discrete action space using deep reinforcement
1Department of Financial Engineering, Ajou University, Yeongtong-gu, Suwon, Republic of Korea.
Plos One
|July 28, 2020
Summary
This study introduces an action-specialized expert ensemble method for reinforcement learning trading systems. The novel approach significantly enhances trading efficiency and cumulative returns compared to traditional methods.
Area of Science:
- * Quantitative Finance
- * Machine Learning
- * Algorithmic Trading
Background:
- * Existing reinforcement learning (RL) trading systems require performance enhancements.
- * Current ensemble methods lack action-specific optimization.
Purpose of the Study:
- * To propose a novel action-specialized expert ensemble method for RL trading.
- * To improve the efficiency and robustness of RL-based trading strategies.
- * To evaluate the impact of different reward functions on trading performance.
Main Methods:
- * Developed action-specialized expert models for 'buy,' 'hold,' and 'sell' actions.
- * Defined reward values correlated with specific actions and market conditions.
- * Compared proposed system's profits against single and common ensemble trading systems.
- * Analyzed performance across S&P500, Hang Seng Index, and Eurostoxx50 data.
- * Tested sensitivity using profit, Sharpe ratio, and Sortino ratio reward functions.
Main Results:
- * The proposed model demonstrated 39.1% and 21.6% greater efficiency than single and common ensemble models, respectively.
- * Expanding the action space from 3 to 11 and 21 actions increased cumulative returns by 427.2% and 856.7%.
- * Models trained with Sharpe and Sortino ratios outperformed those using profit alone, with Sortino ratio showing a slight edge.
Conclusions:
- * The action-specialized expert ensemble method offers a significant improvement for RL trading systems.
- * The approach enhances trading efficiency, robustness, and cumulative returns.
- * Reward function selection impacts performance, with risk-adjusted ratios like Sortino and Sharpe proving effective.
Related Concept Videos
State Space Representation
438
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
438
Fixed Action Patterns
17.2K
A fixed action pattern (FAP) is a specific, hard-wired sequence of behaviors that occurs in response to an external stimulus, called a sign stimulus. The behavior is “fixed” because it is essentially unchangeable—proceeding similarly across individuals of a species every time it occurs.
17.2K
Reinforcement
696
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
696
Reinforcement Schedules
365
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
365
Multi-input and Multi-variable systems
307
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
307
Generalization, Discrimination, and Extinction
1.2K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.2K

