Related Experiment Video
Updated: Jul 13, 2025

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Asynchronous Deep Double Dueling Q-learning for trading-signal execution in limit order book markets
Peer Nagy1, Jan-Peter Calliess1, Stefan Zohren1,2,3
1Department of Engineering Science, Oxford-Man Institute of Quantitative Finance, University of Oxford, Oxford, United Kingdom.
Deep reinforcement learning (RL) trains a trading agent to optimize limit order placement. The RL agent effectively manages inventory and outperforms benchmark strategies in realistic market simulations.
Area of Science:
- Quantitative Finance
- Machine Learning
- Algorithmic Trading
Background:
- High-frequency trading (HFT) relies on sophisticated algorithms to execute orders rapidly.
- Limit order books (LOBs) present complex dynamics that challenge traditional trading strategies.
- Reinforcement learning (RL) offers a powerful framework for developing adaptive trading agents.
Purpose of the Study:
- To develop and evaluate a deep reinforcement learning agent for optimal limit order placement in HFT.
- To assess the agent's performance in a realistic simulated NASDAQ equity trading environment.
- To investigate the agent's ability to manage inventory and maximize trading returns.
Main Methods:
- Utilized the ABIDES limit order book simulator to create a realistic trading environment.
- Developed an RL agent using Deep Dueling Double Q-learning with APEX architecture.
- Trained the agent on historical order book data and synthetic alpha signals with varying noise levels.
Main Results:
- The RL agent learned an effective trading strategy for inventory management and order placement.
- The agent demonstrated superior performance compared to a heuristic benchmark strategy.
- The approach proved robust even with noisy trading signals.
Conclusions:
- Deep reinforcement learning is a viable approach for creating adaptive and profitable HFT strategies.
- RL agents can effectively learn optimal order placement policies in complex LOB environments.
- This study highlights the potential of RL for advancing automated trading systems.
Related Concept Videos
Dynamic Equilibrium
Second Order systems II
Basic Continuous Time Signals
The unit step function, denoted u(t), is zero for negative time values and one for positive time values, exhibiting a discontinuity at t=0. This function often represents abrupt changes, such as the step voltage introduced when turning a car's...
Observational Learning
Second Order systems I
By reinterpreting the system, one can derive the closed-loop transfer function, which...
Signal and System

