Related Experiment Video
Updated: Oct 3, 2025

The Modular Design and Production of an Intelligent Robot Based on a Closed-Loop Control Strategy
Published on: October 14, 2017
MSPM: A modularized and scalable multi-agent reinforcement learning-based system for financial portfolio management
Zhenhan Huang1, Fumihide Tanaka2
1Graduate School of Science and Technology, University of Tsukuba, Tsukuba, Ibaraki, Japan.
This study introduces a novel multi-agent reinforcement learning system for financial portfolio management, enhancing scalability and reusability in dynamic markets. The proposed modular system significantly boosts returns compared to traditional strategies.
Area of Science:
- Artificial Intelligence
- Computational Finance
- Machine Learning
Background:
- Existing reinforcement learning (RL) approaches for financial portfolio management (PM) lack scalability and reusability for dynamic market conditions.
- Current RL systems struggle with varying asset numbers and heterogeneous data, requiring ad-hoc agent training.
- A modular design is needed for adaptable and reusable asset-specific agents in RL-based PM.
Purpose of the Study:
- To propose a novel multi-agent RL-based system for financial portfolio management (MSPM) that addresses scalability and reusability challenges.
- To introduce a modular architecture with reusable asset-dedicated agents for enhanced adaptability in financial markets.
Main Methods:
- Developed a multi-agent RL system (MSPM) with asynchronously updated Evolving Agent Modules (EAMs) and Strategic Agent Modules (SAMs).
- EAMs utilize Deep Q-network (DQN) agents for heterogeneous data processing and signal generation for specific assets.
- SAMs employ Proximal Policy Optimization (PPO) agents for portfolio optimization, connecting multiple EAMs for asset reallocation.
Main Results:
- MSPM demonstrated significant outperformance over five baselines on 8-year U.S. stock market data, improving Accumulated Rate of Return (ARR) by at least 186.5% compared to Constant Rebalanced Portfolio (CRP).
- The modular design and reusable EAMs proved crucial, with EAM-enabled MSPMs showing at least 1341.8% higher ARR than EAM-disabled versions across four different portfolios.
- MSPM effectively addresses scalability and reusability, outperforming baselines in ARR, Daily Rate of Return (DRR), and Sortino Ratio (SR).
Conclusions:
- The proposed MSPM system, with its modular architecture and reusable EAMs, effectively enhances scalability and reusability in RL-based financial portfolio management.
- MSPM offers a robust solution for adapting to ever-changing markets and heterogeneous data, significantly improving investment returns.
- The findings underscore the importance of modularity and agent reusability for advancing RL applications in complex financial domains.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Observational Learning
Reinforcement Schedules
Once a behavior is learned,...

