Related Experiment Video
Updated: Jan 8, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
Deep Q-Managed: a new framework for multi-objective deep reinforcement learning.
Richardson Menezes1,2, Thiago Henrique Freire de Oliveira3, Luiz Paulo de Souza Medeiros1
1Postgraduate Program in Electrical and Computer Engineering, Federal University of Rio Grande do Norte, Natal, Brazil.
Deep Q-Managed, a novel multi-objective reinforcement learning (MORL) algorithm, discovers all Pareto Front policies. It uses deep learning to enhance multi-objective optimization in deterministic environments.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Reinforcement Learning
Background:
- Multi-objective optimization presents challenges like the curse of dimensionality and overestimation bias.
- Existing algorithms struggle to efficiently identify the entire Pareto Front in complex scenarios.
Purpose of the Study:
- Introduce Deep Q-Managed, a novel multi-objective reinforcement learning (MORL) algorithm.
- Enable the discovery of all policies within the Pareto Front.
- Enhance multi-objective optimization using deep learning techniques.
Main Methods:
- Integrate deep learning techniques, specifically Double and Dueling Networks, into the MORL framework.
- Mitigate dimensionality and overestimation bias through advanced network architectures.
- Apply the algorithm to deterministic episodic environments.
Main Results:
- Deep Q-Managed successfully attains non-dominated multi-objective policies across various Pareto Front complexities (convex, concave, mixed).
- Consistent achievement of maximum hypervolume values on standard MORL benchmarks (DST, BST, MBST).
- Demonstrated ability to locate all Pareto Front points effectively.
Conclusions:
- Deep Q-Managed is a proficient algorithm for discovering comprehensive Pareto Fronts in deterministic settings.
- The algorithm shows robustness and versatility for applications in robotics, finance, and healthcare.
- Future work will focus on extending the algorithm to stochastic environments.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...