Related Experiment Video
Updated: Jan 15, 2026

Pavlovian Conditioned Approach Training in Rats
Published on: February 4, 2016
Generalizable Offline Multiobjective Reinforcement Learning via Preference-Conditioned Diffuser
Abstract:
Multiobjective reinforcement learning (MORL) addresses sequential decision-making problems with multiple objectives by learning policies optimized for diverse pReferences. While traditional methods necessitate costly online interaction with the environment, recent approaches leverage static datasets containing precollected trajectories, making offline MORL the preferred choice for real-world applications. However, existing offline MORL techniques suffer from limited expressiveness and poor generalization on out-of-distribution (OOD) preferences. To overcome these limitations, we propose diffusion-based MORL (DiffMORL), a generalizable diffusion-based planning frame work for MORL. Leveraging the strong expressiveness and generation capability of diffusion models, DiffMORL further boosts its generalization through offline data mixup, which mitigates the memorization phenomenon and facilitates feature learning by data augmentation. By training on the augmented data, DiffMORL is able to condition on a given preference, whether in-distribution or OOD, to plan the desired trajectory and extract the corresponding action. Evaluations conducted on the datasets for MORL (D4MORL) benchmark demonstrate that DiffMORL achieves state-of-the-art results across nearly all tasks. Notably, it surpasses the best baseline on 14 out of 18 metrics for OOD generalization, underscoring its remarkable generalization ability in offline MORL scenarios.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Once a behavior is learned,...
Observational Learning
Multi-input and Multi-variable systems
In the absence of...
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Associative Learning
Classical conditioning, also known...

