Related Experiment Video
Updated: Jun 25, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Global and Coalition Cognition Graph Modeling for Interpretable Multiagent Reinforcement Learning
Abstract:
Multiagent reinforcement learning (MARL) has achieved impressive success across various domains, yet the neural network policies produced with MARL fail to explain their behaviors in a human-understandable way. To address this challenge, we propose global and coalition cognition graph modeling (GC2GM), a novel framework that equips agents with explicit cognitive representations integrating both coalition-level relational reasoning and individualized global state cognition. Specifically, we first introduce an information exchange graph network (IEGN) to build coalition cognition that enables agents to dynamically reason about the intentions and behaviors of teammates. We further enhance individualized global cognition of agents via imposing a mutual-information regularization that aligns the global state with its local perspective, which produces globally informed yet locally grounded representations of the environment. During execution, agents rely solely on their local information and consciousness to make decisions, adhering to the centralized training with decentralized execution (CTDE) paradigm. Additionally, GC2GM is modular and can be seamlessly integrated into existing CTDE MARL methods. Experiments on a wide range of challenging MARL benchmarks demonstrate that GC2GM achieves desirable results, striking a favorable balance between performance and interpretability.
Related Concept Videos
Observational Learning
Associative Learning
Classical conditioning, also known...
Cognitive Learning
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Multi-input and Multi-variable systems
In the absence of...
Collisions in Multiple Dimensions: Problem Solving
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...