Related Experiment Video
Updated: Mar 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Hierarchical multi-agent reinforcement learning for retrieval-augmented industrial document question answering
Yihong Qian1, Baoli Han2, Yufeng Yuan2
1State Grid Zhejiang Electric Power Co., Ltd., Shaoxing Power Supply Company, Shaoxing, Zhejiang, China. qianyihong9760@163.com.
Abstract:
Multimodal industrial documents-such as operation manuals, circuit diagrams, and parameter tables-contain domain knowledge distributed across text, images, and document layout. However, most existing retrieval-augmented generation (RAG) frameworks rely on static retrieval and fusion policies with fixed modality weights and uniform retrieval depth, making them less adaptable to diverse query intents and dynamic cross-modal dependencies. As a result, they often retrieve incomplete evidence and yield suboptimal reasoning in complex long-document scenarios. To address these challenges, we propose MARL-RAGDoc, a hierarchical multi-agent reinforcement learning framework for multimodal retrieval-augmented reasoning. A high-level coordinator agent dynamically allocates modality weights and retrieval depth based on query characteristics, while specialized text, image, and table agents perform fine-grained evidence selection within their respective candidate pools. A collaborative reasoning module integrates the retrieved evidence and provides hierarchical reward signals to continuously optimize retrieval policies. Experimental results on multiple multimodal document benchmarks demonstrate that MARL-RAGDoc consistently outperforms baselines in both retrieval accuracy and reasoning performance, while remaining computationally efficient. Our code and dataset are publicly available at https://github.com/Yihong-Q/MARL-RAGDoc .
Related Concept Videos
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Associative Learning
Classical conditioning, also known...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Reinforcement Schedules
Once a behavior is learned,...
Observational Learning