Related Experiment Video
Updated: Jul 16, 2026

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends.
Yihao Ding1, Siwen Luo1, Yue Dai2
1The University of Western Australia, Crawley, Australia.
Summary
This survey explores Multimodal Large Language Models (MLLMs) for Visually Rich Document Understanding (VRDU). It highlights techniques, training strategies, and challenges for advancing MLLM-based VRDU systems.
Area of Science:
- Computer Science
- Artificial Intelligence
- Natural Language Processing
Background:
- Visually Rich Document Understanding (VRDU) is crucial for interpreting complex documents.
- Multimodal Large Language Models (MLLMs) show promise for VRDU, using both OCR-based and OCR-free methods.
Purpose of the Study:
- To survey recent advances in MLLM-based VRDU.
- To identify emerging trends and future research directions in the field.
Main Methods:
- Review of MLLM techniques for feature representation and integration (textual, visual, layout).
- Analysis of MLLM training paradigms (pretraining, instruction tuning, strategies).
- Examination of challenges and emerging trends (data scarcity, multi-page/lingual documents, Retrieval-Augmented Generation, agentic frameworks).
Main Results:
- MLLMs offer powerful capabilities for information extraction from diverse document types.
- Key advancements lie in multimodal feature fusion and sophisticated training methodologies.
- Addressing challenges is vital for developing robust and adaptable VRDU systems.
Conclusions:
- MLLM-based VRDU is a rapidly evolving field with significant potential.
- Further research is needed to overcome current limitations and enhance system scalability and reliability.
- This survey provides a roadmap for future advancements in MLLM-driven document understanding.