A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends.

Yihao Ding1, Siwen Luo1, Yue Dai2

  • 1The University of Western Australia, Crawley, Australia.

Findings of ACL. ACL
|July 15, 2026
PubMed
Summary

This survey explores Multimodal Large Language Models (MLLMs) for Visually Rich Document Understanding (VRDU). It highlights techniques, training strategies, and challenges for advancing MLLM-based VRDU systems.

Related Concept Videos