Related Experiment Video
Updated: Apr 30, 2026

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
An image and legend text alignment model based on position guided enhanced attention mechanism
ZiQuan Wang1, Jin Han2, Qiang Li3
1Software College, Nanjing University of Information Science and Technology, Nanjing, 210044, Jiangsu, China. 202312490701@nuist.edu.cn.
This study introduces a novel fusion model for fine-grained image-text matching in document images. The proposed model significantly improves accuracy in understanding document structure by effectively aligning image regions with legend texts.
Area of Science:
- Computer Science
- Artificial Intelligence
- Document Image Analysis
Background:
- Understanding document image structure is vital for information extraction.
- Current methods struggle with fine-grained semantic matching between image regions and legend texts.
- Existing models primarily focus on overall layout analysis, lacking detailed image-text correlation.
Purpose of the Study:
- To develop an effective model for fine-grained matching between image regions and legend texts in document images.
- To enhance the understanding of document image structure through accurate image-text relationship identification.
- To address the limitations of current methods in detailed semantic alignment.
Main Methods:
- A fusion model utilizing a position-guided enhanced attention mechanism.
- Separate processing of image and legend text features using heterogeneous structures.
- Integration of region coordinates as auxiliary features for local structure perception.
- Application of a cross-attention mechanism for deep feature fusion between modalities.
Main Results:
- Achieved matching accuracy rates of 90.34% on the Ancient Book Digitization Dataset.
- Attained a matching accuracy rate of 68.5% on the DocBank Dataset.
- Outperformed existing mainstream methods in fine-grained structural understanding tasks.
Conclusions:
- The proposed fusion model demonstrates superior performance in fine-grained image-text matching.
- The position-guided enhanced attention and cross-attention mechanisms are effective for detailed document structure analysis.
- This approach significantly advances the field of document image understanding and information extraction.
More Related Videos
13:00Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Related Concept Videos
Gestalt Principles of Perception
The Anchoring-and-Adjustment Heuristic