Related Experiment Video
Updated: May 25, 2025

09:44
Author Spotlight: Advancing Large-Scale Neural Dynamics Through HD-MEA Technology
Published on: March 8, 2024
4.6K
ZS-MNET: A zero-shot learning based approach to multimodal named entity typing.
Baohang Zhou1, Ying Zhang1, Kehui Song2
1College of Computer Science, VCIP, DISSec, TMCC, TBI Center, Nankai University, Tianjin 300350, China.
Summary
This study introduces a zero-shot multimodal named entity typing (NET) model, ZS-MNET, that uses both text and images. It effectively recognizes new entity types without retraining, outperforming traditional methods.
Area of Science:
- Natural Language Processing
- Computer Vision
- Machine Learning
Background:
- Named Entity Typing (NET) on social media typically uses only text, ignoring valuable visual information.
- Existing NET methods struggle with evolving entity types and require retraining for new categories.
- Multimodal data presents challenges and opportunities for recognizing diverse and emerging named entities.
Purpose of the Study:
- To develop a novel zero-shot learning based multimodal NET (ZS-MNET) model.
- To leverage both textual and visual modalities for improved named entity recognition.
- To enable the recognition of previously unseen entity types without additional training.
Main Methods:
- The ZS-MNET model integrates text and image data using pre-trained transformer-based models like BERT and ViT.
- Fine-grained multimodal representations are generated to capture semantic correlations between data and entity types.
- A fusion approach is employed to combine different multimodal representations for enhanced feature modeling.
Main Results:
- Multimodal data significantly enhances performance in the NET task.
- The proposed ZS-MNET model demonstrates superior performance compared to traditional zero-shot NET (ZS-NET) approaches.
- The model successfully recognizes previously unseen named entity types in a zero-shot manner.
Conclusions:
- Integrating visual and textual modalities is crucial for advancing Named Entity Typing.
- The ZS-MNET model offers a robust solution for recognizing diverse and emerging entity types.
- Zero-shot learning combined with multimodal data provides a scalable approach for NET in dynamic environments.

