Related Experiment Video
Updated: Jan 8, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.0K
CNER-Omni: A unified dynamic modality learning framework for Chinese named entity recognition across text and speech
Jinzhong Ning1, Wenxuan Mu1, Songtao Li1
1School of Information Science and Technology, Dalian Maritime University, 1 Linghai Road, Dalian, 116026, Liaoning, China.
Summary
CNER-Omni unifies Chinese Named Entity Recognition (CNER) across text, speech, and multimodal data. This integrated approach enhances cross-modal generalization and achieves state-of-the-art results with reduced complexity.
Area of Science:
- Natural Language Processing
- Artificial Intelligence
- Machine Learning
Background:
- Multimodal data is increasingly common in real-world applications.
- Existing Chinese Named Entity Recognition (CNER) methods treat text, speech, and multimodal inputs independently, leading to inefficiencies and poor cross-modal generalization.
- There is a need for a unified framework for Integrated Multimodal Named Entity Recognition (IMNER).
Purpose of the Study:
- To propose CNER-Omni, a unified framework for IMNER that consolidates text-based, speech-based, and multimodal CNER into a single model.
- To introduce IMAGE, an Integrated Multimodal Generation framework that treats NER as an entity-aware sequence generation task.
- To improve cross-modal representation learning and model adaptability to various input configurations.
Main Methods:
- Designed a unified input representation schema for multimodal data.
- Developed IMAGE, an Integrated Multimodal Generation framework utilizing pseudo-modal inputs for cross-modal learning.
- Incorporated a modality-composition-aware mixture-of-experts (MoE) module for dynamic adaptation.
- Conducted experiments on AISHELL-NER, CNERTA, and MSRA benchmarks.
Main Results:
- CNER-Omni achieved state-of-the-art performance across text, speech, and multimodal NER tasks.
- The unified framework significantly reduced model complexity compared to independent approaches.
- Demonstrated strong performance in both flat and nested NER settings.
- Showcased robustness in low-resource and cross-modal scenarios.
Conclusions:
- CNER-Omni provides an effective and efficient unified framework for Integrated Multimodal Named Entity Recognition.
- The IMAGE framework enables superior cross-modal representation learning and adaptability.
- The proposed method sets a new standard for CNER across diverse data modalities and challenging conditions.
Related Concept Videos
Components of Language
721
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
721
Language and Cognition
688
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
688

