CNER-Omni:一个统一的动态模式学习框架,用于在文本和语音中识别中文命名实体
Jinzhong Ning1, Wenxuan Mu1, Songtao Li1
1School of Information Science and Technology, Dalian Maritime University, 1 Linghai Road, Dalian, 116026, Liaoning, China.
概括
在文本,语音和多式联网数据中,CNER-Omni统一了中国命名实体识别 (CNER). 这种综合方法增强了跨模式的概括性,并以减少复杂性实现了最先进的结果.
科学领域:
- 自然语言处理自然语言处理.
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 多式联网数据在现实应用中越来越常见.
- 现有的中文命名实体识别 (CNER) 方法独立处理文本,语音和多模式输入,导致效率低下和跨模式概括性差.
- 需要一个统一的综合多式联运命名实体认可 (IMNER) 框架.
研究的目的:
- 提出CNER-Omni,这是IMNER的统一框架,将基于文本的,基于语音的和多式联网的CNER整合到一个单一的模型中.
- 引入IMAGE,一个集成多式联机生成框架,将NER视为实体意识的序列生成任务.
- 改进跨模态表示学习和模型适应性,以各种输入配置.
主要方法:
- 为多式联运数据设计了一个统一的输入表示方案.
- 开发了IMAGE,这是一个集成的多模式生成框架,利用伪模式输入进行跨模式学习.
- 整合了一种模式-成分-意识的专家混合 (MoE) 模块,用于动态适应.
- 在AISHELL-NER,CNERTA和MSRA基准测试中进行了实验.
主要成果:
- 在文本,语音和多模式NER任务中,CNER-Omni实现了最先进的性能.
- 与独立方法相比,统一框架显著降低了模型的复杂性.
- 在平面和嵌套NER设置中表现出强的性能.
- 在低资源和跨模式场景中展示了稳健性.
结论:
- CNER-Omni提供了一个有效和高效的统一框架,用于集成的多式联运命名实体识别.
- IMAGE框架能够实现卓越的跨模式表示学习和适应性.
- 拟议的方法为CNER在各种数据模式和具有挑战性的条件中设定了新的标准.
相关概念视频
Components of Language
721
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
721
Language and Cognition
688
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
688


