使用新型python包生成的语法意识的短语数据集
Ebisa A Gemechu1, G R Kanagachidambaresan1
1Department of Computer Science and Engineering, Vel Tech Rangarajan Dr. Sagunthala R&D Institute of Science and Technology, Chennai 600062, Tamil Nadu, India.
本研究介绍了"Oromo-grammar",这是一个Python软件包,可以自动创建Oromo语言数据集. 它提取动词并生成语法短语,帮助NLP应用程序和语言研究.
科学领域:
- 计算语言学 计算语言学
- 自然语言处理自然语言处理.
- 非洲语言 非洲语言
背景情况:
- 手动准备数据集是劳动密集型的,容易出现错误.
- 现有的数据采集网络抓取方法也会导致数据不准确.
- 开发自动化工具对于高效的语言数据处理至关重要.
研究的目的:
- 介绍"奥罗莫语法",一个新的Python包用于自动生成奥罗莫语言数据集.
- 克服手动和网络取数据准备的局限性.
- 创建适用于NLP和语言研究的语法丰富的数据集.
主要方法:
- "奥罗莫语法"包接受原始文本文件作为输入.
- 它提取根动词并生成相应的动词干.
- 该算法将语法短语与附加词和代词合成,表示语法特征,如数字,性别和大小写.
主要成果:
- 开发了一个新的Python包"Oromo-grammar".
- 该软件包成功生成了丰富的奥罗莫语句语法数据集.
- 生成的数据集包括语法信息,如数字,性别和案例.
结论:
- "奥罗莫语法"包为奥罗莫语言数据集创建提供了一种高效和可重复的方法.
- 生成的数据集支持先进的自然语言处理 (NLP) 应用程序,包括机器翻译和语法检查.
- 这种方法可以通过系统分析和轻微修改适应其他语言.
更多相关视频
05:54Eye-tracking to Distinguish Comprehension-based and Oculomotor-based Regressive Eye Movements During Reading
Published on: October 18, 2018
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
相关概念视频
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Quantifying and Rejecting Outliers: The Grubbs Test
Detection of Gross Error: The Q Test
Data Collection I
Genetic Lingo
Data Collection by Experiments
An example of the experimental method is a public...
