一个罗马乌尔都文本的数据集,其中有拼写变化,用于句子级别情绪分析
Mudasar Ahmed Soomro1, Rafia Naz Memon2, Asghar Ali Chandio3,4
1Department of Information Technology, Quaid-e-Awam University of Engineering, Science & Technology, Nawabshah, Pakistan.
Data in brief
|December 31, 2024
概括
这项研究收集了罗马乌尔都语数据集,包括5244个字的拼写变化和28090条评论. 这些资源解决了罗马乌尔都语的挑战.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 数字人文学科 数字人文学科
背景情况:
- 罗马乌尔都语是广泛用于在线非正式通信的脚本.
- 在罗马乌尔都语中缺乏标准化的拼写,这给自然语言处理 (NLP) 任务带来了挑战.
- 现有的NLP工具经常与罗马乌尔都文本的多样性和不一致性作斗争.
研究的目的:
- 为罗马乌尔都语策划全面的数据集.
- 促进罗马乌尔都语NLP的研究和开发.
- 为了解决罗马乌尔都语拼写的非标准化性质.
主要方法:
- 收集了5244个罗马乌尔都语单词的数据集,每个单词都有1-5个拼写变化.
- 从七个在线来源收集了28,090个罗马乌尔都语评论的数据集.
- 由乌尔都语语言领域专家进行注释的评论情绪 (非常积极,积极,中立,负面,非常负面).
主要成果:
- 建立了一个独特的罗马乌尔都语单词数据集,捕捉拼写变化.
- 创建了一个大规模的,多类情感注释的罗马乌尔都文评论数据集.
- 为罗马乌尔都语的计算分析提供了宝贵的资源.
结论:
- 开发的数据集对于推进罗马乌尔都语NLP研究至关重要.
- 这些资源将有助于为罗马乌尔都语建立更强大的情绪分析和语言处理模型.
- 该研究强调了解决数字通信中的语言差异的重要性.
相关概念视频
Improving Translational Accuracy
2.5K
2.5K
Variation
6.6K
An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation, which is the square root of variance.
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
6.6K
Mean Absolute Deviation
2.5K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
2.5K
Alternative RNA Splicing
3.6K
3.6K
What is Variation?
11.0K
Apart from the measures of central tendency, distribution, outliers, and the changing characteristics of data with time, an important characteristic of any data set is its variation or spread. In some data sets, the data values are concentrated closely near the mean; in others, the data values are more widely spread out from the mean.
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
11.0K
Data Validation
4.7K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
4.7K


