巴巴萨:一个大规模的孟加拉语方面基于情绪分析数据集.
1North South University (register @ northsouth edu), Plot # 15, Block # B, Bashundhara R/A, Dhaka 1229, Bangladesh.
Data in brief
|March 9, 2026
概括
研究人员开发了BABSA,这是一个大型的孟加拉语数据集,用于基于方面的情绪分析 (ABSA). 该资源支持孟加拉语语言模型的细粒度情感分类和方面提取.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
背景情况:
- 孟加拉语基于方面情感分析 (ABSA) 的研究受到缺乏全面,高质量的数据集的阻碍.
- 孟加拉语在数字通信中的广泛使用需要专门的情绪分析资源.
研究的目的:
- 介绍BABSA,这是孟加拉ABSA的一个新型数据集.
- 促进方面提取和方面特定的情绪分类在孟加拉语.
- 为孟加拉语解决细粒度情绪分析资源的缺口.
主要方法:
- 从五个不同的来源编译了15860个实例,包括从网络上取的新闻数据.
- 实施了严格的三通手动注释协议,对方面术语和情感有明确的指导方针.
- 确保了高标注质量,标注者间协议得分为0.84 (科恩的卡帕).
主要成果:
- 开发了BABSA,一个覆盖21个领域的大型数据集,对方面术语和情绪进行了详细的注释.
- 数据集包括用于高级下游分析的元数据,并分为列车测试集.
- 达成了高度的注释者间协议,确保注释的可靠性.
结论:
- 巴巴萨为推进孟加拉ABSA研发提供了至关重要的资源.
- 数据集的规模,多样性和细粒度的注释适合训练和评估各种NLP模型.
- 公共可用性BABSA促进可复制性和进一步研究孟加拉语理解.
相关概念视频
Attitudes
33.5K
Attitude is our evaluation of a person, an idea, or an object. We have attitudes for many things ranging from products that we might pick up in the supermarket to people around the world to political policies. Typically, attitudes are favorable or unfavorable: positive or negative (Eagly & Chaiken, 1993). And, they have three components: an affective component (feelings), a behavioral component (the effect of the attitude on behavior), and a cognitive component (belief and knowledge;...
33.5K
Cluster Sampling Method
15.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.3K
Mean Absolute Deviation
3.6K
The mean absolute deviation is also a measure of the variability of data in a sample. It is the absolute value of the average difference between the data values and the mean.
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
3.6K
Sampling Distribution
18.7K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
18.7K
Classification of Signals
1.5K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.5K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K

