DDSUD:在不平衡的中国情绪分析中,动态检测后续不确定性和多样性,以积极学习
Shufeng Xiong1, Yibo Si1, Guipei Zhang1
1College of Information and Management Science, Henan Agricultural University, Zhengzhou, China.
PeerJ. Computer science
|September 24, 2025
概括
本研究介绍了动态检测后续不确定性和多样性 (DDSUD),这是中国情绪分析的积极学习框架. DDSUD有效地训练使用较少标记数据的模型,在不平衡的数据集上表现优于现有的方法.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 人工智能的人工智能
背景情况:
- 对中国情绪分析的监督深度学习方法需要大量的标记数据,这是昂贵的和耗时的.
- 现有的积极学习方法与不平衡的数据集以及有效利用有限的标记数据作斗争.
研究的目的:
- 提出动态检测次序不确定性和多样性 (DDSUD),基于变压器 (BERT) 的双向编码器表示的积极学习框架.
- 为了应对数据稀缺和不平衡数据集在中国情绪结构分析中的挑战.
- 为了提高顺序标记任务的积极学习的效率和有效性.
主要方法:
- DDSUD集成了后续不确定性检测,多样性驱动的样本选择和动态加权.
- 该框架在整个积极学习代过程中以适应的方式平衡不确定性和多样性.
- 使用来自变压器的双向编码器表示 (BERT) 来进行特征提取.
主要成果:
- DDSUD使用仅50%的标记数据,实现了与完全监督方法相匹配的性能.
- 在使用相同数量的标记数据时,优于最先进的主动学习方法.
- 在低资源和不平衡的场景中表现出强大的适应性和通用性,提高了少数群体的阶级认可.
结论:
- 对于中国人的情绪分析,DDSUD提供了一种高效和有效的解决方案,特别是在资源不足和不平衡的环境中.
- 该框架对不确定性和多样性的动态调整提高了模型性能和概括性.
- 减少对大型标记数据集的需求,使先进的情绪分析更容易获得.
相关概念视频
Uncertainty: Confidence Intervals
10.2K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
10.2K
Survival Tree
388
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
388
Variability: Analysis
448
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
448
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Uncertainty: Overview
1.6K
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
1.6K
