基于BERT基础模型的电影评级研究
Weijun Ning1, Fuwei Wang2, Weimin Wang1
1School of Artificial Intelligence and Software, LiaoNing Petrochemical University, Fushun, 113001, China.
Scientific reports
|March 18, 2025
概括
这项研究增强了电影评论分类的BERT模型,通过解决远程依赖和数据偏差来提高准确性和公平性. 修改后的模型在IMDb数据集上表现更好.
科学领域:
- 自然语言处理 (NLP) 是一种自然语言处理.
- 深度学习 (Deep Learning) 是一种深度学习.
- 情绪分析 情绪分析
背景情况:
- 电影评论对于用户的电影选择和内容推至关重要.
- 手动审查分类是耗时的,劳动密集的,主观的.
- 像BERT这样的深度学习模型提供自动分类,但在处理长文本和潜在偏差方面存在局限性.
研究的目的:
- 提高BERT模型在电影评审分类中的表现.
- 在漫长的审查中,解决捕捉远程依赖和局部特征的挑战.
- 为了减轻因性别和种族等敏感属性而产生的模型偏见.
主要方法:
- 实现了基于注意力的动态位置偏移编码机制.
- 引入了一个动态加权的聚变聚合策略,结合了平均,最大和自我注意力聚合.
- 在预处理过程中减轻了敏感属性 (性别,种族),并用于中性样本的数据增强 (EDA,噪声注入).
主要成果:
- 改进的BERT模型证明了功能提取和位置信息处理的改进.
- 偏差减小技术和数据增强增强模型概括.
- 在F1得分上获得了0.73%的提高,在IMDb数据集上获得了0.90%的准确性改善.
结论:
- 拟议的改进有效地提高了BERT对电影评审分类的能力.
- 该研究成功地解决了与远程依赖和数据偏差相关的局限性.
- 修改后的模型为自动化审查分析提供了更准确,更公平,更强大的解决方案.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
23
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
23
Friedman Two-way Analysis of Variance by Ranks
127
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
127
Ranks
214
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
214
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
37
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
37
Review and Preview
6.9K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
6.9K


