在餐厅评论中预测情绪极性,使用基于进化XGBoost的顺序回归方法
Dana A Al-Qudah1,2, Ala' M Al-Zoubi2, Alexandra I Cristea3
1King Abdullah II School for Information Technology, The University of Jordan, Amman, Jordan.
PeerJ. Computer science
|February 3, 2025
概括
本研究介绍了一种新的PSO-XGBoost模型,用于分析约旦食品行业的客户评论. 先进的情绪分析模型准确地分类阿拉伯语和英语反,改善业务洞察力.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 业务分析 业务分析
背景情况:
- 数字化转型需要对移动应用程序的多语言客户反进行高级分析.
- 约旦的食品和餐饮行业在理解在线评论方面面临着挑战.
- 情绪分析和机器学习为从文本数据中提取有价值的商业见解提供了机会.
研究的目的:
- 开发和评估一种新的情绪分析模型,用于对约旦食品和餐厅行业的客户评论进行分类.
- 解决自动文本分析和用户情绪状态识别方面的挑战.
- 为企业增强对阿拉伯语和英语客户反的理解.
主要方法:
- 利用情绪极性和生物识别技术进行全面的审查分析.
- 从Talabat应用程序收集并准备阿拉伯语和英语评论.
- 开发了一个四个阶段的模型:数据收集,准备,构建和评估.
- 提出了一个整数回归模型,将极端梯度提升 (XGBoost) 和粒子群集优化 (PSO) 结合起来,用于多语言预测.
主要成果:
- 拟议的PSO-XGBoost算法在与支持矢量机 (SVM) 和其他方法相比显示出更高的性能.
- 在英语和阿拉伯语数据集中实现较低的根平均平方误差 (RMSE).
- 具体来说,PSO-XGB模型在阿拉伯评论中产生了0.7722的RMSE,超过了PSO-SVM (0.9988).
结论:
- PSO-XGBoost模型提供了一种有效的方法,用于对多语言在线评论的情绪分析.
- 这项研究为企业提供了一个强大的工具,以更好地了解客户反和改进服务.
- 这些发现突显了先进机器学习技术在应对现实世界商业挑战方面的潜力.
相关概念视频
Predicting Reaction Outcomes
8.2K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
8.2K
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Ordinal Level of Measurement
22.9K
The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. For analysis, data are classified into four levels of measurement—nominal, ordinal, interval, and ratio.
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
Data measured using an ordinal scale are similar to nominal scale data, but there is one major difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks...
22.9K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K


