集体学习方法用于区分人类和计算机生成的阿拉伯文评论
Fatimah Alhayan1, Hanen Himdi2
1Department of Information Systems, College of Computer and Information Sciences, Princess Nourah Bint Abdulrahman University, Riyadh, Saudi Arabia.
PeerJ. Computer science
|December 9, 2024
概括
这项研究使用机器学习区分了人类和人工智能生成的阿拉伯评论. 计算机生成的评论显示出明显的语言模式,有助于检测假评论并提高消费者信任.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 计算语言学 计算语言学
背景情况:
- 客户评论对企业至关重要,但人工智能生成的假评论会侵蚀消费者的信任.
- 关于检测人工智能生成文本的现有研究主要集中在英语,阿拉伯语存在差距.
- 对阿拉伯语假评论进行分类的集体学习 (EL) 技术尚未得到充分探索.
研究的目的:
- 开发和评估用于分类人类与计算机生成的阿拉伯文评论的模型.
- 在这个分类任务中,研究集合学习,特别是软投票的有效性.
- 识别区分人类和人工智能生成的阿拉伯语评论的语言特征.
主要方法:
- 采用传统的机器学习,深度学习和变压器模型.
- 利用集体技术,包括软投票,结合逻辑回归 (LR) 和卷积神经网络 (CNN) 模型.
- 进行了文本分析,重点关注语音部分 (POS),情绪和语言模式.
主要成果:
- 在分类方面取得了很高的准确性,LR和CNN组合达到89.70%,相当于AraBERT的90.0%.
- 确定了显著的语言差异:人工智能评论中含有比人类评论 (0.46%) 的形容词比例要高得多 (6.3%).
- 证明了整体方法在改善假评论检测方面的有效性.
结论:
- 该研究成功地区分了人类和人工智能生成的阿拉伯评论,为企业提供了有价值的工具.
- 语言分析为假评论的特征提供了关键的见解,有助于检测.
- 推进阿拉伯自然语言处理 (NLP),并为维护市场完整性和消费者信任提供实用意义.
相关概念视频
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K


