使用RoBERTa-Large语言模型和深层嵌入的新闻标题中的点击诱检测.
Fawaz Khaled Alarfaj1, Amara Muqadas2, Hikmat Ullah Khan3
1Department of Management Information Systems, School of Business, King Faisal University, Al Ahsa, Saudi Arabia. falarfaj@kfu.edu.sa.
Scientific reports
|December 2, 2025
概括
本研究介绍了RoBERTa-Large,一个变压器模型,用于高级点击诱标题检测. 它达到97%的准确性,在数字新闻分析中表现优于传统的机器学习和深度学习方法.
科学领域:
- 人工智能的人工智能
- 自然语言处理 (Natural Language Processing) 是一种自然语言处理.
- 数字新闻分析 数字新闻分析
背景情况:
- 点击诱惑标题检测是数字新闻分析中的一个具有挑战性的研究领域.
- 现有的研究主要使用传统的机器学习 (ML) 和深度学习 (DL) 模型.
- 需要先进的模型来捕捉复杂的语言细微差别.
研究的目的:
- 引入基于变压器的架构RoBERTa-Large,用于自动检测点击诱标题.
- 评估RoBERTa-Large与最先进的ML和DL方法的有效性.
- 使用可解释AI (XAI) 方法增强模型的解释性.
主要方法:
- 使用了RoBERTa-Large,这是一个具有自我注意力机制的变压器架构.
- 采用了各种各样的文本功能:TF-IDF,语音部分标记,n-grams,word2Vec,快速文本和句子嵌入.
- 根据已建立的ML和DL模型评估分类性能.
- 应用于可解释AI (XAI) 的应用LIME和SHAP.
主要成果:
- 罗伯塔-Large实现了97%的分类准确度.
- 该模型显著优于现有的ML和DL方法.
- XAI 方法为模型的决策过程提供了洞察力.
结论:
- RoBERTa-Large在点击引诱标题检测方面表现出卓越的表现.
- 基于变压器的模型在捕获上下文和语义信息方面具有优势.
- 可解释的人工智能增强了自动新闻分析系统的可信度和理解力.
相关概念视频
Detection of Black Holes
2.5K
Although black holes were theoretically postulated in the 1920s, they remained outside the domain of observational astronomy until the 1970s.
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
Their closest cousins are neutron stars, which are composed almost entirely of neutrons packed against each other, making them extremely dense. A neutron star has the same mass as the Sun but its diameter is only a few kilometers. Therefore, the escape velocity from their surface is close to the speed of light.
Not until the 1960s, when the first neutron...
2.5K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
lncRNA - Long Non-coding RNAs
3.5K
3.5K
lncRNA - Long Non-coding RNAs
9.7K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
9.7K
Regulated mRNA Transport
3.3K
3.3K
