研究文章分类的整体方法:人工智能的案例研究
Min Lu1, Lie Tang1, Xianke Zhou2
1Hangzhou Science and Technology Information Institute, Hangzhou, Zhejiang, China.
PeerJ. Computer science
|February 3, 2025
概括
本研究引入了一个深度学习组合模型,用于对人工智能 (AI) 等新兴领域的研究文章进行分类. 这种新的方法显著改善了与人工智能相关的研究的识别,优于传统的关键词方法.
科学领域:
- 计算机科学 计算机科学
- 信息科学 信息科学 信息科学
背景情况:
- 新兴科学领域的文本分类是具有挑战性的,因为边界不断变化和跨学科性质.
- 传统的基于关键字的方法由于不完整的术语列表而受到很低的回忆.
研究的目的:
- 开发和评估基于深度学习的整体方法,用于在动态研究领域准确的文章分类.
- 以人工智能 (AI) 为案例研究,解决传统方法在捕捉新兴科学领域的全部范围方面的局限性.
主要方法:
- 结合决策树,SciBERT和正则表达式匹配的集合模型被开发出来.
- 使用支持矢量机器 (SVM) 来整合单个模型的结果.
- 该方法在来自Web of Science (WoS) 库的手动标记数据集上进行了评估.
主要成果:
- 整体模型在识别AI相关文章方面实现了97%的回忆率和0.92的精度.
- 这与现有的基于搜索术语的方法相比,F1得分增加了0.15.
- 一项废弃研究证实了每个组合组件的贡献,SciBERT在其他BERT模型中表现出卓越的性能.
结论:
- 提出的深度学习组合方法有效地提高了快速发展的科学领域的研究文章的分类.
- 这种方法比传统方法有了显著的改进,特别是在跨学科和新兴领域.
- 在集合中,SciBERT显示出强大的有效性,用于识别相关的科学文献.
相关概念视频
Non-equilibrium in the Cell
4.1K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
4.1K
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K


