预测撤回的研究:一个数据集和机器学习方法
Aaron H A Fletcher1, Mark Stevenson2
1School of Computer Science, The University of Sheffield, Regent Court, Sheffield, S1 4DP, UK. ahafletcher1@sheffield.ac.uk.
Research integrity and peer review
|June 10, 2025
概括
机器学习模型可以帮助识别撤回的科学文章,提高研究完整性. 这项研究开发了一个数据集和评估模型,发现传统分类器和Llama 3.2在预测收缩方面具有竞争力.
科学领域:
- 图书统计学 图书统计学
- 科学出版业的科学出版.
- 机器学习是机器学习.
背景情况:
- 收回的科学文章损害了研究的完整性,并可能延续错误信息.
- 开发识别撤回的方法对于保持可靠的科学记录至关重要.
研究的目的:
- 创建一个全面的数据集的收回的文章与文献资料元数据.
- 训练和评估机器学习 (ML) 模型,用于预测文章撤回.
- 通过废弃性研究评估ML分类器中的特征重要性.
主要方法:
- 通过整合Retraction Watch和OpenAlex数据构建了一个开放访问数据集.
- 一个经过案例控制的设计将收缩的物品与非收缩的对应物品配对.
- 包括传统分类器和语言模型在内的各种ML模型被训练并使用准确性,精度,回忆和F1分数进行评估.
主要成果:
- 拉玛3.2基本模型的整体准确性很高.
- 随机森林在未收缩的物品中获得了0.687的精度;拉玛3.2在收缩的物品中获得了0.683的精度.
- 传统的ML分类器通常表现优于语境语言模型,Llama 3.2显示出竞争性表现.
结论:
- 机器学习有效地帮助识别撤回的研究,尽管没有一个单一的模型主导了所有指标.
- 这些发现支持出版商和审稿人开发自动化工具,以检测有问题的出版物.
- 需要进一步的研究来完善模型,并纳入额外的功能,以提高预测准确度.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Longitudinal Research
11.9K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
11.9K
Archival Research
16.0K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
16.0K
Survival Tree
73
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
73
Retrieval
101
Retrieval is the process of getting information out of memory storage and back into conscious awareness. This ability is essential for daily tasks like brushing hair and teeth, driving to work, and performing job duties. Retrieval occurs in three ways: recall, recognition, and relearning.
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
101


