基于机器学习的自然历史研究在罕见疾病的识别:迈向了解疾病发展和结果的一步
Kelly Chen1, Minghui Ao1, Sungrim Moon1
1National Center for Advancing Translational Sciences, Rockville, Maryland 20850, United States.
Journal of rare diseases (Berlin, Germany)
|November 17, 2025
概括
这项研究引入了一种机器学习方法,可以在PubMed.中自动识别自然历史研究 (NHS). 这种方法通过有效收集关键疾病进展数据,加速罕见疾病研究和药物开发.
科学领域:
- 医疗信息学 医疗信息学
- 计算生物学 计算生物学
- 罕见疾病研究 罕见疾病研究
背景情况:
- 自然史研究 (NHS) 对罕见疾病研究至关重要,有助于流行率估计和生物标志物识别.
- 系统地识别NHS对于支持药物开发的大规模分析至关重要.
- 目前用于识别NHS的方法通常是手动的,耗时的.
研究的目的:
- 开发和评估基于机器学习的方法,用于从PubMed.com自动识别NHS.
- 为在罕见病药物开发中进行大规模NHS分析奠定基础.
- 为了比较NHS识别的二进制与多类分类的性能.
主要方法:
- 使用NHS手工策划的数据库来训练和评估机器学习和深度学习模型.
- 测试了二进制 (NHS相关与NHS无关) 和多类 (无关,无关,二级,初级) 分类方法.
- 使用PubMedBERT-base-uncased-abstract模型进行分类.
主要成果:
- 二元分类模型在识别NHS时显著优于多类模型.
- 在二进制分类中,PubMedBERT-base-uncased-abstract模型获得了最高的性能 (精度=0.8171,回忆=0.8079,F1=0.8125,AUCPR=0.8768).
- 深度学习模型展示了自动化NHS识别的可行性.
结论:
- 使用深度学习的自动化NHS识别是可行的和有效的.
- 二元分类是一种非常有效的初始策略,用于识别NHS.
- 这种方法可以加速用于NHS分析的数据收集,提高对罕见和常见疾病疾病进展的理解.
相关概念视频
Steps in Outbreak Investigation
476
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
476
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K


