Related Experiment Video
Updated: Jan 11, 2026

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
2.3K
Machine Learning-Based Identification of Natural History Studies in Rare Diseases: A Step toward Understanding
Kelly Chen1, Minghui Ao1, Sungrim Moon1
1National Center for Advancing Translational Sciences, Rockville, Maryland 20850, United States.
Journal of Rare Diseases (Berlin, Germany)
|November 17, 2025
Summary
This study introduces a machine learning method to automatically identify natural history studies (NHS) in PubMed. This approach accelerates rare disease research and drug development by efficiently collecting crucial disease progression data.
Area of Science:
- Medical Informatics
- Computational Biology
- Rare Disease Research
Background:
- Natural history studies (NHS) are vital for rare disease research, aiding in prevalence estimation and biomarker identification.
- Systematic identification of NHS is crucial for large-scale analysis supporting drug development.
- Current methods for identifying NHS are often manual and time-consuming.
Purpose of the Study:
- To develop and evaluate a machine learning-based approach for automated identification of NHS from PubMed.
- To establish a foundational step for large-scale NHS analysis in rare disease drug development.
- To compare the performance of binary versus multiclass classification for NHS identification.
Main Methods:
- A manually curated corpus of NHS was used to train and evaluate machine learning and deep learning models.
- Both binary (NHS-relevant vs. NHS-irrelevant) and multiclass (Unrelated, Irrelevant, Secondary, Primary) classification approaches were tested.
- The PubMedBERT-base-uncased-abstract model was utilized for classification.
Main Results:
- Binary classification models significantly outperformed multiclass models in identifying NHS.
- The PubMedBERT-base-uncased-abstract model achieved the highest performance in binary classification (Precision=0.8171, Recall=0.8079, F1=0.8125, AUCPR=0.8768).
- Deep learning models demonstrate the feasibility of automated NHS identification.
Conclusions:
- Automated NHS identification using deep learning is feasible and effective.
- Binary classification is a highly effective initial strategy for identifying NHS.
- This approach can accelerate data collection for NHS analysis, enhancing understanding of disease progression in rare and common diseases.
Related Concept Videos
Steps in Outbreak Investigation
476
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
476
Genome-wide Association Studies-GWAS
15.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.3K

