Related Experiment Video
Updated: Jul 11, 2026

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.4K
Identification and prioritization of disease candidate genes using biomedical named entity recognition and random
Sushrutha Raj1, Vindhya Namdeo2, Payal Singh2
1Amity Institute of Integrative Sciences and Health, Amity University Haryana, Amity Education Valley, Gurgaon, 122413, India.
Computers in Biology and Medicine
|May 11, 2025
Summary
This study introduces a novel two-tier machine learning protocol to identify disease-associated genes from abstracts. The method accurately classifies gene associations, discovering new candidate genes for complex diseases like Alzheimer's, breast cancer, and type 2 diabetes.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Identifying candidate genes is crucial for understanding complex diseases, enabling early diagnosis, personalized treatment, and drug discovery.
- Traditional methods for disease gene identification face limitations, requiring large sample sizes and high statistical power, which is challenging for complex diseases.
- Existing disease gene prioritization methods often rely on computational predictions lacking experimental validation.
Purpose of the Study:
- To introduce an innovative two-tier machine learning protocol for distilling disease-gene association details from disease-specific abstracts.
- To classify disease-gene associations into Positive, Negative, and Ambiguous categories using advanced text mining.
- To address the limitations of existing tools in disease gene identification and prioritization.
Main Methods:
- Utilized a two-tier machine learning protocol employing Random Forest for text classification.
- Level 1 classifier distinguished the presence or absence of disease-gene associations.
- Level 2 classifier further categorized associations into Positive, Negative, and Ambiguous classes, trained and validated on Alzheimer's, Breast Cancer, and Type 2 Diabetes datasets.
Main Results:
- Achieved high accuracy: 97.29% for level 1 and 98.14% for level 2 classification.
- Successfully extracted 2769, 3220, and 740 positively associated genes for Alzheimer's, Breast Cancer, and Type 2 Diabetes, respectively.
- Identified a significant number of novel genes (1008 for Alzheimer's, 670 for Breast Cancer, 165 for Type 2 Diabetes) not present in existing databases, expanding genetic exploration.
Conclusions:
- The developed protocol offers a comprehensive roadmap for exploring genetic diseases.
- The protocol demonstrated versatility and robustness across different disease contexts (Alzheimer's, Breast Cancer, Type 2 Diabetes).
- The findings provide promising candidates for targeted interventions and highlight ambiguous associations for further investigation, with broad applicability in biomedical research.

