Related Experiment Video
Updated: Jul 11, 2026

Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Identification and prioritization of disease candidate genes using biomedical named entity recognition and random
Sushrutha Raj1, Vindhya Namdeo2, Payal Singh2
1Amity Institute of Integrative Sciences and Health, Amity University Haryana, Amity Education Valley, Gurgaon, 122413, India.
Background And Objective:
The elucidation of candidate genes is fundamental to comprehending intricate diseases, vital for early diagnosis, personalized treatment, and drug discovery. Traditional Disease Gene Identification methods encounter limitations, necessitating substantial sample sizes and statistical power, particularly challenging for complex diseases. Conversely, Disease Gene Prioritization methods leverage biological knowledge but rely on computational predictions, often lacking experimental validation. Addressing existing tool challenges, this study introduces an innovative two-tier machine-learning protocol that distils Disease Gene Association details from disease-specific abstracts, incorporating diverse findings. Employing advanced text mining, the model classifies disease-gene associations from the abstracts into Positive, Negative, and Ambiguous classes.
Methods:
Leveraging Random Forest as a robust text classification tool, this study demonstrates its efficacy in navigating complexities within biomedical texts. In the developed 2-tiered protocol, the level 1 classifier categorizes information into two classes, distinguished by the presence or absence of disease-gene associations, whereas the level 2 classifier further classifies into three classes: Positive, Negative, and Ambiguous associations. The developed classifier underwent rigorous training and cross-validation on different gold standard datasets - Alzheimer's, Breast Cancer and Type 2 Diabetes. Its performance across these varied disease contexts underscores its versatility and robustness without succumbing to overfitting.
Results:
Achieving an average accuracy of 97.29 % and 98.14 % for level 1 and level 2 classification, the protocol successfully extracted 2769, 3220 and 740 genes associated positively with Alzheimer's, Breast Cancer and Type 2 Diabetes. From the identified positive genes, a substantial number-1008, 670, and 165 genes, respectively-were not reported in established databases, thus expanding the genetic exploration of these diseases. These identified genes offer promising opportunities for targeted interventions, while ambiguous genes warrant further investigation to unravel deeper disease associations.
Conclusions:
This research significantly contributes to the understanding of genetic diseases by offering a comprehensive roadmap for their intricate exploration. Beyond the study's focus on Alzheimer's, Breast Cancer, and Type 2 Diabetes, the protocol's applicability extends to diverse biomedical landscapes, demonstrating its versatility and impactful potential for comprehensive disease exploration.

