Related Experiment Video
Updated: Mar 15, 2026

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
2.9K
Data Mining and Pattern Recognition Models for Identifying Inherited Diseases: Challenges and Implications
Lahiru Iddamalgoda1, Partha S Das2, Achala Aponso1
1Department of Computing, Informatics Institute of Technology, University of Westminster Colombo, Sri Lanka.
Frontiers in Genetics
|August 26, 2016
Summary
Data mining and pattern recognition improve genetic studies for inherited diseases. Combining gene prioritization, protein interaction (PPI), and K-nearest neighbors methods can accurately identify causal genetic variants.
Area of Science:
- Genomics and Bioinformatics
- Computational Biology
- Biomedical Data Mining
Background:
- Genetic studies are crucial for understanding inherited diseases.
- Accurate prioritization of single nucleotide polymorphisms (SNPs) linked to diseases remains a significant challenge.
- Existing data mining models for biomedical applications require refinement for disease-causal variant identification.
Purpose of the Study:
- To review current data mining and pattern recognition models for identifying inherited diseases.
- To discuss the necessity of binary classification and scoring-based prioritization methods for causal variant determination.
- To propose an integrated approach for enhanced genetic factor categorization in disease causation.
Main Methods:
- Review of state-of-the-art data mining and pattern recognition techniques.
- Analysis of binary classification and scoring-based prioritization methods for SNPs.
- Exploration of gene prioritization and protein-interaction (PPI) network analysis.
- Application of the K-nearest neighbors algorithm.
Main Results:
- Current data mining models show promise but face challenges in precise SNP prioritization.
- Binary classification and scoring methods offer distinct advantages and disadvantages for variant identification.
- Gene prioritization and PPI methods, when combined with K-nearest neighbors, demonstrate potential for accurate genetic categorization.
Conclusions:
- An integrated approach combining gene prioritization, PPI networks, and K-nearest neighbors offers a robust strategy for identifying genetic factors in inherited diseases.
- This integrated method can improve the accuracy of categorizing genetic variants contributing to disease causation.
- Further research into these combined methodologies is warranted for advancing precision medicine.
Related Concept Videos
Pedigree Analysis
90.4K
Overview
90.4K
Genome-wide Association Studies-GWAS
16.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.5K
EPS and iPS Cells in Disease Research
3.5K
Embryonic and induced pluripotent stem cells are excellent models for disease research because of their ability to self-renew and differentiate into most cell types. Somatic cells from a patient are isolated and reprogrammed into induced pluripotent stem cells or iPSCs. These iPSCs are later differentiated into the desired cell type, which mirrors the diseased cell of the patient. In this way, disease models have been created for investigating diseases such as Down syndrome, type I diabetes,...
3.5K
Genomic Imprinting and Inheritance
38.5K
Diploid organisms inherit genetic material through chromosomes from both parents. Copies of the same gene are known as alleles. In most cases, both alleles are simultaneously expressed and allow various cellular processes to function optimally. If one of the alleles is missing or mutated, the expression of the other allele can compensate; however, this is not true for all genes.
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
38.5K

