Breast cancer prediction using genome wide single nucleotide polymorphism data.
BMC Bioinformatics
|November 26, 2013
Summary
This study developed a breast cancer prediction model using SNP profiles, achieving 59.55% accuracy in internal validation and 60.25% in external validation, outperforming baseline models.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- Breast cancer prediction remains a challenge.
- Genome-wide association studies (GWAS) offer potential for predictive modeling.
- Single Nucleotide Polymorphism (SNP) profiles are key genetic markers.
Purpose of the Study:
- To develop and validate a predictive model for breast cancer risk using SNP data.
- To evaluate the efficacy of feature selection and machine learning methods for breast cancer prediction.
- To compare the developed model against existing methods and datasets.
Main Methods:
- Genotyping of 696 female subjects (348 cases, 348 controls) using Affymetrix Human SNP 6.0 arrays.
- Application of EIGENSTRAT for population stratification correction and stringent SNP filtering.
- Utilized MeanDiff feature selection combined with K-Nearest Neighbors (KNN) learning algorithm.
Main Results:
- Achieved 59.55% Leave-One-Out Cross-Validation (LOOCV) accuracy with the developed model, significantly better than the 51.52% baseline.
- External validation on the CGEMS dataset yielded 60.25% LOOCV accuracy, surpassing its 50.06% baseline.
- The MeanDiff and KNN combination outperformed other tested feature selection and learning method combinations.
Conclusions:
- The developed SNP-based predictive model shows promising accuracy for breast cancer risk assessment.
- Future improvements may involve larger cohorts, detailed phenotyping, and integration of diverse genomic and non-genetic data.
- Further research is needed to enhance prediction accuracy by accounting for breast cancer heterogeneity and other risk factors.
Related Concept Videos
Genome-wide Association Studies-GWAS
12.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.7K
Single Nucleotide Polymorphisms-SNPs
14.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.8K
Comparing Copy Number Variations and SNPs
11.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.6K


