Related Experiment Videos
An adaptive threshold determination method of feature screening for genomic selection.
Guifang Fu1, Gang Wang2, Xiaotian Dai2
1Department of Mathematics and Statistics, Utah State University, Logan, 84322, UT, USA. guifang.fu@usu.edu.
BMC Bioinformatics
|April 14, 2017
Summary
Selecting influential single nucleotide polymorphisms (SNPs) is crucial for genomic prediction. Our novel Backward Elimination Iterative Distance Correlation (BE-IDC) method efficiently identifies key SNPs, improving accuracy and speed in genomic selection.
Area of Science:
- Genomics
- Statistical Genetics
- Bioinformatics
Background:
- Complex traits are influenced by a small subset of single nucleotide polymorphisms (SNPs) within a vast genome.
- Accurate and efficient selection of these influential SNPs is a significant challenge in genomic studies.
- Traditional feature screening methods often struggle with unclear thresholds for SNP selection.
Purpose of the Study:
- To develop a novel procedure for selecting a parsimonious set of influential SNPs for complex trait prediction.
- To address the limitations of existing feature screening approaches, particularly the issue of unclear thresholds.
- To enhance both prediction accuracy and computational efficiency in genomic selection.
Main Methods:
- Proposed a Backward Elimination Iterative Distance Correlation (BE-IDC) procedure.
- BE-IDC utilizes an adaptive threshold for SNP selection.
- Evaluated the method through six simulations and application to Arabidopsis thaliana genome-wide data.
Main Results:
- The adaptive threshold estimated by BE-IDC consistently outperformed fixed threshold methods in simulations.
- BE-IDC successfully identified four influential SNPs from over 216,000 SNPs in the Arabidopsis thaliana dataset.
- The identified SNPs included the FRIGIDA gene, consistent with findings from traditional methods.
Conclusions:
- The BE-IDC procedure offers a robust approach for selecting SNPs that are crucial for predicting complex traits.
- BE-IDC balances high prediction accuracy with computational efficiency, meeting key demands in genomic selection.
- This method provides a valuable tool for identifying important genetic markers in large-scale genomic datasets.