Related Experiment Videos
High-throughput identification, database storage and analysis of SNPs in EST sequences
1Delaware Biotechnology Institute, University of Delaware, 15 Innovation Way, Newark, DE 19711, USA. useche@capsl.udel.edu
Genome Informatics. International Conference on Genome Informatics
|January 16, 2002
Summary
This study introduces an in-silico pipeline for discovering single nucleotide polymorphisms (SNPs) in maize EST data. The developed software efficiently identifies SNPs and insertion/deletions, aiding genetic research.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Single nucleotide polymorphisms (SNPs) are frequent DNA variations crucial for genetic markers.
- Experimental SNP discovery is complex and costly.
- In-silico SNP discovery offers a cost-effective alternative using existing large datasets.
Purpose of the Study:
- To design and implement an in-silico SNP detection software pipeline.
- To address challenges in large-scale in-silico SNP discovery.
- To facilitate preliminary analysis and basic statistics of SNP data.
Main Methods:
- Developed an integrated pipeline for data processing from sequence collection to SNP database.
- Optimized PolyBayes parameters for SNP detection in maize expressed sequence tag (EST) data.
- Integrated PHRAP and CAT assemblers, and a Bayesian engine (PolyBayes) for SNP detection.
Main Results:
- Detected 2439 SNPs and 822 insertion/deletions (INDELs) in 68,000 maize ESTs with high confidence (PolyBayes probability > 0.99).
- Implemented a user interface for preliminary data analysis and statistics.
- Ensured smooth data transition between pipeline components using data interfaces.
Conclusions:
- The developed in-silico pipeline effectively addresses challenges in large-scale SNP discovery.
- The pipeline facilitates efficient identification and preliminary analysis of SNPs and INDELs.
- This approach aids in gaining insights into polymorphism distribution and significance prior to experimental validation.