Related Experiment Videos
Perspectives: sequence data base searching in the era of large-scale genomic sequencing
1Department of Molecular and Human Genetics, Baylor College of Medicine, Houston, Texas 77030 USA. Randall_F_Smith@sbphrd.com
Genome Research
|August 1, 1996
Summary
Genomic sequencing advances promise to reveal biochemical functions but pose challenges. New tools and curated databases are crucial for managing large-scale sequence data effectively.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Large-scale genome sequencing projects are generating vast amounts of biological data.
- Existing sequence database search methods face significant challenges with increasing data volume.
Purpose of the Study:
- To identify the adverse effects of expanding sequence databases on data searching.
- To propose solutions for managing and analyzing large-scale genomic data.
Main Methods:
- Analysis of current sequence database search limitations.
- Identification of key challenges including search time, false positives, and annotation accuracy.
Main Results:
- Increased database size leads to longer search times and more irrelevant matches.
- Coding region prediction accuracy is compromised, impacting protein database searches.
- Limited initial annotation hinders the determination of biological relevance for database hits.
Conclusions:
- Improved database annotation tools are essential for effective data analysis.
- Smaller, representative, and highly-annotated databases are needed for initial analyses.
- Proactive strategies are required to manage the influx of genomic sequence data.