Related Experiment Video
Updated: Apr 12, 2026

12:01
3' End Sequencing Library Preparation with A-seq2
Published on: October 10, 2017
11.2K
Investigation into the annotation of protocol sequencing steps in the sequence read archive
Jamie Alnasir1, Hugh P Shanahan1
1Department of Computer Science, Royal Holloway, University of London, Egham, TW20 0EX UK.
Gigascience
|May 12, 2015
Summary
High-throughput sequencing data preparation lacks detailed annotation in the Sequence Read Archive (SRA). This limits the ability to systematically study and quantify bias introduced during critical library preparation steps.
Area of Science:
- Genomics
- Bioinformatics
- Molecular Biology
Background:
- High-throughput sequencing data generation involves a complex workflow with multiple protocol steps.
- Key library preparation stages, including DNA fractionation, blunting, phosphorylation, adapter ligation, and library enrichment, require accurate quantification of associated bias.
- The extent of bias introduced by these specific protocol steps remains largely undetermined.
Purpose of the Study:
- To assess the level of annotation for critical sequencing protocol steps within the Sequence Read Archive (SRA) database.
- To determine if experimental metadata adequately describes key library preparation stages for downstream analysis.
Main Methods:
- Utilized SQL relational database queries on the SRAdb SQLite database.
- Searched for keywords associated with DNA fragmentation, adapter ligation, and library enrichment across SRA records.
- Analyzed metadata from studies submitted to the Sequence Read Archive.
Main Results:
- Only 7.10% (fragmentation), 5.84% (ligation), and 7.57% (enrichment) of SRA records contained keywords for at least one of these three protocol steps.
- A mere 4.06% of all SRA records included keywords for all three critical library preparation steps.
- This indicates a low level of detailed annotation for essential sequencing workflow components.
Conclusions:
- The current annotation level in the SRA database hinders systematic investigations into protocol-specific biases.
- Lack of detailed metadata prevents the quantification of bias in high-throughput sequencing data.
- Future meta-analyses and comparative studies using SRA data are susceptible to unquantified sources of bias.
Related Concept Videos
RNA-seq
12.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.6K
Genome Annotation and Assembly
22.2K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
22.2K

