Experimental design-based functional mining and characterization of high-throughput sequencing data in the sequence
Takeru Nakazato1, Tazro Ohta, Hidemasa Bono
1Database Center for Life Science (DBCLS), Research Organization of Information and Systems (ROIS), Tokyo, Japan.
Plos One
|October 30, 2013
Summary
High-throughput sequencing data in the Sequence Read Archive (SRA) is hard to search. This study created lists of cited SRA entries and associated diseases to improve access to quality sequencing data.
Area of Science:
- Genomics and Bioinformatics
- Biotechnology
- Data Science
Background:
- High-throughput sequencing, or next-generation sequencing (NGS), generates vast amounts of genomic, transcriptomic, and epigenetics data.
- The Sequence Read Archive (SRA) is a public repository for this sequencing data, experiencing rapid growth.
- Challenges exist in searching and accessing high-quality data from SRA due to complex structure and inconsistent data quality.
Purpose of the Study:
- To improve the accessibility and quality assessment of data within the Sequence Read Archive (SRA).
- To establish curated lists of SRA entries linked to publications and disease information.
- To facilitate research utilizing -omics data for disease mechanism elucidation.
Main Methods:
- Focused on SRA entries cited in journal articles as a proxy for data quality.
- Extracted SRA IDs and PubMed IDs (PMIDs) to create SRA-PMID pairs and a publication list.
- Characterized SRA entries by disease keywords (MeSH terms) from associated articles, creating SRA-MeSH disease term pairs and a disease list.
Main Results:
- Successfully retrieved 2748 SRA ID-PMID pairs and 989 SRA ID-MeSH disease term pairs.
- Constructed a publication list and a disease list referencing SRA data.
- Developed hyperlinks between SRA diseases and existing disease feature profiles in the Gendoo system.
Conclusions:
- The developed DBCLS SRA web service enhances the discoverability of high-quality, cited sequencing data.
- This resource aids researchers in navigating SRA for -omics analyses, particularly for understanding disease mechanisms.
- Improved data accessibility from SRA supports advancements in genomic research and personalized medicine.


