Related Experiment Video
Updated: Oct 26, 2025

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project
Published on: April 8, 2017
A Large-Scale and Serverless Computational Approach for Improving Quality of NGS Data Supporting Big Multi-Omics Data
Dariusz Mrozek1, Krzysztof Stępień1, Piotr Grzesik1
1Department of Applied Informatics, Silesian University of Technology, Gliwice, Poland.
This study introduces a scalable cloud-based Data Lake and a library for cleaning next-generation sequencing (NGS) data. The solution efficiently processes large multi-omics datasets for personalized medicine applications.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Next-generation sequencing (NGS) generates vast amounts of DNA/RNA data crucial for multi-omics analyses and personalized medicine.
- Existing Big Data tools lack simple, declarative methods for large-scale NGS data quality improvement.
- Efficient storage and processing infrastructure are essential for handling large sequencing projects and molecular profiling.
Purpose of the Study:
- To adapt the Data Lake concept for storing and processing big NGS data.
- To develop a dedicated library for cleaning DNA/RNA sequences from single-read and paired-end techniques.
- To create a scalable, cloud-based solution for efficient NGS data management and analysis.
Main Methods:
- Implementation of a Data Lake architecture for big NGS data storage and processing.
- Development of a specialized library with U-SQL extensions for DNA/RNA sequence cleaning.
- Leveraging cloud scalability for flexible adjustment to data processing volumes.
- Utilizing declarative U-SQL for simplified data extraction, processing, and storage workflows.
Main Results:
- The proposed solution demonstrates scalability on the Cloud, adapting to varying data volumes.
- The library effectively cleans DNA/RNA sequences, improving data quality for downstream analyses.
- The integrated system supports ample storage and highly parallel, scalable processing for NGS multi-omics data.
- Experiments confirm the solution's capability to meet the demands of large-scale sequencing data analysis.
Conclusions:
- The Data Lake and dedicated library provide a scalable and efficient infrastructure for big NGS data.
- This approach simplifies NGS data cleaning and processing, supporting multi-omics research and personalized treatment.
- The solution addresses the critical need for robust data management in the era of large-scale sequencing.
More Related Videos
13:24Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies
Published on: April 11, 2016
10:41Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Genomics