Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

RNA-seq03:21

RNA-seq

10.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases. 
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.7K
Genomics02:02

Genomics

38.2K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
38.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multi-Task Deep Learning for Surface Metrology.

Sensors (Basel, Switzerland)·2025
Same author

Effect of Carbon-Based Modifications of Polydicyclopentadiene Resin on Tribological and Mechanical Properties.

Materials (Basel, Switzerland)·2025
Same author

The Influence of Resin Volume Fraction on Selected Properties of Polymer Concrete.

Materials (Basel, Switzerland)·2025
Same author

Utilization of lightweight ceramic aggregates based on waste materials in the production of lightweight polymer concrete as a component of sustainable architecture.

Scientific reports·2024
Same author

Tribological Properties of Composites Based on Single-Component Powdered Epoxy Matrix Filled with Graphite.

Materials (Basel, Switzerland)·2024
Same author

The Influence of Graphite Filler on the Self-Lubricating Properties of Epoxy Composites.

Materials (Basel, Switzerland)·2024

Related Experiment Video

Updated: Oct 26, 2025

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project
10:19

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project

Published on: April 8, 2017

17.6K

A Large-Scale and Serverless Computational Approach for Improving Quality of NGS Data Supporting Big Multi-Omics Data

Dariusz Mrozek1, Krzysztof Stępień1, Piotr Grzesik1

  • 1Department of Applied Informatics, Silesian University of Technology, Gliwice, Poland.

Frontiers in Genetics
|July 30, 2021
PubMed
Summary

This study introduces a scalable cloud-based Data Lake and a library for cleaning next-generation sequencing (NGS) data. The solution efficiently processes large multi-omics datasets for personalized medicine applications.

Keywords:
OMICS databig datacloud computingdata lakedata qualitynext-generation sequencingqueryingserverless

More Related Videos

Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies
13:24

Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies

Published on: April 11, 2016

12.0K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.4K

Related Experiment Videos

Last Updated: Oct 26, 2025

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project
10:19

Transcriptomic Analysis of C. elegans RNA Sequencing Data Through the Tuxedo Suite on the Galaxy Project

Published on: April 8, 2017

17.6K
Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies
13:24

Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies

Published on: April 11, 2016

12.0K
Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
10:41

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms

Published on: May 9, 2017

9.4K

Area of Science:

  • Bioinformatics
  • Computational Biology
  • Genomics

Background:

  • Next-generation sequencing (NGS) generates vast amounts of DNA/RNA data crucial for multi-omics analyses and personalized medicine.
  • Existing Big Data tools lack simple, declarative methods for large-scale NGS data quality improvement.
  • Efficient storage and processing infrastructure are essential for handling large sequencing projects and molecular profiling.

Purpose of the Study:

  • To adapt the Data Lake concept for storing and processing big NGS data.
  • To develop a dedicated library for cleaning DNA/RNA sequences from single-read and paired-end techniques.
  • To create a scalable, cloud-based solution for efficient NGS data management and analysis.

Main Methods:

  • Implementation of a Data Lake architecture for big NGS data storage and processing.
  • Development of a specialized library with U-SQL extensions for DNA/RNA sequence cleaning.
  • Leveraging cloud scalability for flexible adjustment to data processing volumes.
  • Utilizing declarative U-SQL for simplified data extraction, processing, and storage workflows.

Main Results:

  • The proposed solution demonstrates scalability on the Cloud, adapting to varying data volumes.
  • The library effectively cleans DNA/RNA sequences, improving data quality for downstream analyses.
  • The integrated system supports ample storage and highly parallel, scalable processing for NGS multi-omics data.
  • Experiments confirm the solution's capability to meet the demands of large-scale sequencing data analysis.

Conclusions:

  • The Data Lake and dedicated library provide a scalable and efficient infrastructure for big NGS data.
  • This approach simplifies NGS data cleaning and processing, supporting multi-omics research and personalized treatment.
  • The solution addresses the critical need for robust data management in the era of large-scale sequencing.