Related Experiment Video
Updated: Jan 23, 2026

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
Moving Just Enough Deep Sequencing Data to Get the Job Done
Nicholas Mills1, Ethan M Bensman2, William L Poehlman3
1Holcombe Department of Electrical and Computer Engineering, Clemson University, Clemson, SC, USA.
Processing only a subset of large DNA sequencing data can significantly reduce transfer times. This approach maintains sufficient biological signal detection for RNA-Seq analysis, optimizing data handling costs.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- High-throughput DNA sequencing generates massive datasets, leading to high storage and transfer costs.
- These costs can limit data processing capabilities, especially for institutions with limited resources.
- Processing only essential data subsets is a potential solution to mitigate these costs.
Purpose of the Study:
- To investigate the feasibility of processing partial DNA sequence datasets.
- To determine if data reduction strategies can maintain biological information integrity.
- To assess the impact of processing data subsets on RNA-Seq analysis and transfer times.
Main Methods:
- Utilized four high-throughput DNA sequence datasets from two species with varying sequencing depths.
- Employed an RNA-Seq workflow to evaluate the effect of partial dataset processing on transcript detection.
- Determined a data cutoff point based on transcript detection sufficiency.
- Physically transferred minimal partial datasets and compared transfer times with full datasets.
Main Results:
- Processing partial datasets in an RNA-Seq workflow demonstrated the ability to detect RNA transcripts.
- A reduction of approximately 25% in total transfer time was observed when transferring minimal partial datasets compared to full datasets.
- The study successfully identified a data subset that sufficiently detects the biological signal.
Conclusions:
- Transferring minimal partial DNA sequence datasets is a viable strategy to reduce data transfer times.
- This method can accelerate analysis pipelines for large sequencing datasets.
- Optimizing data transfer through subsetting can make genomic data processing more accessible.
Related Concept Videos
Cis-regulatory Sequences
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Magnetic Field due to Moving Charges
Consider a point charge moving with a constant velocity. Like the electric field, the magnetic field at any point is directly proportional to the magnitude of the charge and inversely proportional to the square of the distance between the source point and the field point. However, unlike the electric field, the magnetic field is always perpendicular to the plane containing the line...
Sanger Sequencing
Data Reporting and Recording

