Related Experiment Video
Updated: Feb 7, 2026

07:45
Quasi-light Storage for Optical Data Packets
Published on: February 6, 2014
11.3K
Broom: application for non-redundant storage of high throughput sequencing data
Levent Albayrak1,2, Kamil Khanipov1,2,3, George Golovko1,2
1Department of Pharmacology and Toxicology, University of Texas Medical Branch - Galveston, Galveston, TX, USA.
Bioinformatics (Oxford, England)
|July 17, 2018
Summary
High-throughput sequencing (HTS) generates massive data, posing storage challenges for researchers. Broom efficiently compresses and stores only high-quality sequencing reads, offering a practical solution for data management.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- High-throughput sequencing (HTS) technology has rapidly advanced, increasing data output while decreasing costs.
- The exponential growth in HTS data presents significant storage and transfer challenges for smaller research entities.
- Current HTS data storage strategies need reevaluation in light of improved sequencing quality and coverage.
Purpose of the Study:
- To develop a novel application for efficient storage of high-throughput sequencing data.
- To address the data management challenges faced by small labs and individual researchers using HTS.
- To leverage advancements in sequencing quality for optimized data storage solutions.
Main Methods:
- Developed Broom, a stand-alone C++ application.
- Implemented algorithms for selecting and storing only high-quality sequencing reads.
- Enabled high compression rates for reduced data footprint.
- Supported single and paired-end reads in FASTQ and FASTA formats.
Main Results:
- Broom achieves extremely high compression rates for HTS data.
- The application effectively selects and stores only high-quality sequencing reads.
- Data can be decompressed into FASTA format.
- The software is available as a stand-alone application.
Conclusions:
- Broom offers a viable solution for managing large HTS datasets.
- The application significantly reduces storage requirements for sequencing data.
- This tool can facilitate HTS data accessibility and analysis for researchers with limited resources.
Related Concept Videos
Storage
416
A schema is a mental framework that helps individuals organize and interpret information. Schemata, formed from previous experiences, influence how we process new information: how we encode it, the inferences we make, and how we retrieve it. For instance, a schema for what a typical classroom looks like might include desks, a teacher's desk, a whiteboard, and students in such an environment. This expectation helps us quickly understand and navigate new classrooms without needing to analyze...
416
ATP Energy Storage and Release
14.5K
ATP is a highly unstable molecule. Unless quickly used to perform work, ATP spontaneously dissociates into ADP and inorganic phosphate (Pi), and the free energy released during this process is lost as heat. The energy released by ATP hydrolysis is used to perform work inside the cell and depends on a strategy called energy coupling. Cells couple the exergonic reaction of ATP hydrolysis with endergonic reactions, allowing them to proceed.
One example of energy coupling using ATP involves a...
One example of energy coupling using ATP involves a...
14.5K
Sugars as Energy Storage Molecules
9.9K
Sugar (a simple carbohydrate) metabolism (chemical reactions) is a classic example of the many cellular processes that use and produce energy. Living things consume sugar as a major energy source because sugar molecules have considerable energy stored within their bonds. Consumed carbohydrates have their origins in photosynthesizing organisms like plants. During photosynthesis, plants use the energy of sunlight to convert carbon dioxide gas into sugar molecules, like glucose. Because this...
9.9K
Fats as Energy Storage Molecules
27.1K
Triglycerides are a form of long-term energy storage molecules. They are made of glycerol and three fatty acids. To obtain energy from fat, triglycerides must first be broken down by hydrolysis into their two principal components, fatty acids and glycerol. This process, called lipolysis, takes place in the cytoplasm. The resulting fatty acids are oxidized by β-oxidation into acetyl-CoA, which is used by the Krebs cycle. The glycerol that is released from triglycerides after lipolysis...
27.1K
Cis-regulatory Sequences
11.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.9K
How Data are Classified: Categorical Data
44.8K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.8K

