Related Experiment Video
Updated: Jul 9, 2025

09:34
Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
33.8K
ENLIGHTENMENT: A Scalable Annotated Database of Genomics and NGS-Based Nucleotide Level Profiles
IEEE/ACM Transactions on Computational Biology and Bioinformatics
|December 6, 2023
Summary
Researchers developed the Enlightenment database to integrate human genome data from multiple sources. This platform aids in identifying cancer biomarkers and understanding drug responses for cancer research.
Area of Science:
- Genomics
- Bioinformatics
- Cancer Research
Background:
- Advances in sequencing technologies have led to a surge in available whole-genome data.
- Understanding the human genome's role in cancer requires aggregating and standardizing this large dataset.
Purpose of the Study:
- To create a unified, scalable platform for integrating and analyzing human genomic data.
- To facilitate the identification of cancer-specific and pharmacogenetic biomarkers and their association with drug responses.
Main Methods:
- Implementation of the Enlightenment database, integrating data from eight public databases and H. sapiens DNA sequencing profiles.
- Annotation of genomic data for cancer biomarkers, pharmacogenetic biomarkers, drug response variability, and novel copy number variants.
- Deployment of the database on HBase distributed over a Hadoop cluster for scalability and processing of Next-Generation Sequencing (NGS) data.
Main Results:
- A centralized platform providing access to integrated genomic data and annotated results.
- Identification of cancer-specific biomarkers, pharmacogenetic biomarkers, and novel copy number variants.
- A scalable infrastructure capable of handling large-scale genomic datasets and integrating other omics data.
Conclusions:
- The Enlightenment database provides a valuable resource for cancer research by integrating and analyzing vast amounts of genomic data.
- The scalable HBase and Hadoop-based architecture addresses the challenges of storing and processing NGS data.
- This platform has the potential to accelerate cancer research and guide towards its eradication by enabling further data integration.
Related Concept Videos
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Next-generation Sequencing
89.0K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
89.0K
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Sanger Sequencing
754.6K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.6K

