Related Experiment Video
Updated: Aug 2, 2026

09:14
Informatic Analysis of Sequence Data from Batch Yeast 2-Hybrid Screens
Published on: June 28, 2018
DDBJ in collaboration with mass-sequencing teams on annotation
1Center for Information Biology and DNA Data Bank of Japan, National Institute of Genetics, Research Organization of Information and Systems, Yata, Mishima, 411-8540, Japan. ytateno@genes.nig.ac.jp
Nucleic Acids Research
|December 21, 2004
Summary
DNA Data Bank of Japan (DDBJ) released over 1 million genetic entries, including chimpanzee chromosomes and silkworm genomes. They also hosted workshops for human and mouse gene annotation, making valuable data publicly available.
Area of Science:
- Genomics and Bioinformatics
- Molecular Biology
- Gene Expression Analysis
Background:
- The DNA Data Bank of Japan (DDBJ) plays a crucial role in archiving and disseminating biological sequence data.
- Accurate and comprehensive annotation of full-length cDNAs is essential for understanding gene function and regulation.
- Advancements in sequencing technologies necessitate efficient data management and release strategies.
Purpose of the Study:
- To report the data collection and release activities of DDBJ over the past year.
- To highlight the integration of new data types, such as Cap Analysis Gene Expression (CAGE) data, for genome annotation.
- To emphasize DDBJ's commitment to facilitating research through public data accessibility and annotation workshops.
Main Methods:
- Collection and release of over 1 million genetic sequence entries, totaling over 718 million bases.
- Hosting workshops for human and mouse full-length cDNA annotation.
- Collaboration with RIKEN to establish a new Mass Sequences for Genome Annotation (MGA) data category for CAGE data.
Main Results:
- Release of diverse genomic data, including the complete chimpanzee chromosome 22 and whole-genome shotgun sequences of silkworm.
- Public availability of annotated human and mouse full-length cDNA data.
- Establishment of the MGA category to accommodate CAGE data, enhancing genome annotation resources.
Conclusions:
- DDBJ significantly expanded its public genetic database with substantial new entries and diverse species data.
- Active participation in annotation workshops and data integration initiatives improves the quality and utility of genomic information.
- The newly incorporated CAGE data under the MGA category will provide valuable insights into gene expression control mechanisms.
Related Concept Videos
Complementary DNA
Overview
Complementary DNA
Overview
Sanger Sequencing
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.

