科纳:韩国核酸档案作为核酸序列数据的新数据库
Gunhwan Ko1, Jae Ho Lee1, Young Mi Sim1
1Korea Bioinformation Center, Korea Research Institute of Bioscience & Biotechnology, Daejeon 34141, Republic of Korea.
Genomics, proteomics & bioinformatics
|June 11, 2024
概括
韩国核酸档案 (KoNA) 解决了管理大规模测序数据的挑战. 它提供了核酸序列的国家存储库,促进了全球基因组学研究.
科学领域:
- 基因组学和生物信息学
- 数据科学与管理数据科学与管理
背景情况:
- 高通量测序数据的指数增长在数据访问,传输,存储和共享方面带来了重大挑战.
- 越来越需要强大的基础设施来支持数据驱动的生物研究,特别是管理大型基因组数据集.
研究的目的:
- 介绍韩国核酸档案 (KoNA),这是一个核酸序列数据的国家存储库.
- 描述实施的基础设施和程序,以管理和提供对大规模基因组数据的访问.
- 提高研究人员提交,访问和分析核酸序列数据的经验.
主要方法:
- 作为韩国生物数据站 (K-BDS) 的一部分,成立了KoNA,用于存储政府资助的生物数据.
- 实施了符合国际标准的标准操作程序,包括数据和元数据的自动和手动质量控制.
- 利用高速数据传输系统 (GBox) 和集成的云计算服务 (Bio-Express) 实现无数据处理.
主要成果:
- 截至2022年7月,KoNA的韩国阅读档案拥有超过477TB的下一代原始测序数据.
- 科纳,GBox和Bio-Express的综合系统简化了数据提交,访问和分析.
- KoNA成功地解决了对国家序列存储库的需求,并支持全球基因组学研究.
结论:
- 在韩国,KoNA为核酸序列数据提供了一个重要的国家资源.
- 集成平台增强了数据管理和可访问性,为基因组学的进步做出了贡献.
- 科纳促进全球数据共享,并支持更广泛的科学界.
相关概念视频
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Nucleic Acids and Nucleotides
9.0K
Nucleic acids are the most important macromolecules for the continuity of life. They carry the cell's genetic blueprint and have instructions for its functioning. The two main types of nucleic acids are deoxyribonucleic acid (DNA) and ribonucleic acid (RNA).
Deoxyribonucleic Acid (DNA)
DNA is the genetic material in all living organisms, ranging from single-celled bacteria to multicellular mammals. It is in the nucleus of eukaryotes and the organelles such as chloroplasts and mitochondria....
Deoxyribonucleic Acid (DNA)
DNA is the genetic material in all living organisms, ranging from single-celled bacteria to multicellular mammals. It is in the nucleus of eukaryotes and the organelles such as chloroplasts and mitochondria....
9.0K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Next-generation Sequencing
88.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.6K
Nucleic Acid Structure
6.1K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
6.1K
Sanger Sequencing
754.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.0K


