PHA4GE质量控制上下文数据标签:用于共享具有已知的质量问题的公共卫生序列数据集的标准化注释,以促进测试和培训
Emma J Griffiths1, Inês Mendes2, Finlay Maguire3
1Centre for Infectious Disease Genomics and One Health, Faculty of Health Sciences, Simon Fraser University, Burnaby, British Columbia, Canada.
Microbial genomics
|June 11, 2024
概括
公共卫生实验室现在有标准化的标签来标记低质量的病原体基因组数据,提高其可发现性和重用性. 这些上下文数据标签增强了基因组流行病学和可重现性的数据共享.
科学领域:
- 基因组流行病学基因组流行病学
- 生物信息学是一种生物信息学.
- 公共卫生监督是对公共卫生的监督.
背景情况:
- 公共卫生实验室正在扩大对病原体监测的基因组测序和生物信息学能力.
- 对于基因组数据分析来说,湿干实验室程序的强有力的验证,培训和优化至关重要.
- 分享低质量或有目的生成的数据集对于可重复性至关重要,但缺乏标准化的机制.
研究的目的:
- 为应对共享非最佳病原体序列数据的挑战.
- 开发标准化的上下文数据标签,用于标记和提高低质量数据集的可发现性.
- 提高基因组数据在公共存储库中的实用性,可搜索性,可访问性和重用性.
主要方法:
- 基因组流行病学公共卫生联盟 (PHA4GE) 通过社区商开发了标准化的上下文数据标签,包括国际核酸序列数据协作 (INSDC).
- 标签使用本体学进行标准化,确保有机体和测序技术不可知论.
- 开发的标签由FDA的GenomeTrakr实验室网络测试和实施,用于SARS-CoV-2废水监测.
主要成果:
- PHA4GE已经成功开发和标准化了一套上下文数据标签,用于标记已知质量问题的病原体序列数据.
- 这些标签是无生物体的,适用于各种测序技术和合成数据,并由PHA4GE维护.
- 美国食品和药物管理局的GenomeTrakr网络的实施证明了这些标签在例行监控提交中的实际实用性.
结论:
- 标准化的PHA4GE上下文数据标签改善了公共存储库中关于质量控制的沟通.
- 这些标签使质量可变的数据集更容易识别,促进更好的数据共享和重用.
- 标签旨在随着社区需求的发展而发展,提供反和建议的机制.
相关概念视频
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Sanger Sequencing
754.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.0K
Next-generation Sequencing
88.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.6K
Maxam-Gilbert Sequencing
11.2K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.2K
Single Nucleotide Polymorphisms-SNPs
15.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.0K


