很容易找到短读变异的区域,从 pangenome 数据中调用
Heng Li1,2,3
1Department of Biomedical Informatics, Harvard Medical School, 10 Shattuck St, Boston, MA 02215, USA.
ArXiv
|August 13, 2025
概括
新的样本不可知容易区域改善了人类基因组的短读变异调用精度. 该资源增强了用于临床和研究应用的变异过,克服了以前方法的局限性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 人口遗传学 人口遗传学
背景情况:
- 短读变异调用基准通常使用预定义的自信区域,可能低估非参考人体样本中的错误率.
- 现有的"容易区域"集可能无法考虑样本特定的变化,或者受到特定的对齐器和短读数据的偏差.
- 不自信地区的高错误率可能会阻碍在临床和研究环境中准确识别变异.
研究的目的:
- 开发一套全面的样本不可知易区集,用于准确的短读变异,跨越多样化的人类基因组进行调用.
- 在临床和研究环境中提供一个强大的资源来过虚假变异调用.
- 建立一种方法,为其他物种或组件生成类似区域.
主要方法:
- 利用数百个高质量的人类基因组组件来识别和定义样本不可知的容易区域.
- 评估了这些衍生区域内的变种呼叫的性能.
- 评估了关键基因组特征的覆盖范围,包括编码区域和致病变异.
主要成果:
- 开发了样本不可知的简单区域,使高精度的短读变量调用成为可能.
- 这些区域覆盖了人类基因组的很大一部分 (88.2%的GRCh38),包括92.2%的编码区域和96.3%的ClinVar病原变异.
- 确定的地区在基因组覆盖率和变异调用容易度之间提供了有利的平衡.
结论:
- 衍生出来的容易区域提供了一种强大而方便的方法来过人类样本中不准确的变异调用.
- 该资源适用于临床诊断和基础研究,提高了变异数据的可靠性.
- 该方法可用于为其他物种或人类组合生成类似的资源.
相关概念视频
Next-generation Sequencing
92.6K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
92.6K
RNA-seq
10.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.4K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K


