定义一个并联重复目录和变异集群,用于全基因组分析
Ben Weisburd1,2, Egor Dolzhenko3, Mark F Bennett4,5,6
1Program in Medical and Population Genetics, Broad Center for Mendelian Genomics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.
bioRxiv : the preprint server for biology
|November 24, 2025
概括
这项研究介绍了TRExplorer目录v1.0,这是一个用于并联重复 (TR) 基因型定型的新资源. 它解决了现有的TR目录中的不一致性,并为短读和长读分析提供了全面的注释数据集.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 人口遗传学 人口遗传学
背景情况:
- 串联重复 (TR) 目录对于重复基因型定型,定义基因组位置和动图至关重要.
- 现有的TR目录在位置边界上经常存在分歧,这阻碍了交叉研究的比较和数据的重用.
- 开发 TR 变异的公共数据库可能会导致由于位置定义的分歧而导致碎片化.
研究的目的:
- 为了比较现有的并联重复目录,并确定一个全面的全基因组目录的理想特征.
- 为大规模分析和人口数据库提供一个新的,富有注释的TR目录 (TRExplorer v1.0).
- 引入一种新的算法来定义使用长读序列数据的变异集群.
主要方法:
- 对现有的双重重复目录进行比较分析.
- TRExplorer目录v1.0的开发和注释,包括STR和VNTR.
- 应用一种新的算法,利用长时间读取的HiFi测序数据来识别和定义变异集群.
主要成果:
- TRExplorer目录v1.0包含490万个TR位点,包括以前目录中缺少的新型多态位点.
- TRs被分为分层,用于复制数分析的分离重复和用于序列级分析的变异集群.
- 人类基因组中至少发现了25,000个复杂变异集群,通常在多态区域内包含多个TRs.
结论:
- TRExplorer目录v1.0提供了一个标准化,全面的资源,用于跨多种测序技术的TR基因类型.
- 新的变异集群定义方法使复杂的多态区域能够更准确地分析.
- trexplorer.broadinstitute.org门户便于访问和使用目录和变异集群数据.
相关概念视频
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Single Nucleotide Polymorphisms-SNPs
17.9K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.9K
RNA-seq
11.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.7K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Genome-wide Association Studies-GWAS
15.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
15.2K
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K


