CNV-Profile回归:在全基因组测序数据中的副本数变异关联分析的新方法
bioRxiv : the preprint server for biology
|December 9, 2024
概括
这项研究引入了一个新的框架来分析副本数变异 (CNV) 和它们对疾病风险的影响. 该方法有效地识别因果CNV,为阿尔茨海默病 (AD) 风险评估提供了更高的准确性.
科学领域:
- 遗传学 遗传学 是一个
- 基因组分析 基因组分析
- 疾病风险评估疾病风险评估
背景情况:
- 副本数变异 (CNV) 是影响疾病风险的DNA变异.
- 由于位置不明,剂量/长度依赖,以及需要对区域CNV进行联合分析,评估CNV影响具有挑战性.
研究的目的:
- 开发一个新的框架,用于对复制数变体 (CNVs) 的关联分析.
- 在基因组区域内模拟个体的整个CNV资料,以捕捉长度和剂量变化.
- 为了使CNV效应的共同估计,并绕过预定义的CNV loci的需要.
主要方法:
- 一个新的框架,使用CNV概况曲线来表示个人的CNV.
- 对于选择与特征相关的CNVs而言,拉索惩罚.
- 权重的L2融合处罚,以鼓励相邻的CNV产生类似的效果.
主要成果:
- 模拟显示了因果CNV的改进识别,控制了假阳性率和精确的效果大小估计.
- 应用到阿尔茨海默氏病测序项目数据确定了与阿尔茨海默氏病 (AD) 相关的额外的CNV.
- 已识别的CNV与已知的AD风险基因重叠,并且在与神经元相关的生物过程中得到丰富.
结论:
- 拟议的框架提供了一种可靠的CNV关联分析方法,改进疾病风险评估.
- 这种方法增强了对导致阿尔茨海默氏症等复杂疾病的遗传变异的识别.
- 这些发现为阿尔茨海默病的遗传结构提供了新的见解.
相关概念视频
Comparing Copy Number Variations and SNPs
17.2K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.2K
Single Nucleotide Polymorphisms-SNPs
14.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
14.0K
Next-generation Sequencing
87.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.4K
Sanger Sequencing
752.8K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
752.8K
Genomics
35.9K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.9K
Genome Copying Errors
4.1K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.1K


