在1000个基因组的基于回归的多SNP分析中使用本地主要组件进行维度缩小,以及加拿大长度老龄化研究 (CLSA)
Fatemeh Yavartanoo1, Myriam Brossard2, Shelley B Bull2,3
1Department of Mathematics Education, Seoul National University, Seoul, South Korea.
Genetic epidemiology
|March 1, 2025
概括
使用本地主要组件 (DRLPC) 的维度缩小有效地解决了遗传关联研究中的多线性. 这种方法提高了回归模型的稳定性,并提高了遗传测试的统计能力.
科学领域:
- 遗传学 是一个遗传学.
- 统计遗传学 统计遗传学
- 生物信息学是一种生物信息学.
背景情况:
- 在使用多个单核酸多态 (SNP) 的遗传关联研究中,多线性构成了重大挑战.
- 这一问题可能导致回归模型的不稳定性和遗传关联分析的失败.
- 现有的方法很难在密集的基因型数据中充分解决严重的多线性.
研究的目的:
- 提出和评估一种新的维度缩小方法,即使用本地主要组件 (DRLPC) 的维度缩小,以解决多线性.
- 提高基因关联研究中基于回归的统计测试的功率和适用性.
- 评估DRLPC在减少变量数量的有效性,同时保持重要的遗传信息.
主要方法:
- DRLPC删除了具有高线性依赖性的SNP,假设剩余的SNP捕捉了它们的影响.
- 差异通胀因子 (VIF) 用于量化对线性,不包括高于VIF值 (例如20) 的SNP.
- 该方法应用于来自1000个基因组项目和加拿大长度衰老研究 (CLSA) 的染色体22 SNPs.
主要成果:
- DRLPC显著降低了回归分析中SNP的数量,特别是对于较大的基因 (平均降低至20%).
- 对于较小的基因,减少程度较低 (平均约为48%),表明了量身定制的有效性.
- 在模拟研究中,DRLPC的应用使多重回归沃尔德测试的功率从60%提高到大约80%.
结论:
- DRLPC是缓解遗传关联研究多线性的一种有效策略.
- 该方法增强了统计测试的力量,导致更强大的遗传关联发现.
- DRLPC为后续的回归分析提供了更好的适用性,特别是使用大型遗传数据集.
关键词:
1000个基因组项目 (第三阶段)加拿大长度老龄化研究 (CLSA)缩小尺寸缩小尺寸的方法多个SNP的统计数据.多重对线性多重对线性主要组成部分分析 (PCA)差异通胀因子 (VIF) 的变化更多相关视频
14:27Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
15.6K
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
12.9K
相关概念视频
Genome-wide Association Studies-GWAS
12.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.3K
Comparing Copy Number Variations and SNPs
17.0K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.0K
Pleiotropy
39.4K
Pleiotropy is the phenomenon in which a single gene impacts multiple, seemingly unrelated phenotypic traits. For example, defects in the SOX10 gene cause Waardenburg Syndrome Type 4, or WS4, which can cause defects in pigmentation, hearing impairments, and an absence of intestinal contractions necessary for elimination. This diversity of phenotypes results from the expression pattern of SOX10 in early embryonic and fetal development. SOX10 is found in neural crest cells that form melanocytes,...
39.4K
Single Nucleotide Polymorphisms-SNPs
13.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
13.8K
