在遗传关联研究中,对线性混合模型的异常精确适合
Yongtao Guan1,2, Daniel Levy1,2
1Framingham Heart Study, 73 Mt. Wayte, Framingham, MA 01702, USA.
Genetics
|August 30, 2024
概括
新的方法,IDUL和IDUL†,有效地适应线性混合模型 (LMMs) 进行遗传关联研究. 这些代分散更新显著超过现有方法,特别是复杂的多组数据.
科学领域:
- 遗传学 遗传学 是一个
- 生物统计学 生物统计学
- 计算生物学 计算生物学
背景情况:
- 线性混合模型 (LMMs) 在遗传关联研究中至关重要,以控制人口结构和样本相关性,最大限度地减少假阳性.
- 目前的LMM研究往往侧重于近似计算,因为准确的方法是计算密集的,缺乏理论保证.
- 涉及数百万个遗传标记物和数千个表型的多组学研究,为LMM适配带来了重大的计算挑战.
研究的目的:
- 引入新的代方法,IDUL和IDUL†,以高效准确地适应线性混合模型 (LMM).
- 解决LMM在大型遗传关联研究中的计算需求,特别是针对多组数据.
- 为现有的LMM安装算法提供理论上合理且实际上高效的替代方案.
主要方法:
- 开发IDUL,一个代分散更新算法,用于LMM配件.
- 引入IDUL†,这是IDUL的修改版,确保更新期间单调的可能性增加.
- 在计算效率和准确性方面,IDUL和IDUL†性能与牛顿-拉普森方法进行比较.
主要成果:
- 与最先进的牛顿-拉普森方法相比,IDUL和IDUL†的效率明显更高.
- 在实践中,IDUL和IDUL†都取得了相同的结果.
- 这些方法在分析额外的表型时表现出极高的效率,使其适合于多组学研究.
结论:
- IDUL和IDUL†提供了一种计算效率高,准确的方法,用于将LMM与遗传关联研究相匹配.
- 由于LMM概率的单模性质,IDUL†的理论性质确保了对比准确性.
- 开发的软件包为研究人员研究复杂的多组特征的遗传基础提供了宝贵的工具.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
Hardy-Weinberg Principle
72.0K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
72.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
45
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
45
Genome-wide Association Studies-GWAS
13.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.2K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Goodness-of-Fit Test
3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K


