FedGMMAT:联合通用线性混合模型关联测试
Wentao Li1, Han Chen1,2, Xiaoqian Jiang1
1McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, Houston, Texas, United States of America.
PLoS computational biology
|July 24, 2024
概括
联合基因关联测试通过安全共享洞察力,而不是敏感数据,使得强大的疾病研究成为可能. FedGMMAT在保护患者隐私和遵守HIPAA的同时实现了聚合分析准确性.
科学领域:
- 遗传学 遗传学 是一个
- 生物信息学是一种生物信息学.
- 统计遗传学 统计遗传学
背景情况:
- 了解遗传疾病决定因素需要大量数据集,但数据共享受到隐私问题 (PHI,HIPAA) 和机构障碍的阻碍.
- 不同质的样本大小需要复杂的统计方法,比如一般化的线性混合效应模型,以解决混因素.
研究的目的:
- 为遗传关联测试开发一种保护隐私的联合方法.
- 为了使高功率的协作研究,而不影响敏感的遗传数据.
主要方法:
- 开发了FedGMMAT,一种使用联合统计方法的联合遗传关联测试工具.
- 为混固定和添加多基因随机效应实施了纠正.
- 保证的遗传数据保留在当地站点;中间统计数据是加密的.
主要成果:
- FedGMMAT的结果与传统的聚合分析相提并论.
- 使用模拟和现实世界的数据集证明了有效性.
- 该框架保护隐私,需要实际的计算资源.
结论:
- FedGMMAT为保护隐私,高功率的遗传关联研究提供了可行的解决方案.
- 根据严格的隐私法规,促进机构合作和数据共享.
- 通过克服数据访问限制,推进遗传流行病学领域.
相关概念视频
Friedman Two-way Analysis of Variance by Ranks
177
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
177
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Bonferroni Test
2.7K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.7K
Two-Way ANOVA
2.6K
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
2.6K
Comparing the Survival Analysis of Two or More Groups
170
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
170


