在一个高维逻辑回归模型中,对异质子群的偏差推断
Hyunjin Kim1, Eun Ryung Lee2, Seyoung Park3
1Department of Statistics, Sungkyunkwan University, Seoul, 100190, South Korea.
Scientific reports
|December 11, 2023
概括
这项研究引入了一种新的统计方法,用于分析癌症细胞系中复杂,异质的数据. 合并组拉索方法有效量化了跨子群体的共变效应,改善了对高维二进制反应的推断.
科学领域:
- 统计 统计 统计 统计
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 数据异质性在科学研究中很常见,特别是在复杂的数据集和子群中.
- 在异质数据中分析高维二进制响应提出了重大的推断挑战.
- 现有的方法往往难以有效量化跨子群体的共同变量效应.
研究的目的:
- 开发一种新的统计推理方法,用于在异质子群的存在下进行高维逻辑回归.
- 调查和量化跨子群体的协变效应的异质性.
- 测试复杂数据集中的共同变量效应的等价性和意义.
主要方法:
- 为整体稀疏性和融合系数提出了一个融合组拉索惩罚方法.
- 开发了一种偏差纠正的统计推理方法.
- 适应近接梯度和交替方向方法的乘法器 (ADMM) 计算效率.
- 为合并的拉索组提供了非对称的分析,并为 debiased 测试统计数据提供了千平方近似.
主要成果:
- 拟议的方法有效地考虑了高维物流回归中的数据异质性.
- 与现有方法相比,模拟显示出更高的性能.
- 合并后的拉索群体实现了跨子群体的系数的稀疏性和融合.
- 经过偏差测试的统计数据被证明允许基平方近似值.
结论:
- 这种新的统计推理方法提供了一个强大的方法,用于分析具有高维二进制数据的异质子群.
- 该方法提供了准确和高效的统计推断,优于现有技术.
- 通过对癌症细胞系百科全书 (CCLE) 数据的分析,证明了其实用性.
更多相关视频
相关概念视频
Test for Homogeneity
2.0K
The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can...
2.0K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
133
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
133
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Choosing Between z and t Distribution
2.8K
The z and the Student t distribution estimate the population mean using the sample mean and standard deviation. However, to decide which distribution to use for a calculation, one needs to determine the sample size, the nature of the distribution, and whether the population standard deviation is known. If the population standard deviation is known and the population is normally distributed, or if the sample size is greater than 30, the z distribution is preferred. The Student t distribution is...
2.8K
Goodness-of-Fit Test
3.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.4K


