在"大n,大p"贝叶斯稀疏回归中加速吉布斯采样的先前预先条件的结合梯度方法.
Akihiko Nishimura1, Marc A Suchard2
1Department of Biostatistics, Johns Hopkins University, Baltimore, MD.
Journal of the American Statistical Association
|March 29, 2024
概括
这项研究引入了一种用于大规模贝叶斯回归的新型算法,在复杂的医疗数据分析中显著加快后置推理. 新方法将计算时间从几周缩短到几天,使得临床共变量更有效地分析.
科学领域:
- 计算统计学 计算统计学
- 生物统计学 生物统计学
- 医疗信息学 医疗信息学
背景情况:
- 现代观测研究利用了大量的医疗保健数据库,包含数百万个观察结果和数万个预测结果.
- 在如此高维的数据中估计大量参数对于传统方法来说是具有计算挑战性的.
- 稀疏回归,特别是贝叶斯式方法与收缩先验,提供了一个解决方案,但面临着计算瓶.
研究的目的:
- 开发一种新的算法,以克服大规模观测研究 (大n和大p设置) 的贝叶斯推理中的计算瓶.
- 为了加速从高维高斯分布中重复采样,需要进行后置计算.
- 为了使复杂的临床数据集能够有效地分析风险评估.
主要方法:
- 引入了一种新的算法,它避免了对高维精度矩阵 (Φ) 的明确计算和分解.
- 利用随机向量 (z) 的生成,并使用结合梯度 (CG) 算法解决线性系统 (Φx = z).
- 开发了一种"预先条件"理论,以保证CG算法的快速融合.
主要成果:
- 这种新的算法证明了后置推理中一个数量级的加速度.
- 应用到对72,489名患者和22,175名共同变量的研究中,计算时间从两周减少到不到一天.
- 成功实现了对两种抗凝固药疗法的不良事件的高效风险评估.
结论:
- 拟议的算法显著提高了贝叶斯回归在大规模观察性健康研究中的可行性和效率.
- 预先条件理论确保了高维高斯取样的计算效率.
- 这一进步有助于更及时,更全面地分析复杂的临床数据,以获得更好的医疗洞察力.
更多相关视频
09:38Generalized Psychophysiological Interaction PPI Analysis of Memory Related Connectivity in Individuals at Genetic Risk for Alzheimer's Disease
Published on: November 14, 2017
14.9K
07:11Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
2.3K
相关概念视频
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
490
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
490
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Random Sampling Method
11.1K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
11.1K
Stratified Sampling Method
12.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.0K
Statistical Hypothesis Testing
1.9K
Hypothesis testing is a critical statistical procedure facilitating informed, evidence-based decisions. It begins with a hypothesis, which is a tentative explanation, or a prediction about a population parameter. This hypothesis can be either a null hypothesis (H0), indicating no effect or difference, or an alternative hypothesis (Ha), suggesting an effect or difference.
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
1.9K
Distributions to Estimate Population Parameter
4.1K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
4.1K
