使用先前数据冲突调整贝叶斯规范回归模型.
Timofei Biziaev1, Karen Kopciuk1,2,3, Thierry Chekouo1,4
1Department of Mathematics and Statistics, University of Calgary, 2500 University Drive NW, Calgary, AB T2N 1N4 Canada.
概括
本研究介绍了一种经验贝叶斯方法来配置贝叶斯调节回归模型. 该方法在高维设置中改善了变量选择,特别是当真实效应很小时.
科学领域:
- 统计 统计 统计 统计
- 计算生物学 计算生物学
- 生物统计学 生物统计学
背景情况:
- 高维回归模型为变量选择提出了计算和理论上的挑战.
- 贝叶斯规则化的回归与收缩先验 (例如,拉普拉斯,尖和板块) 提供有效的变量选择,当先验配置良好时.
研究的目的:
- 提出一个实证贝叶斯配置方法,用于贝叶斯规则化回归中的收缩先验.
- 评估这种方法在高维线性回归模型的变量选择中的性能.
主要方法:
- 开发了一种使用先前数据冲突检查的经验贝叶斯配置.
- 将该方法应用于贝叶斯式LASSO和spike-and-slab priors. 应用该方法对贝叶斯式LASSO和spike-and-slab priors进行了应用.
- 通过高维模拟和分析COVID-19蛋白质组数据来评估性能.
主要成果:
- 拟议的经验贝叶斯配置显示出潜在的超越竞争模型的性能,特别是当真正的回归效应很小时.
- 该方法成功地应用于模拟数据和真实世界的蛋白质组数据.
结论:
- 使用先前数据冲突检查的实证贝叶斯配置为贝叶斯规范回归中的变量选择提供了强大的方法.
- 这种方法提高了复杂数据集中的贝叶斯变量选择技术的可靠性和性能.
相关概念视频
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
113
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
113
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Parametric Survival Analysis: Weibull and Exponential Methods
329
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
329
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
331
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
331
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K


