使用催化先前分布的考克斯回归模型的贝叶斯推理
1Department of Statistics & Data Science, National University of Singapore, 117546, Singapore.
Biometrics
|February 3, 2026
概括
我们在考克斯模型中引入了贝叶斯推理的考克斯催化先验,改善了小样本大小的稳定性. 该方法通过提供强大的替代标准推断技术来增强生存数据分析.
科学领域:
- 统计 统计 统计 统计
- 生物统计学 生物统计学
- 生存分析的分析.
背景情况:
- 考克斯的比例危险模型 (Cox模型) 被广泛用于生存数据.
- 考克斯模型中的标准推断方法在与模型尺寸相对较小的样本大小方面面临挑战.
- 现有的方法可能无法在高维设置中足够稳定复杂的参数模型.
研究的目的:
- 提出一种新的贝叶斯方法,Cox催化先验,用于增强Cox模型推理.
- 为了解决在小样本,高维度场景中标准最大部分概率推理的局限性.
- 为考克斯模型提供稳定和一致的估计方法.
主要方法:
- 使用合成数据和替代基线危险的Cox催化先验的配方.
- 从更简单的拟合模型的预测分布生成合成数据.
- 导出一个近似的边际后端模式作为一个规则化的日志局部概率估计器.
主要成果:
- 建议的Cox催化先验已被证明在温和条件下是适当的.
- 由此产生的估计器证明了一致性.
- 与标准的最大部分概率推断相比,模拟研究显示出更高的性能,与现有的收缩方法相比,结果可比.
结论:
- 考克斯催化先验为考克斯模型推理提供了一个强大的和有效的贝叶斯方法,特别是在挑战小样本大小的场景中.
- 该方法提供了稳定和一致的估计器,性能优于传统技术.
- 该方法适用于现实世界的生存数据分析.
相关概念视频
Regression Toward the Mean
7.0K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.0K
Multiple Regression
4.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.0K
Correlation and Regression
3.4K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.4K
Regression Analysis
8.4K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.4K
Turnover Number and Catalytic Efficiency
21.6K
The turnover number of an enzyme is the maximum number of substrate molecules it can transform per unit time. Turnover numbers for most enzymes range from 1 to 1000 molecules per second. Catalase has the known highest turnover number, capable of converting up to 2.8×106 molecules of hydrogen peroxide into water and oxygen per second. Lysozyme has the lowest known turnover number of half a molecule per second.
Chymotrypsin is a pancreatic enzyme that breaks down proteins during digestion....
Chymotrypsin is a pancreatic enzyme that breaks down proteins during digestion....
21.6K
Catalytically Perfect Enzymes
5.1K
The theory of catalytically perfect enzymes was first proposed by W.J. Albery and J. R. Knowles in 1976. These enzymes catalyze biochemical reactions at high-speed. Their catalytic efficiency values range from 108-109 M-1s-1. These enzymes are also called 'diffusion-controlled' as the only rate-limiting step in the catalysis is that of the substrate diffusion into the active site. Examples include triose phosphate isomerase, fumarase, and superoxide dismutase.
Most enzymes...
Most enzymes...
5.1K


