对一般混合的反复事件数据的回归分析.
Ryan Sun1, Dayu Sun2, Liang Zhu3
1Department of Biostatistics, University of Texas MD Anderson Cancer Center, Houston, TX, USA. rsun3@mdanderson.org.
Lifetime data analysis
|July 12, 2023
概括
本研究引入了一种新的回归分析方法,用于含有混合重复事件数据的复杂生物医学数据. 建议的最大概率方法有效处理不完整的数据,提高反复事件分析的准确性和效率.
科学领域:
- 生物统计学 生物统计学
- 流行病学 流行病学
- 医疗信息学 医疗信息学
背景情况:
- 生物医学数据集经常包含不完整的反复结果数据.
- 循环事件信息通常是循环事件,面板计数和面板二进制数据的混合,称为一般混合循环事件数据.
- 现有的方法缺乏这种组合数据结构的既定回归分析,导致具有潜在缺点的临时解决方案.
研究的目的:
- 为一般混合反复事件数据提出一种新的最大概率回归估计程序.
- 确定拟议的估计器的非对称性质.
- 扩展该方法,以适应终端事件在循环事件分析.
主要方法:
- 开发一个最大概率估计程序,用于组合的一般混合反复事件数据.
- 为开发的估计器理论上建立了非对称性质.
- 方法的概括,以纳入终端事件.
主要成果:
- 建议的最大概率方法为分析混合的反复事件数据提供了一个强大的方法.
- 该程序在数值模拟中表现出良好的性能.
- 适用于儿童癌症幸存者研究验证了该方法的实际实用性.
结论:
- 开发的最大概率回归程序有效地解决了分析一般混合反复事件数据的挑战.
- 该方法比临时技术提供了改进,提高了稳定性和效率.
- 一般化方法适用于复杂的生物医学数据,包括那些具有终端事件的生物医学数据.
相关概念视频
Regression Analysis
5.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.8K
Censoring Survival Data
144
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
144
Comparing the Survival Analysis of Two or More Groups
226
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
226
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


