聚类当前状态数据的联合回归分析与潜在变量
Yanqin Feng1, Sijie Wu1,2, Jieli Ding1
1School of Mathematics and Statistics, Wuhan University, Wuhan, P.R. China.
Statistical methods in medical research
|October 23, 2024
概括
这项研究引入了集群当前状态数据的新联合建模方法,有效地处理未观察到的因素和生存分析中的信息集群大小. 该方法提高了分析复杂健康和环境研究的准确性.
科学领域:
- 生物统计学 生物统计学
- 生存分析的分析.
- 统计建模 统计建模
背景情况:
- 聚类当前状态数据在生存研究中很常见.
- 未观察到的因素 (隐性变量) 和集群大小可以影响结果.
- 现有的方法可能无法完全解释这些复杂性.
研究的目的:
- 为分析聚类当前状态数据提出联合建模方法.
- 纳入潜在变量和潜在的信息集群大小.
- 为具有复杂依赖性的生存数据提供一个强大的统计框架.
主要方法:
- 一个联合模型,结合了潜在变量因子分析和添加性危险脆弱性模型.
- 使用预期最大化算法和加权估计方程进行参数估计.
- 建立理论属性,包括估计器的一致性和非对称的正常性.
主要成果:
- 拟议的方法有效地使用替代标记器模拟潜在变量.
- 它解释了集群内部的相关性和信息集群大小.
- 模拟研究表明有限样本表现良好.
结论:
- 开发的联合建模方法为集群的当前状态数据提供了一个强大的工具.
- 它适用于各种领域,包括毒理学和健康研究.
- 该方法在复杂的生存数据中为共变量效应提供了可靠的估计.
相关概念视频
Regression Analysis
5.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.6K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Friedman Two-way Analysis of Variance by Ranks
150
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
150
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Correlation and Regression
1.2K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.2K
Comparing the Survival Analysis of Two or More Groups
155
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
155


