集群多状态当前状态数据的伪值回归,具有信息性的集群大小
Samuel Anyaso-Samuel1, Dipankar Bandyopadhyay2, Somnath Datta1
1Department of Biostatistics, University of Florida, Gainesville, FL, USA.
Statistical methods in medical research
|June 16, 2023
概括
这项研究引入了一种新的统计方法来分析复杂的健康数据,提高了多状态当前状态数据的准确性,并提供了信息集群大小. 该方法增强了对聚集群体疾病进展的理解.
科学领域:
- 生物统计学 生物统计学
- 流行病学 流行病学
- 统计建模 统计建模
背景情况:
- 多州当前状态数据涉及复杂的审查和信息集群大小的潜在偏见.
- 现有的方法可能无法解释过渡结果和集群大小之间的关系,从而导致不准确的推断.
- 牙周病研究经常产生如此复杂的集群数据.
研究的目的:
- 开发一种统计方法来分析集群的多状态当前状态数据,并提供信息集群大小.
- 在信息集群大小的存在下,估计对国家占有概率的共变量效应.
- 解决未经调整的信息集群大小导致的统计推断中的潜在偏差.
主要方法:
- 对聚类多态当前状态数据的伪值方法的扩展.
- 使用非参数回归计算边际状态占用概率的计算.
- 重新加权估计方程与集群大小函数,以调整信息性.
主要成果:
- 模拟研究在各种信息性情景下证明了拟议的伪值回归的特性.
- 该方法有效地调整信息集群大小,减少推断偏差.
- 通过应用到牙周病数据集的验证.
结论:
- 拟议的伪值扩展提供了一种可靠的方法,用于分析集群的多状态当前状态数据,并提供信息集群大小.
- 这种方法提高了复杂的健康研究中共变量效应估计的准确性.
- 该方法适用于现实世界的临床数据,例如牙周病进展.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K
Estimating Population Mean with Unknown Standard Deviation
8.2K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
8.2K
Testing a Claim about Mean: Known Population SD
2.8K
A complete procedure of testing the hypothesis about a population mean is explained here.
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
2.8K


