一种基于异常值去除的空间扫描统计的新型共变量调整方法
Sheng Li1, Xuelin Li2, Wei Wang3
1West China School of Public Health and West China Fourth Hospital, Sichuan University, Chengdu, Sichuan, China.
Computers in biology and medicine
|April 8, 2025
概括
空间扫描统计 (SSS) 检测到地理疾病集群,但受到共变量的影响. 一种新的方法,集群基于异常值的协变量调整 (COCA),通过重新估计协变量效应,提高了集群检测的准确性.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 地理信息系统 (GIS) 是指地理信息系统.
背景情况:
- 空间扫描统计 (SSS) 对疾病监测和流行病学集群检测至关重要.
- 共变量分布可以显著影响传统的SSS方法的准确性.
- 现有的共同变量调整技术可能不准确地估计集群内外的影响.
研究的目的:
- 开发和评估一种用于增强空间集群检测的新型协变量调整方法.
- 解决传统方法在计算特定集群内外共变量效应方面的局限性.
- 提高地理疾病集群识别的准确性和可靠性.
主要方法:
- 引入了基于异常值的聚类共变量调整 (COCA) 方法.
- 可卡 (COCA) 代地将最初检测到的集群作为异常值删除,并重新估计共变系数.
- 更新了预期病例数和重新应用集群检测以改进结果.
主要成果:
- 与传统的共变量调整 (TRA-CA) 和基于通用线性模型 (GLM) 的SSS相比,COCA表现优越.
- 在关键指标中观察到更好的准确性:灵敏度,特异性,正预测值 (PPV) 和错误分类.
- 该方法显示增强了对真实集群的检测,并减少了假阳性.
结论:
- 在存在非随机分布的共变量时,COCA为空间集群检测提供了更准确的方法.
- 这种方法很容易使用标准的统计软件 (例如R) 和专用工具 (如SatScan) 来实现.
- 对于需要精确地理集群识别的疾病监测和流行病学研究,建议使用COCA.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Detection of Gross Error: The Q Test
4.8K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
4.8K
What Are Outliers?
3.6K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.6K
Outliers and Influential Points
3.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
3.9K
Modified Boxplots
9.1K
A standard box and whisker plot informs us about the spread of the data in a given sample. One can identify the minimum value, maximum value, first quartile value, second quartile or median value, and third quartile.
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
However, the box plot does not tell the reader about outliers - values that lie far from the center of the data. We can modify the standard box and whisker plot to identify the outliers and visualize the actual spread of the data in a sample.
Initially, we calculate the adjusted...
9.1K
Cluster Sampling Method
11.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.5K


