在地理加权回归中适应空间异质性
Tengdi Zheng1, Rong Li2, Mixia Wu1
1Department of Statistics and Data Science, School of Mathematics, Statistics and Mechanics, Beijing University of Technology, Beijing, China.
Statistics in medicine
|August 20, 2025
概括
这项研究引入了地理加权群体激光回归 (GWGPL) 来分析来自多个地点的复杂调查数据. 该方法有效地处理空间差异,并确定关键的健康成本变量,改善估计和预测.
科学领域:
- 统计数据
- 空间分析
- 健康经济学
背景情况:
- 分析多个地点的调查数据越来越常见.
- 现有的方法往往无法解释空间异质性,导致效果不佳.
- 需要能够执行变量选择并适应空间变化的方法.
研究的目的:
- 开发一种新的统计方法,即地理加权群拉索回归 (GWGPL),用于分析多个地点的调查数据.
- 有效地解决空间异质性,并在不同地理位置进行共享变量选择.
- 严格证明拟议的GWGPL方法的选择和估计特性.
主要方法:
- 开发地理加权群拉索回归模型 (GWGPL).
- 理论分析以证明GWGPL的选择和估计特性.
- 将GWGPL应用于中国社会调查数据 ("一千人,一百个村庄").
主要成果:
- GWGPL有效地适应空间异质性,并执行共享变量选择.
- 模拟研究表明GWGPL在竞争优势方面优于其他方法.
- 对健康调查数据的分析确定了与住院,门诊和自我治疗费用相关的关键变量.
结论:
- 拟议的GWGPL方法为分析空间异质数据提供了可靠的方法.
- 在分组结构,系数估计流性和预测准确性方面,GWGPL表现出卓越的性能.
- 已识别的变量提供了对受研究群体医疗费用影响因素的见解.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
279
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
279
Selected Data About Geographic Locations
66
Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
66
One-Way ANOVA: Unequal Sample Sizes
5.9K
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
5.9K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
One-Way ANOVA: Equal Sample Sizes
3.4K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.4K
Multiple Regression
3.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.2K


