使用XGBoost和SHAP解释公民在政策支持中对荷兰重新实施COVID-19措施的差异
Jose Ignacio Hernandez1,2, Sander van Cranenburgh2, Marijn de Bruin3,4
1Center of Economics for Sustainable Development (CEDES), Faculty of Economics and Government, Universidad San Sebastian, Concepción, Chile.
Quality & quantity
|March 25, 2025
概括
了解公民对COVID-19政策的支持,需要考虑个人差异. 这项研究揭示了基于年龄,感知风险和意见权重的支持显著差异,为定制的公共卫生战略提供了信息.
科学领域:
- 社会科学 社会科学 社会科学
- 公共卫生 公共卫生
- 计算社会科学 计算社会科学
背景情况:
- 之前的研究确定了公众支持COVID-19政策的驱动因素.
- 然而,对这些驱动因素如何对个体产生不同影响的细致理解仍然未被探索.
- 这种差距可能会导致误导政策决策,因为公众论的异质性没有得到解决.
研究的目的:
- 确定影响个人层面对COVID-19措施支持的因素.
- 分析这些影响因素在不同公民群体和政策类型中的分布.
- 为了解决以前分析的局限性,这些分析忽视了政策支持的个体变化.
主要方法:
- 利用XGBoost,一个监督的机器学习算法,用于预测建模.
- 采用SHAP (Shapley添加式解释) 来解释模型预测并确定关键驱动因素.
- 分析了一项参与式价值评估 (PVE) 实验的二次数据,该实验涉及1888名荷兰公民,涉及四种风险场景.
主要成果:
- 在公民对各种COVID-19措施的支持中发现了显著的异质性.
- 导致这种异质性的关键因素包括年龄组,公众论的感知重要性以及个人对感染COVID-19的风险感知.
- 机器学习分析揭示了传统数据分析方法中无法发现的变异.
结论:
- 公民对COVID-19政策的支持并不统一,并且在个人层面上有很大差异.
- 政策制定者可以利用这些发现来设计更有针对性和有效的公共卫生干预措施.
- 根据人口因素和风险感知量身定制措施可以提高整体政策的接受和遵守.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
107
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
107
Bias in Epidemiological Studies
112
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
112
Residuals and Least-Squares Property
7.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.2K
Statistical Methods for Analyzing Epidemiological Data
266
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
266
Confounding in Epidemiological Studies
121
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
121
Social Proof
27.3K
Social proof is a form of persuasion based on comparison and conformity. People compare their behavior and actions to what others are doing and will change to conform to do what their peers do.
27.3K


