基于机器学习的自杀预测和美国各县自杀脆弱性指数的开发
Vishnu Kumar1, Kristin K Sznajder2, Soundar Kumara3
1Department of Industrial and Manufacturing Engineering, The Pennsylvania State University, University Park, PA, USA. vbk5101@psu.edu.
Npj mental health research
|April 12, 2024
概括
美国的自杀率正在上升. 一个新的机器学习模型使用人口和人口统计数据预测县级自杀脆弱性,帮助有针对性的预防工作.
科学领域:
- 公共卫生 公共卫生
- 数据科学数据科学数据科学
- 计算流行病学计算流行病学
背景情况:
- 自杀在美国是一个显著且日益增长的公共卫生挑战.
- 了解和预测自杀模式对于有效的控制和预防战略至关重要.
研究的目的:
- 分析自杀趋势和2010-2019年美国各县的地理分布.
- 开发和验证用于县级自杀预测的机器学习模型.
- 确定影响自杀率的主要人口特征.
主要方法:
- 利用涵盖2010-2019年所有3140个美国县的公开数据.
- 开发了一个XGBoost机器学习模型,包含17个功能.
- 采用了夏普利添加式解释 (SHAP) 来确定特征的重要性.
主要成果:
- 在许多县观察到自杀率的显著增加;大约25%的自杀率至少增加了10%,12%的自杀率超过了50%.
- 该XGBoost模型实现了高预测准确度,R2值为0.98.
- 总人口, % 非洲裔美国人, % 白人, 平均年龄和 % 女性人口被确定为前五大预测因素.
结论:
- 一个新的自杀脆弱性指数 (SVI) 是使用前5个预测特征开发的.
- 该SVI可以确定美国自杀风险较高的县.
- 该工具支持针对性自杀预防和控制计划的知情决策.
相关概念视频
Cancer Survival Analysis
345
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
345
Steps in Outbreak Investigation
125
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
125
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Kaplan-Meier Approach
135
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
135
Applications of Life Tables
62
Life tables are versatile across various fields, providing a quantitative basis for analyzing mortality and survival rates. Whether used by demographers, actuaries, epidemiologists, or sociologists, life tables offer valuable insights into the dynamics of life and death, facilitating informed decisions in public health, insurance, conservation, and beyond. Their broad applicability highlights the interconnectedness of demographic data with practical outcomes in everyday life and strategic...
62
Statistical Methods for Analyzing Epidemiological Data
364
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
364


