机器学习预测波兰自杀企图数量:来自谷歌趋势和历史数据的洞察力
Michał Walaszek1, Zofia Kachlik1, Wojciech Nazar2
1Department of Psychiatry, Faculty of Medicine, Medical University of Gdansk, ul. Smoluchowskiego 17, 80-214 Gdańsk, Poland.
International journal of clinical and health psychology : IJCHP
|October 29, 2025
概括
机器学习使用谷歌趋势数据准确地预测了波兰每月的自杀企图. 随机森林模型显示出高准确度,识别了诸如焦虑和社会隔离等关键预测因素,以改善公共卫生战略.
科学领域:
- 公共卫生 公共卫生
- 数据科学数据科学数据科学
- 精神病学是一个精神病学.
背景情况:
- 自杀是一种复杂的公共卫生问题,具有重要的生物心理社会因素.
- 它是发达国家的主要死亡原因,需要先进的预测方法.
- 机器学习 (ML) 提供了分析趋势和告知预防策略的潜力.
研究的目的:
- 利用机器学习 (ML) 预测波兰每月的自杀数量.
- 评估谷歌趋势数据与ML一起用于自杀预测的有效性.
- 识别表明自杀风险的关键搜索术语.
主要方法:
- 波兰国家警察 (2013-2023) 每月自杀企图数据的分析.
- 自杀数与40个自杀和心理健康术语的相对搜索量 (RSV) 的相关性.
- 四个ML模型的比较:线性回归,随机森林,支向量回归 (SVR) 和XGBoost回归.
主要成果:
- 随机森林回归表现出卓越的表现,在一般人群中达到0.909的PCC,在成年人群中达到0.853.
- 发现的关键预测因素包括与"焦虑障碍"",精神科医生"和"社会隔离"相关的术语.
- 随机森林模型的平均绝对百分比误差 (MAPE) 在一般人群中低至6.78%.
结论:
- 这项研究强调了谷歌趋势数据和ML在国家一级预测自杀企图方面的潜力.
- 这些发现可以为预防自杀的公共卫生策略提供信息和增强.
- 建议使用更高分辨率数据进行进一步的研究,以完善预测准确性.
相关概念视频
Steps in Outbreak Investigation
485
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
485
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Comparing the Survival Analysis of Two or More Groups
548
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
548
Residuals and Least-Squares Property
9.0K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.0K
Survival Tree
382
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
382
Statistical Methods for Analyzing Epidemiological Data
889
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
889

