机器学习算法的构建和验证,用于在大规模的COVID-19爆发期间预测家庭隔离人员的抑郁症:基于Adaboost模型
Yiwei Zhou1,2,3, Zejie Zhang4, Qin Li5
1Business School, University of Shanghai for Science and Technology, 200093, Shanghai, China.
BMC psychology
|April 24, 2024
概括
这项研究开发了一种Adaboost机器学习模型,用于预测COVID-19期间在家隔离人员的抑郁症. 该模型在识别抑郁症水平方面表现出高准确性和有效性.
科学领域:
- 机器学习 机器学习
- 公共卫生 公共卫生
- 流行病学 流行病学
背景情况:
- COVID-19 流行病与增加的抑郁率有关.
- 准确识别家庭隔离人口的抑郁症至关重要.
- 现有的方法可能无法充分捕捉与流行病相关的心理健康挑战的细微差别.
研究的目的:
- 为在COVID-19疫情期间在家中隔离的个人构建和验证基于机器学习的抑郁症预测模型.
- 为了比较用于抑郁症预测的多个机器学习算法的性能.
- 确定一个强大的和可通用的模型来识别抑郁风险.
主要方法:
- 一项横截面研究调查了在家隔离的个人,采集了社会人口统计数据,COVID-19的影响和使用PHQ-9尺度的抑郁症.
- 一组数据被随机分为培训 (70%) 和验证 (30%) 组.
- 通过十倍交叉验证评估了15个机器学习模型,通过准确性,精度,ROC曲线和AUC来评估性能.
主要成果:
- 在家庭隔离人员中,抑郁症的患病率为31.66%.
- 该模型的准确度为0.7180和AUC为0.7803,超过了其他15种模型.
- 在验证组中,Adaboost表现强,AUC超过0.83.
结论:
- 在疫情期间,Adaboost机器学习算法有效地预测了家庭隔离人员的抑郁症.
- 开发的模型表现出卓越的机器学习性能,有效性,稳定性和通用性.
- 该模型为在公共卫生危机期间的心理健康监测和干预提供了有价值的工具.
相关概念视频
Cancer Survival Analysis
345
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
345
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Statistical Methods for Analyzing Epidemiological Data
363
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
363
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Statistical Software for Data Analysis and Clinical Trials
540
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
540
Study Design in Statistics
8.0K
A study design is a set of techniques that allow a researcher to collect and analyze data from different variables defined for a specific research problem. Statistics is commonly for effective study design and more robust experiments,
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
8.0K


