使用可解释机器学习预测和解释反复发生的虐待儿童:来自韩国国家级虐待儿童数据的证据 (2017-2020年)
Donghun Kim1, Ting Jiang2, Kihyun Kim3
1School of Information Management, Nanjing University, No.163, Xianlin Road, Qixia District, Nanjing, China.
Social science & medicine (1982)
|November 23, 2025
概括
机器学习模型预测了虐待儿童的复发风险. 确定的关键因素包括犯罪者的年龄和虐待类型,有助于针对性干预儿童保护.
科学领域:
- 儿童保护研究 儿童保护研究
- 机器学习在社会科学中的应用.
- 公共卫生监督是对公共卫生的监督.
背景情况:
- 儿童虐待的复发对儿童福利构成重大风险.
- 了解反复虐待的预测因素对于有效干预至关重要.
- 国家数据为分析虐待儿童模式提供了坚实的基础.
研究的目的:
- 开发机器学习模型来预测反复发生的虐待儿童风险.
- 确定每个虐待类型 (身体,情感,性,忽视) 特定的关键风险因素.
- 解释个别的复发风险,并为定制的预防策略提供信息.
主要方法:
- 利用国家儿童虐待数据,包括儿童,事者和报告的案件.
- 开发了针对经常性身体虐待 (RPA),经常性情绪虐待 (REA),经常性性性虐待 (RSA) 和经常性疏忽 (RN) 的独特预测模型.
- 分析了各种各样的因素,包括儿童人口统计,事者特征,虐待事件细节和服务干预.
主要成果:
- 实现了复发性疏忽的最高预测准确度 (AUC-ROC 0.793),其次是RSA (0.749),REA (0.702) 和RPA (0.700).
- 年轻的犯罪者年龄与所有虐待类型的风险增加有关.
- 对儿童和犯罪者的咨询减轻了RPA,REA和RSA的风险.
- 同类虐待的复发对RPA,REA和RN有很大的影响;犯罪者的性别 (男性) 对RSA具有重要意义.
- 育儿态度,知识,技能和家庭环境冲突对RSA至关重要.
结论:
- 机器学习模型为预测虐待儿童的复发提供了有价值的工具.
- 识别特定的风险因素,可以定制预防和干预策略.
- 调查结果支持儿童保护机构在早期发现和为处于危险的家庭提供定制支持.
相关概念视频
Steps in Outbreak Investigation
472
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
472
Residuals and Least-Squares Property
8.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.9K
