使用机器学习通过警方数据和超级学习来预测家庭杀人事件
Jacob Verrey1, Barak Ariel2,3, Vincent Harinam2
1Institute of Criminology, University of Cambridge, Sidgwick Ave, Cambridge, CB3 9DA, UK. jjv31@cam.ac.uk.
Scientific reports
|December 22, 2023
概括
这项研究表明,机器学习可以使用警察数据预测国内杀人事件. 一种新的"超级学习者"模型与现有方法相比,显著提高了预测准确性.
科学领域:
- 犯罪学 犯罪学
- 数据科学数据科学数据科学
- 机器学习 机器学习
背景情况:
- 现有的国内杀人预测工具往往缺乏预测有效性或专注于非致命的复发.
- 需要准确和实用的方法来预测家庭杀人事件.
研究的目的:
- 探索在警察数据上使用机器学习的可行性,以预测国内杀人事件.
- 开发和评估一个"超级学习者"模型,以提高预测准确度.
- 仅使用警察记录来评估机器学习模型的实际实用性.
主要方法:
- 实现一个"超级学习者"整体机器学习模型.
- 使用专门来自伦敦大都会警察局的数据集.
- 将超级学习者的表现与现有的国内杀人预测工具进行比较.
主要成果:
- 超级学习者模型实现了77.64%的回忆和18.61%的精度得分.
- 该模型显示曲线下面面积 (AUC) 为71.04%,表明出色的预测性能.
- 超级学习者显著超过了所有现有的国内杀人预测工具.
结论:
- 机器学习,特别是超级学习者方法,是预测国内杀人的可行和有效方法.
- 该模型的表现表明,它对改善对家庭暴力的理论理解,未来的研究和实际干预有重大影响.
- 仅仅使用警察数据就足以构建一个高效的预测模型,这突显了它的实用性.
更多相关视频
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.3K
08:20Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
1.5K
相关概念视频
Steps in Outbreak Investigation
131
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
131
End Point Prediction: Gran Plot
329
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
329
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Statistical Methods for Analyzing Epidemiological Data
371
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
371
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Correlation and Regression
1.3K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
1.3K
