犯罪再発率をリアルタイムで予測するための公平性スケール
Jacob Verrey1, Peter Neyroud1, Lawrence Sherman1,2
1Institute of Criminology, University of Cambridge, Sidgwick Ave, Cambridge, CB3 9DA UK.
Neural computing & applications
|September 4, 2025
まとめ
機械学習モデルは 再犯を正確に予測し 既存の方法を上回ります 刑事司法と犯罪の削減に偏ったアプローチを提供することで 人口統計学的な公平性を実現します
科学分野:
- 刑事司法
- 機械学習
- データサイエンス
背景:
- 犯罪の再発を予測することは 資源の配分と公衆の安全にとって極めて重要です
- 既存の予測モデルは 社会的バイアスを永続させ 不公平な結果につながります
- 警察国家コンピュータ (PNC) のデータセットは,再犯研究のための大規模な基盤を提供します.
研究 の 目的:
- 一般的および暴力的な再犯を予測するための機械学習モデルを開発し評価する.
- これらの予測モデル内の社会的バイアスを評価し,軽減します.
- 公正で効果的な再犯予測ツールの導入のための枠組みを提案する.
主な方法:
- イギリス警察の全国コンピュータの有罪判決データ (346,685件の記録) を使って12の機械学習モデルを生成した.
- カーブ下の面積 (AUC) スコアに焦点を当てたモデル評価の5倍クロス検証を使用しました.
- モデル予測における人口格差を定量化して対処するための新しい公平性スケールを開発しました.
主要な成果:
- 最良のモデルはAUCスコア0. 8660 (一般) と0. 8375 (暴力的な再犯) を達成し,最先端を上回った.
- 偏差のないモデルでは 公平性の定義が5つとも満たされ 人口統計学的差は1%以内でした
- 構造的なバイアスの影響が減る可能性があることが示されています.
結論:
- 機械学習は 再犯を効果的に予測し 公平性を大幅に改善します
- 開発された公平性のスケールは,刑事司法モデルをデバイズするための貴重なツールを提供します.
- 提案された導入には,保護措置とランダム化制御試験が 犯罪と偏見を減らすための実用的な効果を検証する.
関連する概念動画
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Hindsight Biases
3.9K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.9K
Residuals and Least-Squares Property
7.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.8K
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Confidence Coefficient
7.8K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.8K
Confidence Interval for Estimating Population Mean
8.0K
A point estimate of the population mean is obtained from a single sample. Such a point estimate does not represent a population well because it needs to account for variability in the population. Single point estimate can also be biased despite the sample being selected randomly. Thus, a point estimate is often unreliable. A confidence interval is needed to reduce this unreliability.
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
8.0K


