Related Experiment Video
Updated: Sep 16, 2026

Image Recognition and Parameter Analysis of Concrete Vibration State Based on Support Vector Machine
Published on: January 5, 2024
From Alarms to Probabilities: Stratified Human Review, Label-Noise Correction, and Calibrated Risk Grading for
Tao Feng1, Kun Chen1, Jing Wang1
1School of Computer and Artificial Intelligence, Beijing Technology and Business University, Beijing 102488, China.
Abstract:
Industrial anomaly detectors emit binary alarms, but production lines need graded dispositions. Calibrating alarms into fault probabilities requires ground truth, coming only from noisy human review treated as exact. We report a deployed closed loop on a reciprocating-compressor line (46,023 units, nine test campaigns, four fused detector legs). A stratified review of 1161 units audited a fusion-score grading and refuted its assumed monotonicity: precision was 25.2%/13.8%/26.1% for high/medium/low tiers. The cause was correlated false positives: two legs firing on shared broadband transients agreed on 480 units at 21.0% precision-detector agreement is not independent evidence-whereas one periodicity feature was monotone (27.8% → 60.0% → 100%). A blind test against seeded fault units (hardware ground truth) measured reviewer sensitivity at 0.905 and specificity at 0.421 on hard cases; Rogan-Gladen correction restored monotonicity (95% of bootstrap replicates; 83% under campaign-cluster resampling), exposed the low-tier advantage as a label-noise artifact, and re-estimated no-alarm prevalence at 4-10% versus the observed 13.3%. Corrected evidence drove a redeployed rule calibrating tiers at ≈67%/27%/17% under review budgets (≤1%/≤3%/≤9% of production). The methodology-stratified audit, seeded-fault blind testing, and evaluation-side prevalence correction-transfers to any human-verified system.