Related Experiment Videos
Measuring Expert Inter-Rater Agreement with a Semi-Automated Clinical Decision Support System: Who Agrees with What?
Yasmin Benhalim1, Matthew W Hodgman2, John R Charpie3
1University of Michigan Medical School, Michigan, United States, Ann Arbor.
Abstract:
OBJECTIVE: Assess inter-rater agreement on clinician- and clinical decision support (CDS)-driven diuretic titrations and risk classifications to benchmark a semi-autonomous CDS (OTTO-FM).
Abstract:
METHODS: Secondary analysis of prospectively collected porcine data modeling postoperative fluid overload. Three pediatric cardiac intensivists rated items (reasonable/unsure/unreasonable) for clinician-driven (human phase; 29 items) and CDS-driven titrations with risk classification (CDS phase; 44 items). Inter-rater agreement was estimated with ordinal-weighted Gwet's AC2.
Abstract:
RESULTS: Dosing inter-rater agreement was similar between phases (human AC2: 0.79 [0.61-0.92]; CDS: 0.83 [0.68-0.92]), though unanimity was reached on only 19 of 29 (65.5%) human-phase and 31 of 44 (70.5%) CDS-phase items. Risk-label inter-rater agreement stratified sharply by label: high-risk, 0.92 (0.74-0.95) versus low-risk, 0.54 (0.28-0.75). Disagreement also differed in structure between tasks: 19 of 20 nonunanimous risk items were minority dissent (one rater against a two-rater majority), including 16 of 17 among low-risk items, whereas 7 of 13 nonunanimous CDS dosing items were three-way splits.
Abstract:
CONCLUSION: Inter-rater agreement on dosing was substantial and similar across phases, but not unanimous, suggesting that perfect reasonableness consensus with CDS may be unrealistic. Agreement on risk classification was lower, and the structure of disagreement differed between CDS dose and risk. Empirical inter-rater agreement baselines may contextualize CDS evaluation, and these differences may have implications for the design of semi-autonomous CDS if confirmed in larger studies.