Related Experiment Video
Updated: Mar 21, 2026

Systematic Assessment of Mammalian Skull Specimens for Dental and Temporomandibular Joint Pathology
Published on: August 22, 2022
Interrater Variability in the Application of the Goldman Criteria to Medical Autopsies
Context.—:
The Goldman criteria are widely used to categorize diagnostic discrepancies in autopsy reports for hospital quality assurance. Despite their adoption, little is known about how consistently they are applied across institutions.
Objective.—:
To evaluate the interrater reliability among academic pathologists applying the Goldman criteria to classify missed diagnoses in autopsy cases.
Design.—:
Thirty autopsy cases from 2 academic hospitals were reviewed by 5 experienced autopsy pathologists. Each reviewer independently assigned Goldman classification scores (classes I-IV) to missed diagnoses based on the original 1983 definitions. Consensus was defined as at least 60% agreement (≥3 of 5 raters). Fleiss κ and Gwet AC1 were calculated to assess interrater reliability.
Results.—:
A total of 196 missed diagnoses were identified by at least 1 rater. Only 26% (51 of 196) achieved a consensus Goldman classification, and just 4% (8 of 196) had unanimous classification agreement. Most diagnoses (65%; 127 of 196) were not identified by at least 3 raters. Among consensus classification ratings, 49% (25 of 51) were class IV. Fleiss κ was 0.176 (slight agreement) while Gwet AC1 was 0.390 (fair agreement). Collapsing major (classes I-II) and minor (classes III-IV) categories modestly improved agreement.
Conclusions.—:
Interrater agreement in applying the Goldman criteria is only fair to moderate, suggesting significant subjectivity in their application. As the criteria are used for institutional comparison and quality metrics, efforts to standardize definitions and provide reviewer training are warranted to improve reproducibility and utility.

