Related Experiment Videos
International variation in histologic grading is large, and persistent feedback does not improve reproducibility
Peter N Furness1, Nicholas Taub, Karel J M Assmann
1Clinical Sciences Laboratories, Leicester General Hospital, Leicester, UK. peter.furness@le.ac.uk
The American Journal of Surgical Pathology
|May 27, 2003
Summary
Histologic grading systems show significant international variation, exceeding typical interobserver differences. Efforts to improve reproducibility of the Banff classification for renal allografts through feedback and standardized images were largely unsuccessful.
Area of Science:
- Nephrology
- Pathology
- Transplantation Immunology
Background:
- Histologic grading systems are crucial for diagnosis, therapy, and auditing in medicine.
- Reproducibility of these systems is often assessed in small, homogenous groups, potentially underestimating international variability.
Purpose of the Study:
- To evaluate the international reproducibility of the established Banff classification for renal allograft pathology across Europe.
- To assess the effectiveness of interventions, including individual feedback and standardized image grading, in improving reproducibility.
Main Methods:
- The Banff classification was applied to renal allograft biopsies by pathologists across Europe.
- Individual feedback was provided after small case batches.
- Participants graded standardized photographs of selected slides to control for slide-specific variations.
Main Results:
- Kappa values for Banff classification features were lower than previously reported, indicating substantial international variation.
- Interventions, including prolonged feedback and grading of photographs, showed limited success in improving reproducibility for most features.
- Features defined by "area affected" demonstrated particular resistance to improvement.
Conclusions:
- International variation in applying the Banff classification is significant and greater than previously assessed interobserver variation.
- Standardized feedback and image grading methods were insufficient to overcome this variation.
- The findings highlight the risk of relying on grading systems with inconsistent application across institutions, potentially impacting patient care.