Related Experiment Videos
Interobserver Agreement in Acute Toxicity Assessment During Head and Neck Radiotherapy/Chemoradiotherapy:
Shwetabh Sinha1, Anuj Kumar1, Samarpita Mohanty1
1Department of Radiation Oncology, ACTREC, Tata Memorial Centre, Homi Bhabha National Institute, Navi Mumbai, India.
Background:
Acute toxicity is a key endpoint in head and neck radiotherapy (RT) and chemoradiotherapy (CRT) trials and is most commonly reported as maximum toxicity per patient. However, the reproducibility of clinician-reported toxicity grading remains incompletely characterized. We evaluated interobserver agreement in acute toxicity assessment, focusing on maximum toxicity as the primary endpoint.
Methods And Materials:
This predefined subgroup analysis was conducted within a prospective randomized controlled trial of patients with head and neck squamous cell carcinoma (HNSCC) undergoing curative-intent RT/CRT. Acute toxicities were assessed weekly using CTCAE v5.0 by two independent, blinded radiation oncologists. Four domains were evaluated: dermatitis, oral mucositis, dysphagia, and xerostomia. The primary endpoint was patient-level maximum toxicity, dichotomized as ≥Grade 2 and ≥Grade 3. Interobserver agreement was assessed using Cohen's κ.
Results:
Seventy-four patients were included. For maximum ≥Grade 2 toxicity, incidence (O1 vs O2) was: dermatitis 55.4% vs 45.9%, mucositis 91.9% vs 70.2%, dysphagia 95.9% vs 91.9%, and xerostomia 55.4% vs 58.1%, with κ of 0.33, 0.35, 0.41, and 0.23, respectively (fair to moderate). For ≥Grade 3 toxicity, κ values were: dermatitis 0.79, mucositis 0.41, dysphagia 0.63, and xerostomia 0.25. Pooled weekly agreement for ≥Grade 2 toxicity was higher (κ = 0.49-0.89), with peak concordance during weeks 5-6.
Conclusions:
Interobserver agreement for maximum acute toxicity was fair to moderate across domains in head and neck RT/CRT, with the largest discordance for oral mucositis even between observers at the same centre. Although maximum toxicity remains a pragmatic and widely used endpoint, caution is warranted in its interpretation, and weekly assessment may offer a more reproducible complementary endpoint, particularly for cross-trial comparisons.