Related Experiment Video
Updated: Apr 21, 2026

Implantation of Electroencephalogram and Electrocardiogram Telemetry Devices in Neonatal Rabbit Kits
Published on: February 28, 2025
Inter-rater reliability between inexperienced and experienced raters using a rabbit sedation scale
Patricia Ruíz-López1, Vicente Soler-Rubio2, Juan Morgaz3
1Small Animal Department, Faculty of Veterinary Medicine, University of Ghent, Merelbeke, Belgium.
Objective:
To assess the reliability of a published rabbit sedation scale when used by inexperienced raters compared with experienced raters.
Study Design:
Experimental, randomised and blinded study.
Animals:
A group of 16 rabbits, 3-4 months old and 1.9 ± 0.5 kg (mean ± standard deviation).
Methods:
Rabbits were premedicated intramuscularly with midazolam 1 mg kg-1 and butorphanol 0.5 mg kg-1 (n = 8) or midazolam 1 mg kg-1 and methadone 1 mg kg-1 (n = 8). After a short explanation of the sedation scale, two experienced raters (one Diplomate of the European College of Veterinary Anaesthesia and Analgesia, and one Diplomate of the European College of Zoological Medicine) and four inexperienced raters (veterinary students) applied the sedation scale before sedation and at 10 and 20 minutes after sedation. Weighted kappa (κw) with linear weights and Gwet's agreement coefficient AC2 (Gwet's AC2) were used to determine inter-observer agreement between the two Diplomates, and Gwet's AC2 and Krippendorff's α (αk) coefficients were used to evaluate inter-observer agreement between the four students for each of the eight items. Agreement was considered when κw = 0.81-1.000. Lin's concordance correlation coefficient was used to evaluate the concordance. A paired t-test and Gwet's AC2 coefficient were used to compare the groups.
Results:
There was agreement within experienced [κw = 0.981 (0.970-0.992), p = 0.001; Gwet's AC2 = 0.973 (0.943-1.000), p = 0.001] and inexperienced raters [αk = 0.658 (0.618-0.699), p = 0.001; Gwet's AC2 = 0.741 (0.705-0.778), p = 0.001]. There were statistically significant differences depending on the experience of the rater (p = 0.001). The sedation protocol did not affect the agreement between raters (p = 0.084).
Conclusions And Clinical Relevance:
The rabbit sedation scale was reliable when used by experienced or inexperienced raters.

