Related Experiment Video
Updated: May 26, 2026

Software-Assisted Quantitative Measurement of Osteoarthritic Subchondral Bone Thickness
Published on: March 18, 2022
Assessing the Reliability of the SICOT Orthopaedic Research Award Selection Process: An Analysis of the Score and the
Shamim Beigh1, Jason P Cheung2, Rolando Del Cruz3
1Anaesthesia, Al Jalila Children's Speciality Hospital, Dubai, ARE.
Abstract:
This study examined the reliability and consistency of the evaluation process used in selecting recipients for the annual Société Internationale de Chirurgie Orthopédique et de Traumatologie (SICOT) Research Award. The primary aim was to assess both inter-rater and intra-rater reliability among the assessors and to evaluate the internal consistency and discriminative capacity of the scoring criteria. A total of 45 anonymised research applications were scored by eight members of the SICOT Research Award Selection Committee. Each application was evaluated across 10 criteria using a 5-point Likert scale, with a maximum total score of 50. The study employed intraclass correlation coefficients (ICCs) to assess reliability, Cronbach's alpha to measure internal consistency, and discrimination indices to evaluate the performance of individual scoring items. The results revealed significant variability between assessors. Inter-rater reliability was poor, with ICCs of 0.289 and 0.190 across two evaluation rounds, suggesting inconsistent scoring among different assessors. In contrast, intra-rater reliability, measuring the consistency of individual assessors over time, showed moderate to good agreement, with an ICC of 0.705. The scoring system exhibited excellent internal consistency (Cronbach's alpha = 0.926), which improved to 0.954 when the "Language" criterion was excluded, indicating that this item may reduce the coherence of the scale. While high internal consistency (Cronbach's alpha > 0.90) indicates reliability, it may also reflect redundancy and a potential halo effect, and should therefore be interpreted as a limitation rather than a strength. All 10 criteria demonstrated acceptable discriminative power, with the "Aesthetic" item scoring the lowest and the total score showing the highest discrimination index. These findings highlight the strengths and limitations of the current evaluation framework. While the scale itself is internally consistent and capable of differentiating between high- and low-quality applications, inconsistencies in inter-rater scoring point to a need for clearer guidelines and more structured assessor training. The study recommends refining the scoring rubric and exploring AI-based tools to support more objective and transparent evaluation in future selection cycles.
