Related Experiment Video
Updated: Aug 13, 2026

In Vivo Quantification of Hip Arthrokinematics during Dynamic Weight-bearing Activities using Dual Fluoroscopy
Published on: July 2, 2021
Inadequate structural validity of the Oxford Hip Score: evidence from Rasch and confirmatory factor analysis in a
Christian Fugl Hansen1, Michael Rindom Krogsgaard2, Claus Varnum3
1Department of Orthopaedic Surgery, Bispebjerg and Frederiksberg Copenhagen University Hospital, Bispebjerg Bakke 23, DK-2400, Copenhagen NV, Denmark.
Background And Objective:
The Oxford Hip Score (OHS) is widely used as a patient reported outcome measure in clinical studies and national quality databases. The content of the questionnaire suggests a potential two-subscale structure, a structure sparsely evaluated. Thus, the aim was to evaluate the validity of both the original unidimensional structure (ie, one total score) of OHS and a two-dimensional structure reflecting separate "Pain" and "Function" subscales.
Methods:
Preoperative and postoperative OHS item scores were available from a prospective cohort of patients undergoing total hip arthroplasty (THA). Eight homogeneous subsamples, each comprising 100 randomly chosen patients with hip osteoarthritis were separately assessed using confirmatory factor analysis (CFA) exact-fit and close-fit indices, and Rasch analysis item fit statistics. Pooled analyses were added for robustness. Both unidimensional models and two-dimensional models were considered for each subsample. Item thresholds, targeting plots, internal consistency, and floor and ceiling effects were also evaluated.
Results:
Exact CFA fit was rejected in 7 out of 8 subsamples for both the unidimensional and two-dimensional structures using pooled data (χ2; P-values <0.007). Close fit indices showed substantial variability across subsamples. Two of 8 subsamples reached acceptable thresholds for the unidimensional structure, and 3 of 8 subsamples for the two-dimensional structure (the root mean square error of approximation [0.031-0.120]; comparative fit index [0.876-0.997]; Tucker-Lewis index [0.842-0.996]; standardized root mean residual [0.059-0.106]). Rasch analysis consistently identified item misfit for all four preoperative subsamples, and minimal item misfit for the 1-year data. Pooling data increased misfit to the models. Substantial ceiling effects were evident at 1-year follow-up, inflicting disordered item threshold and poor targeting of the scale for 1-year data.
Conclusion:
Consistent fit to the statistical models was not demonstrated for the Danish OHS based on robust psychometric analyses of OHS data from a Danish prospective cohort of patients undergoing primary THA. These findings raise concerns regarding the validity of reporting the OHS as a single total score and separately as two subscales of pain and function.
Plain Language Summary:
The Oxford Hip Score (OHS) is a questionnaire commonly used to measure pain and function in patients undergoing total hip replacement. It is widely applied in clinical research and national quality databases, and results are usually reported as a single total score. Some researchers have suggested that the questionnaire may instead consist of the following two separate parts: one measuring pain and one measuring function. However, this structure has not been thoroughly tested. In this study, we evaluated how well the Danish version of the OHS measures what it is intended to measure. We analyzed responses from patients who completed the questionnaire before surgery and 1 year after surgery. To ensure robust results, we examined several independent patient groups and applied modern statistical methods designed to test whether questionnaire items work together as a valid measurement scale. We found that the OHS did not consistently meet statistical criteria for a well-functioning measurement scale, whether reported as a single total score or divided into separate pain and function subscales. Before surgery, several questions did not behave as expected statistically. One year after surgery, many patients achieved the highest possible score, leaving limited ability to distinguish between patients with very good outcomes. This reduced the precision of the questionnaire at follow-up. These findings suggest that caution is needed when interpreting OHS results, particularly if outcomes are measured and treatment modalities compared 1 year after hip replacement. Reporting pain and function separately is advised, but measurement limitations remain.
