Related Experiment Video
Updated: Mar 3, 2026

A Protocol of Manual Tests to Measure Sensation and Pain in Humans
Published on: December 19, 2016
Inter-rater reliability and usability of CATHIS core for homeopathic intervention studies
Martin Loef1, Petra Weiermayer2, Katharina Gaertner3
1Institute of Integrative Medicine, University of Witten/Herdecke, Germany; Society for Clinical Research, Berlin, Germany.
Background:
The Critical Appraisal Tool for Homeopathic Intervention Studies (CATHIS) core is a streamlined appraisal tool for homeopathic intervention studies focusing on credibility, coherence, and clinical relevance. The aim of the research project was to evaluate its inter-rater reliability, feasibility, and face validity.
Methods:
In a preregistered cross-sectional study, four raters independently applied CATHIS core to 28 trials (21 randomised controlled trials, 7 non-randomised studies on interventions) drawn from reviews on insomnia and hypertension; two external reviewers provided consensus ratings. Inter-rater reliability (IRR) was estimated using percent agreement, Fleiss' κ, and Gwet's AC2 (95% CIs). Feasibility was quantified as rating time and consensus time. Associations among the three domains were explored with correlation analyses and sensitivity checks.
Results:
IRR varied markedly by domain. Credibility showed good agreement (Fleiss' κ=0.66, 95% CI 0.57-0.74; AC2 =0.76, 0.71-0.82). Coherence yielded only poor-to-fair agreement (κ=0.28, 0.16-0.40; AC2 =0.41, 0.30-0.51). Clinical relevance was similarly limited (κ=0.32, 0.23-0.41; AC2 =0.36, 0.28-0.44). Individual ratings required on average 65.8 min, while consensus discussions averaged 17.7 min. Correlation analyses indicated heterogeneous and partly overlapping domain signals with limited interpretability. Face-validity responses reflected moderate-to-high acceptance but difficulties in consistent application.
Conclusion:
CATHIS core yielded reproducible credibility ratings but only fair and operationally fragile agreement for coherence and clinical relevance, alongside non-trivial rating burden. Taken together, the current reliability profile is insufficient for confident use in systematic reviews. Targeted refinement appears warranted before broader implementation.
More Related Videos
08:40Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
Published on: June 12, 2019
07:36Analysis of Raw and Processed Cyperi Rhizoma Samples Using Liquid Chromatography-Tandem Mass Spectrometry in Rats with Primary Dysmenorrhea
Published on: December 23, 2022
Related Concept Videos
Reliability and Validity
Bioavailability Study Design: Healthy Subjects Versus Patients
Drug Product Performance: In Vitro–In Vivo Correlation
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Study Design in Statistics
Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments with proper...
Pharmaceutical Alternatives: Stability-Related Therapeutic Nonequivalence