Related Experiment Video
Updated: Sep 17, 2026

Development of a Virtual Reality Assessment of Everyday Living Skills
Published on: April 23, 2014
Interrater reliability issues in multicenter trials, Part I: Theoretical concepts and operational procedures used in
K Tracy1, L A Adler, J Rotrosen
1Psychiatry Service, VA Medical Center/NYU School of Medicine, NY 10010, USA.
Abstract:
This article describes a standardized method for establishing and maintaining desired levels of interrater reliability (IRR) in multicenter trials. The procedure involves six steps: distribution of procedural guides, distribution of an introduction tape, initial distribution of patient interviews to rate, training at the study kickoff meeting, ongoing IRR monitoring, and group training throughout the study. This method is being used in a national Veterans Affairs Cooperative Study (CS #394), involving nine sites to examine the treatment effects of vitamin E on tardive dyskinesia. The six-step standardized process allowed for early detection of areas of concern in assessment administration. When comparing intraclass correlation coefficients (ICCs) at different points in the initial training, the Barnes Akathisia Scale and Anchored Brief Psychiatric Rating Scale reliability improved from 0.68 to 0.74 and from 0.54 to 0.87, respectively. After analyzing the ratings collected prior to the start of CS #394, data were collected to conduct the first check on Abnormal Involuntary Movement Scale (AIMS) IRR during enrollment; the estimated ICC for the AIMS had decreased from 0.87 to 0.60. Raters were instructed to re-assess the subjects from the first videotape on the AIMS and received additional training. The re-rating indicated very good reliability, 0.84, IRR was measured once for the Global Assessment of Functioning Scale resulting in an ICC of 0.90. The companion article (Part II: Edson et al. 1997, page 59 of this issue) describes the statistical procedures used to measure IRR.
More Related Videos
Related Concept Videos
Blind Procedures
Reliability and Validity
Crossover Experiments
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Bioequivalence studies: Biowaivers
Drug Accumulation During Multiple Dosing: Repetitive IV Injections

