Related Experiment Video
Updated: Aug 8, 2026

10:32
Development of a Virtual Reality Assessment of Everyday Living Skills
Published on: April 23, 2014
The Structured Clinical Interview for DSM-III-R (SCID). II. Multisite test-retest reliability
J B Williams1, M Gibbon, M B First
1Department of Psychiatry, Columbia University, New York, NY.
Archives of General Psychiatry
|August 11, 1992
Summary
The Structured Clinical Interview for DSM-III-R shows good reliability in patient samples, with kappas above .60. However, reliability is lower in nonpatient samples, suggesting potential areas for improvement in diagnostic consistency.
Area of Science:
- Psychiatry
- Psychological Assessment
- Clinical Psychology
Background:
- The Structured Clinical Interview for DSM-III-R (SCID) is a widely used diagnostic tool.
- Assessing the test-retest reliability of diagnostic instruments is crucial for clinical practice and research.
Purpose of the Study:
- To evaluate the test-retest reliability of the Structured Clinical Interview for DSM-III-R across diverse clinical and nonclinical settings.
Main Methods:
- A large-scale study involving 592 subjects across multiple patient and nonpatient sites in the US and Germany.
- Utilized kappa statistics to measure diagnostic agreement for current and lifetime diagnoses.
Main Results:
- High reliability (kappa > .60) was observed for most major diagnostic categories in patient samples (overall weighted kappa: current .61, lifetime .68).
- Lower reliability was found in nonpatient samples (mean kappa: current .37, lifetime .51).
- Findings are comparable to other structured diagnostic instruments.
Conclusions:
- The SCID demonstrates acceptable to good test-retest reliability in patient populations.
- Lower reliability in nonpatients may be influenced by factors such as interviewer training, information variance, and low disorder prevalence.
- Further research into optimizing diagnostic consistency is warranted.
More Related Videos
Related Concept Videos
Reliability and Validity
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Self-Report Tests of Personality
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.

