Related Experiment Video
Updated: Jun 13, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Sample size requirements for training to a kappa agreement criterion on clinical dementia ratings
Rochelle E Tractenberg1, Futoshi Yumoto, Shelia Jin
1Department of Neurology, Georgetown University School of Medicine, Washington, DC 20057, USA. ret7@georgetown.edu
Abstract:
The Clinical Dementia Rating (CDR) is a valid and reliable global measure of dementia severity. Diagnosis and transition across stages hinge on its consistent administration. Reports of CDR ratings reliability have been based on 1 or 2 test cases at each severity level; agreement (kappa) statistics based on so few rated cases have large error, and confidence intervals are incorrect. Simulations varied the number of test cases, and their distribution across CDR stage; to derive the sample size yielding a 95% confidence that estimated is at least 0.60. We found that testing raters on 5 or more patients per CDR level (total N=25) will yield the desired confidence in estimated kappa, and if the test involves greater representation of CDR stages that are harder to evaluate, at least 42 ratings are needed. Testing newly trained raters with at least 5 patients per CDR stage will provide valid estimation of rater consistency, given the point estimate for kappa is roughly 0.80; fewer test cases increases the standard error and unequal distribution of test cases across CDR stages will lower kappa and increase error.
