Related Experiment Video
Updated: Apr 4, 2026

15:00
A Tablet-Based Curriculum-Based Measurement Protocol for Kindergarten Writing
Published on: February 7, 2025
1.2K
Generalizability theory reliability of written expression curriculum-based measurement in universal screening
Milena A Keller-Margulis1, Sterett H Mercer2, Erin L Thomas1
1Department of Psychological, Health, and Learning Sciences.
Summary
Written expression curriculum-based measurement (WE-CBM) shows low reliability for universal screening. Current methods, using short writing samples, do not provide adequate scores for student performance decisions.
Area of Science:
- Educational Psychology
- Psychometrics
- Academic Assessment
Background:
- Curriculum-based measurement (CBM) is widely used for progress monitoring and screening.
- Written expression CBM (WE-CBM) is a common tool, but its reliability for universal screening needs further examination.
- Generalizability theory provides a framework to assess the sources of measurement error.
Purpose of the Study:
- To evaluate the reliability of written expression curriculum-based measurement (WE-CBM) within a universal screening context.
- To apply generalizability theory to understand the variance components in WE-CBM scores.
- To determine the number and duration of WE-CBM probes needed for reliable student performance decisions.
Main Methods:
- A generalizability theory framework was used to analyze WE-CBM reliability.
- Three WE-CBM probes (7 minutes each) were administered to 145 students in grades 2-5 across three time points over one year.
- Writing samples were scored using metrics like correct minus incorrect word sequences (CIWS).
Main Results:
- Nearly half of the variance in WE-CBM scores was attributed to unsystematic error.
- Conventional screening procedures using a single 3-minute sample yielded inadequate reliability for relative or absolute decisions.
- Multiple probes (e.g., three 3-minute or two longer samples) were needed for reliable relative decisions, with three 7-minute samples insufficient for within-year growth decisions.
Conclusions:
- Current WE-CBM screening practices are unreliable for making critical student performance decisions.
- More extensive or frequent sampling is necessary to achieve adequate reliability in WE-CBM.
- Recommendations are provided for improving the reliability and utility of WE-CBM in educational settings.
Related Concept Videos
Reliability and Validity
14.4K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
14.4K
Measures of Intelligence
9.6K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
9.6K
Wechsler's Contribution to Measures of Intelligence
2.5K
David Wechsler, a psychologist who worked with World War I veterans, developed a significant IQ test in 1939 called the Wechsler-Bellevue Intelligence Scale. This test was innovative because it combined several subtests that measured both verbal and nonverbal skills, reflecting Wechsler's belief that intelligence is a global capacity involving purposeful action, rational thinking, and effective interaction with the environment. This test later evolved into the Wechsler Adult Intelligence...
2.5K
Binet's Contribution to Measures of Intelligence
2.1K
Alfred Binet, along with his student Théophile Simon, was tasked by the French Ministry of Education in 1904 to create a method for identifying students who struggled to learn through conventional classroom instruction. This initiative aimed to address overcrowding by placing such students in specialized schools. Binet and Simon developed an intelligence test comprising 30 tasks, ranging from simple commands, like touching one's nose or ear, to more complex tasks, such as drawing...
2.1K
Self-Report Tests of Personality
1.2K
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
1.2K

