Related Experiment Video
Updated: Aug 5, 2026

Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
Mechanisms of Effect Size Differences Between Researcher Developed and Independently Developed Outcomes: An
Joshua B Gilbert1, James Soland2
1Center for School Behavioral Health, Massachusetts General Hospital, Boston, Massachusetts; United States.
Researcher-developed (RD) and independently developed (ID) measures show different effect sizes in education research. Greater item-level variability in treatment effects explains some, but not all, of this observed difference.
Area of Science:
- Education Research
- Psychometrics
- Meta-Analysis
Background:
- Differences in effect sizes between researcher-developed (RD) and independently developed (ID) outcome measures are documented in education research.
- The underlying mechanisms explaining these differences remain poorly understood.
Purpose of the Study:
- To conduct a meta-analysis using item-level outcome data.
- To test potential mechanisms explaining effect size differences between RD and ID outcome measures.
Main Methods:
- Meta-analysis of 54 effect sizes from 32 studies.
- Analysis of item-level outcome data to examine predictors of effect sizes.
- Investigation of the role of standard deviations of item-specific treatment effects.
Main Results:
- Greater standard deviations of item-specific treatment effects predict larger effect sizes.
- This variability reduced the observed difference between RD and ID measures from 0.14 SDs to 0.12 SDs.
- The remaining difference is attributable to factors other than item-level variability.
Conclusions:
- Item properties, specifically variability in item-specific treatment effects, predict educational intervention outcomes.
- Understanding item-level data is crucial for developing theory in education research.
- Further research is needed to identify other factors contributing to the RD-ID effect size difference.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Group Design
Bioequivalence Experimental Study Designs: Repeated Measures, Cross-Over, Carry-Over, and Latin Square Designs
Methods of Medium Optimization
Comparing Experimental Results: Student's t-Test
Regression Toward the Mean