Related Experiment Video
Updated: Jun 16, 2026

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
Best practices in multilevel modeling for within-cluster group comparisons: An evaluation of coding strategies
Qian Zhang1, Xiao Liu2, Zijun Ke3
1Department of Educational Psychology and Learning Systems, Anne Spencer Daves College of Education, Health, and Human Sciences, Florida State University.
Abstract:
In multilevel analysis, Level-1 dichotomous predictors (e.g., treatment vs. control, minority vs. majority; D) are commonly used to examine the existence of average within-cluster mean differences. To address this research question, we examine the best practices for coding D, considering within-cluster group compositions and between-cluster heterogeneity in them, as well as possible between-group heterogeneity in residual variances. We compare traditional dummy coding and unweighted effect coding with two proposed methods: overall and within-cluster weighted effect coding. Each is combined with cluster-mean centering to clearly define the within-cluster mean difference. Analytical results show that the intraclass correlation of D (ICC-D) substantially influences differences across coding methods regarding test results for the average within-cluster effect (group mean difference). When ICC-D is zero, all strategies yield identical test results (i.e., the significance tests based on the point estimates and standard error estimates) for the average within-cluster effect. However, as the ICC-D increases, test statistics (e.g., the t statistic) decrease for all coding strategies. An empirical study and simulations further confirm that ICC-D strongly influences results. Based on the simulations, we recommend dummy coding, unweighted effect coding, and overall weighted effect coding, which generally provide comparable or higher power than within-cluster weighted effect coding. The power contrast is most pronounced when both ICC-D and Level-1 residual variance heterogeneity increase, with the difference in power reaching up to 40%. Furthermore, cluster-mean centering not only disaggregates within- and between-cluster effects but also achieves generally reasonable Type I error rates even when homogeneity assumptions are not met. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
More Related Videos
08:27Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Group Design
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Friedman Two-way Analysis of Variance by Ranks
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...