Categorical Variables in Multiple Regression: Some Cautions
Abstract:
Various methods of coding categorical variables for use as predictors in multiple regression analyses have been presented in the literature. The limitations of two oft-discussed methods -- dummy coding and nonsense coding -- are detailed for several frequently-used regression designs. Two examples involving potentially inappropriate interpretations of the results of analyses involving dummy coding are presented. Researchers are cautioned that the parameter estimate or estimates and test of significance associated with a predictor variable or set of predictor variables in an equation which involves dummy- and/or nonsense-coded predictors represent an effect of interest only in a limited set of circumstances.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Bar Graph
Each bar or column in the multiple bar graph represents a data value. These graphs are used primarily in interrelating two or more sets of data. The categories of different kinds of data are listed along the horizontal or x-axis, whereas...
Two-Way ANOVA
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the...
Multiple Allele Traits


