Related Experiment Video
Updated: Mar 26, 2026

12:27
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
7.4K
Cautionary Remarks on the Use of Clusterwise Regression
Michael J Brusco1, J Dennis Cradit2, Douglas Steinley3
1a Florida State University .
Multivariate Behavioral Research
|January 21, 2016
Summary
Clusterwise regression may overfit data by attributing variation to clustering rather than the regression model itself. A benchmarking procedure is recommended to validate results and prevent misuse of this statistical method.
Area of Science:
- Multivariate statistics
- Statistical modeling
Background:
- Clusterwise linear regression aims to cluster objects while minimizing within-cluster regression errors.
- The standard objective function conflates error from clustering and regression.
Purpose of the Study:
- To demonstrate the potential for overfitting in clusterwise regression.
- To propose a benchmarking procedure to mitigate overfitting risks.
Main Methods:
- Analysis of the clusterwise regression objective function.
- Numerical examples and simulation experiments.
- Empirical application predicting reflective judgment.
Main Results:
- The objective function does not distinguish between clustering and regression error.
- Clustering can explain most variation, leaving little for regression models.
- Overfitting is demonstrated through examples and simulations.
Conclusions:
- Clusterwise regression has a high potential for overfitting.
- A benchmarking procedure comparing empirical data to random permutations is recommended.
- Careful validation is needed to prevent misuse of clusterwise regression.
Related Concept Videos
Cluster Sampling Method
15.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.5K
Multiple Regression
4.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.3K
Regression Analysis
8.9K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.9K
Regression Toward the Mean
7.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.3K
Correlation and Regression
4.1K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
4.1K
Outliers and Influential Points
6.7K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.7K

