Related Experiment Video
Updated: Aug 6, 2026

12:27
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Clustering ensembles of neural network models
1SNN, University of Nijmegen, Geert Grooteplein 21, 6525 EZ Nijmegen, The Netherlands. bartb@mbfys.kun.nl
Summary
Large ensembles of neural network models can be summarized by a few representative models. This model summarization, using clustering, often matches or exceeds full ensemble performance for better function estimation.
Area of Science:
- Machine Learning
- Computational Statistics
Background:
- Large ensembles of neural network models are common in machine learning.
- Techniques like bootstrapping and Bayesian sampling generate these extensive model collections.
- Storing and analyzing these large ensembles can be computationally intensive.
Purpose of the Study:
- To develop an effective method for summarizing large neural network model ensembles.
- To investigate if a smaller set of representative models can achieve comparable or superior performance.
- To enhance the qualitative analysis and insight generation from model ensembles.
Main Methods:
- A novel method for identifying representative models using clustering based on model outputs.
- Application of the clustering method to neural network ensembles generated via bootstrapping.
- Evaluation of ensemble performance using Boston housing data and newspaper sales prediction tasks.
Main Results:
- A small subset of representative models can effectively summarize large ensembles.
- The summarized ensembles often match or surpass the performance of the full ensemble.
- Clustering provides a more manageable representation for qualitative analysis and data insights.
Conclusions:
- It is not necessary to store or process all models in an ensemble.
- Representative model selection via clustering offers computational and analytical advantages.
- This approach yields new insights into data and model behavior, improving function estimation.

