Related Experiment Video
Updated: Jun 23, 2025

Studying Murine Small Bowel Mechanosensing of Luminal Particulates
Published on: March 18, 2022
Distinguishing subsampled power laws from other heavy-tailed distributions
Silja Sormunen1, Lasse Leskelä2, Jari Saramäki1
1Department of Computer Science, Aalto University, 00076 Espoo, Finland.
Detecting power-law distributions is hard, especially with subsampling. This study found that while power-law exponents are estimable from subsamples, correctly classifying the distribution remains challenging for common methods.
Area of Science:
- Statistical analysis
- Data science
- Complex systems
Background:
- Distinguishing power-law distributions from other heavy-tailed distributions (lognormal, stretched exponential) is crucial in many scientific fields.
- Subsampling effects, common in network and biological sciences, can further complicate this detection process.
Purpose of the Study:
- To evaluate the performance of two common methods (Clauset et al.'s maximum likelihood and Voitalov et al.'s extreme value) in identifying power-law distributions from subsampled data.
- To assess how well these methods distinguish subsampled power laws from lognormal and stretched exponential distributions under random subsampling.
Main Methods:
- The study employed a random subsampling method, focusing on frequency distributions of elements (e.g., species, network nodes).
- Two established statistical methods were tested: maximum likelihood estimation and extreme value theory.
- The performance was evaluated across various subsampling depths to assess generalization to original distributions.
Main Results:
- Power-law exponents can be estimated with reasonable accuracy from subsamples, but correct distribution classification is more difficult.
- The maximum likelihood method frequently misclassifies subsampled power laws, rejecting the true hypothesis.
- The extreme value method shows better performance with subsampled power laws but struggles to differentiate them from other heavy-tailed distributions, with limitations often stemming from original sample classification.
Conclusions:
- Subsampling complicates the accurate identification of power-law distributions, impacting the reliability of common detection methods.
- The extreme value method demonstrates resilience to subsampling effects for power-law detection compared to the maximum likelihood method.
- Further research is needed to improve classification accuracy, as current methods show limitations in distinguishing between heavy-tailed distributions, particularly with subsampled data.
Related Concept Videos
Probability Distributions
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Quantifying and Rejecting Outliers: The Grubbs Test
Distributions to Estimate Population Parameter
Sampling Distribution
Outliers and Influential Points

