Related Experiment Videos
Empirical error-confidence curves for neural network and Gaussian classifiers
1Machine Learning & Perception Group, Ricoh California Research Center, Menlo Park, CA 94025, USA. wolff@crc.ricoh.com
International Journal of Neural Systems
|July 1, 1996
Summary
This study empirically tests Probably Almost Bayes (PAB) theory for classifier error-confidence. PAB bounds are conservative, and linear predictions align with linear discriminant performance but not neural networks.
Area of Science:
- Machine Learning
- Statistical Learning Theory
- Computational Statistics
Background:
- Error-Confidence (EC) quantifies classifier error probability relative to Bayes error.
- Probably Almost Bayes (PAB) theory models EC increase with training data.
- Understanding classifier performance scaling is crucial for practical applications.
Purpose of the Study:
- Empirically investigate the relationship between training samples and classifier error-confidence.
- Evaluate the predictive accuracy of PAB bounds and asymptotic statistics.
- Compare performance of linear classifiers versus neural networks.
Main Methods:
- Generated Error-Confidence (EC) curves by varying the number of training patterns (m).
- Compared empirical results with theoretical predictions from PAB and asymptotic statistics.
- Utilized linear classifiers and neural networks on Gaussian problems.
Main Results:
- PAB bounds were found to be highly conservative on Gaussian problems.
- Asymptotic statistics predictions showed good agreement with linear discriminant performance at low Bayes error rates.
- A linear relationship between log average error and log training patterns was observed, but with deviations from theory for neural networks and higher error rates.
Conclusions:
- PAB theory provides conservative bounds for classifier error-confidence.
- Linear classifier performance aligns with asymptotic predictions under certain conditions.
- Neural network performance exhibits greater dependence on classifier capacity, challenging simple linear scaling predictions.