Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

The Bayesian evidence scheme for regularizing probability-density estimating neural networks.

D Husmeier1

  • 1Biomathematics and Statistics Scotland, Scottish Crop Research Institute, Dundee, UK.

Neural Computation
|December 8, 2000
PubMed
Summary

A new regularization method prevents overfitting in neural networks trained with the expectation-maximization (EM) algorithm, especially for sparse data. This Bayesian approach modifies the EM algorithm for improved performance in various applications.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Optimal estimation of drift and diffusion coefficients in the presence of static localization error.

Physical review. E·2019
Same author

Reverse engineering of genetic networks with Bayesian networks.

Biochemical Society transactions·2003
Same author

An empirical evaluation of Bayesian sampling with hybrid Monte Carlo for training neural network classifiers.

Neural networks : the official journal of the International Neural Network Society·2003
Same author

Detection of recombination in DNA multiple alignments with hidden Markov models.

Journal of computational biology : a journal of computational molecular cell biology·2001
Same author

Probabilistic divergence measures for detecting interspecies recombination.

Bioinformatics (Oxford, England)·2001
Same author

Learning non-stationary conditional probability distributions.

Neural networks : the official journal of the International Neural Network Society·2000

Area of Science:

  • Machine Learning
  • Computational Statistics
  • Artificial Intelligence

Background:

  • The expectation-maximization (EM) algorithm is commonly used for training probability-density estimating neural networks.
  • Standard EM training maximizes training set likelihood, leading to overfitting with sparse data.
  • Overfitting compromises the generalization ability of models on unseen data.

Purpose of the Study:

  • To propose a novel regularization method for mixture models trained with neural networks.
  • To address the overfitting issue inherent in the standard expectation-maximization algorithm for sparse datasets.
  • To enhance the robustness and applicability of neural network density estimation models.

Main Methods:

  • A Bayesian evidence approach is employed, optimizing prior hyperparameters via type II maximum likelihood.

Related Experiment Videos

  • Marginalization over model parameters is performed using Laplace approximation.
  • The derivation of the Hessian of the log-likelihood function is a key component.
  • A modified expectation-maximization algorithm incorporating a regularization term is developed.
  • Hyperparameter adaptation is performed online after each EM cycle.
  • Main Results:

    • The proposed method introduces a regularization term into the expectation-maximization algorithm.
    • Hyperparameters are adapted dynamically during the training process.
    • The modified EM algorithm demonstrates improved performance in classification tasks.
    • The scheme is effective for predicting stochastic time series.
    • Applications to latent space models are also presented.

    Conclusions:

    • The proposed regularization method effectively mitigates overfitting in neural network density estimation.
    • The Bayesian approach provides a principled way to regularize mixture models.
    • The modified EM algorithm offers a robust and adaptable training scheme for various machine learning problems.
    • This technique enhances the reliability of models dealing with sparse or complex data distributions.