Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

193
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
193
Biostatistics: Overview01:20

Biostatistics: Overview

355
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
355
Variability: Analysis01:11

Variability: Analysis

182
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
182
Survival Tree01:19

Survival Tree

140
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
140
Types of Selection01:46

Types of Selection

41.3K
Natural selection influences the frequencies of particular alleles and phenotypes within populations in several different ways. Primarily, natural selection can be directional, stabilizing, or disruptive. Directional selection favors one extreme trait and shifts the population towards that phenotype while selecting against individuals displaying alternate traits. Stabilizing selection favors an intermediate trait with a narrow range of variation. Deviation from the optimal phenotype towards an...
41.3K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

1.8K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
1.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Searching for improvements in predicting human eye colour from DNA.

International journal of legal medicine·2021
Same author

Selection Consistency of Lasso-Based Procedures for Misspecified High-Dimensional Binary Model and Random Regressors.

Entropy (Basel, Switzerland)·2020
Same author

Analysis of Information-Based Nonparametric Variable Selection Criteria.

Entropy (Basel, Switzerland)·2020
Same author

A deeper look at two concepts of measuring gene-gene interactions: logistic regression and interaction information revisited.

Genetic epidemiology·2017

Related Experiment Video

Updated: Aug 30, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Information Theoretic Methods for Variable Selection-A Review.

Jan Mielniczuk1,2

  • 1Institute of Computer Science, Polish Academy of Sciences, Jana Kazimierza 5, 01-248 Warsaw, Poland.

Entropy (Basel, Switzerland)
|August 26, 2022
PubMed
Summary

This study enhances information theoretic tools for feature selection in classification. New methods improve conditional mutual information estimation for high-dimensional data, boosting predictive model accuracy.

Keywords:
Markov blanketMöbius expansionconditional independencefeature selectioninteraction information

More Related Videos

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.3K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Related Experiment Videos

Last Updated: Aug 30, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.3K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.6K

Area of Science:

  • Computer Science
  • Machine Learning
  • Information Theory

Background:

  • Feature selection is crucial for classification, especially with discrete features.
  • Empirical conditional mutual information struggles with high-dimensional datasets.
  • Existing methods for estimating conditional dependence have limitations.

Purpose of the Study:

  • To review and enhance information theoretic tools for feature selection in classification.
  • To address the poor performance of empirical conditional mutual information in high-dimensional settings.
  • To introduce novel, robust measures of conditional dependence.

Main Methods:

  • Review of principal information theoretic tools for feature selection.
  • Development of unified methods for constructing conditional mutual information counterparts using truncation and weighing of the Möbius expansion.
  • Application of these measures to feature selection and assessment of predictor vector quality.

Main Results:

  • A unified framework for constructing robust conditional mutual information estimators.
  • Improved methods for feature selection in high-dimensional classification problems.
  • Discussion of recent advances in assessing feature selection quality, including asymptotic distributions and resampling techniques.

Conclusions:

  • The proposed information theoretic methods offer improved feature selection for high-dimensional classification.
  • These advancements provide more reliable ways to assess the quality of selected features.
  • The work contributes to more effective and accurate predictive modeling in machine learning.