Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Introduction to Test of Independence01:21

Introduction to Test of Independence

2.1K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.1K
Determination of Expected Frequency01:08

Determination of Expected Frequency

1.7K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
1.7K
Hypothesis Test for Test of Independence01:16

Hypothesis Test for Test of Independence

6.3K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
6.3K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

7.0K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
7.0K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

4.0K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
4.0K
Significance Testing: Overview01:04

Significance Testing: Overview

10.2K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
10.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Federated feature selection with false discovery rate control.

Journal of the Royal Statistical Society. Series B, Statistical methodology·2026
Same author

Sorafenib Restores Pentose Phosphate Pathway-Related Redox Homeostasis via the c-Raf/HSP90/G6PD Axis in Hepatic Ischemia-Reperfusion Injury.

MedComm·2026
Same author

The dissemination of a broad-host-range ARG-carrying plasmid to putative pathogens across agricultural soils.

Environmental pollution (Barking, Essex : 1987)·2026
Same author

Cadmium Stress Favours Biofilm Cooperation and Polysaccharide-Enriched Matrix Remodelling in Bacterial Consortia.

Environmental microbiology·2026
Same author

Developing an Oxygen-17 Isotope-Coupled WRF-Chem Model for Elucidating Sulfate Formation Mechanisms in China Haze and Beyond: Part I. Model Description and Initial Assessments.

Environmental science & technology·2026
Same author

Variable selection in functional linear Cox model.

Biometrics·2026

Related Experiment Video

Updated: May 4, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K

MARGINAL EMPIRICAL LIKELIHOOD AND SURE INDEPENDENCE FEATURE SCREENING.

Jinyuan Chang1, Cheng Yong Tang1, Yichao Wu1

  • 1Peking University, University of Colorado Denver and North Carolina State University.

Annals of Statistics
|January 14, 2014
PubMed
Summary

This study introduces a novel marginal empirical likelihood method for feature screening in high-dimensional data. It effectively identifies significant variables by assessing their contribution to response variables, improving upon existing methods.

Keywords:
Empirical likelihoodhigh-dimensional data analysislarge deviationsure independence screening

More Related Videos

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.0K
Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

2.9K

Related Experiment Videos

Last Updated: May 4, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.0K
Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
04:35

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach

Published on: July 3, 2020

2.9K

Area of Science:

  • Statistics
  • Econometrics
  • Machine Learning

Background:

  • High-dimensional data analysis presents challenges for traditional statistical methods.
  • Existing feature screening methods often rely solely on estimator magnitudes, potentially overlooking variable uncertainty.

Purpose of the Study:

  • To develop a unified feature screening procedure for linear and generalized linear models.
  • To propose a method that incorporates estimator uncertainty for more robust variable selection.

Main Methods:

  • Utilizing a marginal empirical likelihood approach to analyze variable contributions.
  • Examining marginal empirical likelihood ratios to differentiate significant explanatory variables.
  • Developing a unified screening procedure for linear and generalized linear models.

Main Results:

  • The marginal empirical likelihood ratio evaluated at zero effectively distinguishes contributing explanatory variables.
  • The proposed method incorporates estimator uncertainty, enhancing robustness.
  • The approach is less restrictive on distributional assumptions and adaptable to various models.

Conclusions:

  • The marginal empirical likelihood approach offers a powerful and flexible tool for feature screening in high-dimensional settings.
  • This method provides a valuable extension to existing feature screening techniques.
  • The approach demonstrates strong performance through theoretical analysis and empirical validation.