Jove
Visualize
Contact Us

Related Concept Videos

Bootstrapping01:24

Bootstrapping

792
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
792
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

3.5K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Biostatistics: Overview01:20

Biostatistics: Overview

705
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
705
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
1.1K
Frequency-dependent Selection01:21

Frequency-dependent Selection

23.0K
When the fitness of a trait is influenced by how common it is (i.e., its frequency) relative to different traits within a population, this is referred to as frequency-dependent selection. Frequency-dependent selection may occur between species or within a single species. This type of selection can either be positive—with more common phenotypes having higher fitness—or negative, with rarer phenotypes conferring increased fitness.
23.0K
Prediction Intervals01:03

Prediction Intervals

3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
3.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A single GPCR locus in Drosophila melanogaster partitions stress physiology by sex.

Comparative biochemistry and physiology. Part A, Molecular & integrative physiology·2026
Same author

Environmental microbes as modulators of plant volatile landscapes: Implications for plant-insect chemical communication.

Trends in microbiology·2026
Same author

Microbiome-mediated chemical communication in insects: Implications for pest management.

Pest management science·2026
Same author

Rethinking the role of surgery and radiotherapy in BRAF-mutant papillary craniopharyngioma: a position statement from the neurosurgical perspective.

Neurosurgical focus·2026
Same author

Evolution, multifunctionality, and agricultural potential of insect microbiomes and the holobiont concept.

The ISME journal·2026
Same author

Climatic Niche Contraction and Refugial Persistence of an Invasive Tephritid Pest Across the Arabian Peninsula Under Contrasting Emission Scenarios.

Biology·2026
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jan 10, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K

Gradient boosting with knockoff filters: a biostatistical approach to variable selection.

Amr Mohamed1, Kevin H Lee2

  • 1College of Health Sciences, The University of Memphis, 3720 Alumni Ave, Memphis, TN, 38152, USA. A.Mohamed@memphis.edu.

BMC Bioinformatics
|November 26, 2025
PubMed
Summary

This study introduces a novel variable selection method integrating knockoffs with Light Gradient Boosting Machine (LightGBM) and SHAP values. The approach efficiently identifies significant variables while controlling False Discovery Rate (FDR), outperforming traditional methods in big data scenarios.

Keywords:
Gradient boostingKnockoffsLightGBMSHAPVariable selection

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.6K

Related Experiment Videos

Last Updated: Jan 10, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

1.2K
Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients
07:34

Probing the Limits of Egg Recognition Using Egg Rejection Experiments Along Phenotypic Gradients

Published on: August 22, 2018

8.6K

Area of Science:

  • Statistics
  • Machine Learning
  • Data Science

Background:

  • Increasing data complexity necessitates efficient variable selection methods.
  • Controlling False Discovery Rate (FDR) and maintaining statistical power are key challenges.
  • Knockoff filters offer a robust approach by creating negative controls for inference.

Purpose of the Study:

  • To extend the application of knockoff filters to Light Gradient Boosting Machine (LightGBM).
  • To enhance variable selection accuracy and efficiency in the context of big data.
  • To improve the interpretability of machine learning models using SHAP values.

Main Methods:

  • Integration of knockoff variable generation with LightGBM.
  • Utilization of Shapely Additive Explanations (SHAP) for model interpretability.
  • Extensive experimentation and simulation studies for validation.

Main Results:

  • The proposed method accurately identifies important variables for each class.
  • Demonstrated superior performance, speed, and efficiency compared to traditional methods.
  • Enhanced interpretability of variable importance through SHAP values.

Conclusions:

  • The integration of knockoffs into LightGBM provides a powerful tool for variable selection.
  • This approach effectively addresses challenges in big data analysis.
  • The method advances statistical modeling and machine learning applications by improving performance and interpretability.