Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reliability and Validity01:29

Reliability and Validity

Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
Accuracy and Precision01:52

Accuracy and Precision

Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.  Highly accurate measurements...
Accuracy and Precision01:52

Accuracy and Precision

Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.  Highly accurate measurements...
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Sensitivity, Specificity, and Predicted Value01:13

Sensitivity, Specificity, and Predicted Value

In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Survival Tree01:19

Survival Tree

Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a survival tree begins...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Comments on Xu et al. (2026) "Development of a microscopy-based diagnostic test for alkali injury-induced limbal stem cell deficiency through autofluorescence multispectral imaging" (Exp. Eye Res. 264:110829).

Experimental eye research·2026
Same author

Interpretable integration of SEM and SVM for reliable thyroid nodule classification.

Artificial intelligence in medicine·2026
Same author

Reconsidering regression-based predictors in adults with attention‑deficit/hyperactivity disorder and substance use disorders.

European neuropsychopharmacology : the journal of the European College of Neuropsychopharmacology·2026
Same author

Beyond Parametric Assumptions: Reevaluation of Environmental Interactions in Prenatal Exposure Studies.

Chest·2026
Same author

Correspondence on "Clinical correlation between metabolic biomarkers and chemoresistance in gestational trophoblastic neoplasia" by Kong et al.

International journal of gynecological cancer : official journal of the International Gynecological Cancer Society·2026
Same author

Evaluating Linear Parametric Poisson Regression vs Nonparametric Unsupervised Learning in Ulcerative Colitis Data.

Gastroenterology·2026

Related Experiment Video

Updated: Jul 21, 2026

Measurement of Spatial Stability in Precision Grip
09:36

Measurement of Spatial Stability in Precision Grip

Published on: June 4, 2020

The reliability gap: Why high predictive accuracy doesn't guarantee stable feature importance.

Yoshiyasu Takefuji1

  • 1Faculty of Data Science, Musashino University, 3-3-3 Ariake Koto-ku, Tokyo, 135-8181, Japan.

Marine Pollution Bulletin
|February 8, 2026
PubMed
Summary

Machine learning for environmental data can produce unstable feature importances. Unsupervised methods offer stable, reliable insights into pollutant and shellfish poisoning risk, validating environmental ML practices.

Keywords:
Environmental risk predictionFeature importance stabilitySHAP interpretationSupervised learning limitationsUnsupervised feature selection

More Related Videos

Doppler Ultrasound-Based Leg Blood Flow Assessment During Single-Leg Knee-Extensor Exercise in an Uncontrolled Setting
09:18

Doppler Ultrasound-Based Leg Blood Flow Assessment During Single-Leg Knee-Extensor Exercise in an Uncontrolled Setting

Published on: December 15, 2023

Troubleshooting and Quality Assurance in Hyperpolarized Xenon Magnetic Resonance Imaging: Tools for High-Quality Image Acquisition
09:55

Troubleshooting and Quality Assurance in Hyperpolarized Xenon Magnetic Resonance Imaging: Tools for High-Quality Image Acquisition

Published on: January 5, 2024

Related Experiment Videos

Last Updated: Jul 21, 2026

Measurement of Spatial Stability in Precision Grip
09:36

Measurement of Spatial Stability in Precision Grip

Published on: June 4, 2020

Doppler Ultrasound-Based Leg Blood Flow Assessment During Single-Leg Knee-Extensor Exercise in an Uncontrolled Setting
09:18

Doppler Ultrasound-Based Leg Blood Flow Assessment During Single-Leg Knee-Extensor Exercise in an Uncontrolled Setting

Published on: December 15, 2023

Troubleshooting and Quality Assurance in Hyperpolarized Xenon Magnetic Resonance Imaging: Tools for High-Quality Image Acquisition
09:55

Troubleshooting and Quality Assurance in Hyperpolarized Xenon Magnetic Resonance Imaging: Tools for High-Quality Image Acquisition

Published on: January 5, 2024

Area of Science:

  • Environmental Science
  • Data Science
  • Marine Biology

Background:

  • Machine learning (ML) and explainable AI (XAI) are increasingly used for environmental risk assessment, such as pollutant analysis and shellfish poisoning.
  • Techniques like Principal Component Analysis (PCA) and SHapley Additive exPlanations (SHAP) are common, but linear PCA may fail with nonlinear environmental data, and feature importances are often treated as ground truth without validation.

Purpose of the Study:

  • To critically evaluate the reliability of feature importances derived from supervised ML models in environmental studies.
  • To introduce and validate methods for assessing the stability and consistency of ML-derived feature rankings.
  • To compare the performance and stability of supervised vs. unsupervised ML approaches for environmental risk prediction.

Main Methods:

  • Utilized a Basque coastal dataset (8195 instances, 14 features) with chlorophyll-a as a proxy for paralytic shellfish poisoning risk.
  • Implemented a leave-top1-out cross-validation procedure to assess the stability of feature rankings.
  • Compared supervised models (Random Forest, XGBoost with/without SHAP) against unsupervised and non-target-prediction methods.

Main Results:

  • Supervised models (Random Forest, XGBoost) exhibited significant instability in feature importance rankings, suggesting model-dependent biases.
  • Unsupervised and non-target-prediction methods demonstrated perfect ranking stability.
  • These stable methods matched or exceeded the predictive performance of supervised models.

Conclusions:

  • Feature importance from supervised ML models in environmental science should be interpreted with caution due to potential instability and bias.
  • Unsupervised and non-target-prediction methods offer more robust and stable insights for environmental risk assessment.
  • Routine checks for stability, consistency, dose-response relationships, and linearity are crucial for reliable environmental ML studies.