Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

439
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
439
Parametric Survival Analysis: Weibull and Exponential Methods01:14

Parametric Survival Analysis: Weibull and Exponential Methods

1.1K
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
1.1K
Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

621
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
621
Significance Testing: Overview01:04

Significance Testing: Overview

12.8K
Significance testing is a set of statistical methods used to test whether a claim about a parameter is valid. In analytical chemistry, significance testing is used primarily to determine whether the difference between two values comes from determinate or random errors. The effect of a particular change in the measurement protocol, analyst, or sample itself can cause a deviation from the expected result. In the case of a suspected deviation/outlier, we need to be able to confirm mathematically...
12.8K
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test01:09

Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test

7.0K
In parametric statistics, two fundamental tests stand out for their utility and wide application: the Student's t-test and goodness-of-fit tests. These tests provide researchers with a robust method for drawing insights from data, testing hypotheses, and making informed decisions based on their findings.
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with...
7.0K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

1.0K
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
1.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Lung Cancer Screening Centralization Is Associated With Improved Screening Uptake: The Veterans Healthcare Administration's Experience 2015-2021.

Chest·2026
Same author

Centralization of Lung Cancer Screening and Adherence to First Follow-up Assessment: The Veterans Health Administration Experience.

Annals of the American Thoracic Society·2026
Same author

Deconstructing triglyceride glucose-body mass index in normal-weight women with polyendocrine metabolic ovarian syndrome.

Fertility and sterility·2026
Same author

Estimating Lung Cancer Screening Eligibility in the Veterans Health Administration Using Patient-Reported Smoking Histories.

Journal of general internal medicine·2026
Same author

PSMA PET/CT-Derived Indicators and Outcomes After [<sup>177</sup>Lu]Lu-PSMA-617: A Multicenter Retrospective Analysis from the U.S. Expanded-Access Program.

Journal of nuclear medicine : official publication, Society of Nuclear Medicine·2026
Same author

Phase 2 Prospective Trial of Retreatment with [<sup>177</sup>Lu]Lu-PSMA-617 Molecular Radiotherapy for Metastatic Castration-Resistant Prostate Cancer-RE-LuPSMA.

Journal of nuclear medicine : official publication, Society of Nuclear Medicine·2026

Related Experiment Video

Updated: Feb 17, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

768

A simulation based method for assessing the statistical significance of logistic regression models after common

Tristan R Grogan1, David A Elashoff1

  • 1Department of Medicine Statistics Core, University of California, Los Angeles, CA.

Communications in Statistics: Simulation and Computation
|December 12, 2017
PubMed
Summary

Classification models may appear accurate due to chance, not true relationships. Variable selection methods can yield false positives, overestimating performance. Critical values help identify genuine associations.

Keywords:
AUCLogistic RegressionSimulation StudyValidation methodsVariable selection

More Related Videos

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.6K
Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.9K

Related Experiment Videos

Last Updated: Feb 17, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

768
An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.6K
Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.9K

Area of Science:

  • Statistics
  • Machine Learning
  • Biostatistics

Background:

  • Classification models can exhibit high prediction accuracy without a true relationship between predictors and outcomes.
  • Variable selection procedures are prone to false positives and overestimating model performance.

Purpose of the Study:

  • To investigate the impact of variable selection methods on classification model performance.
  • To identify strategies for reducing false positive variable selections and overestimation of model performance.

Main Methods:

  • A simulation study was conducted using logistic regression.
  • Evaluated forward stepwise, best subsets, and LASSO variable selection methods.
  • Varied total sample sizes (20-200) and numbers of noise predictor variables (3-50).

Main Results:

  • Apparent prediction accuracy can occur without genuine predictor-response relationships.
  • Variable selection methods can lead to false positives and inflated performance metrics.
  • The study identified critical values to help mitigate these issues.

Conclusions:

  • Standard variable selection methods can be unreliable in the presence of noise variables.
  • Proposed critical values can aid in distinguishing true associations from spurious findings.
  • Emphasizes the need for cautious interpretation of model performance and variable importance.