Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Prediction Intervals01:03

Prediction Intervals

3.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
3.5K
Data Validation01:03

Data Validation

7.2K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
7.2K
Data Validation01:15

Data Validation

3.4K
Method validation is a crucial process in analytical chemistry designed to confirm that a given method consistently produces reliable and high-quality results. This process is essential when a method is applied to different sample matrices or when procedural modifications are made, ensuring that the results meet acceptable standards across various applications.
Key parameters for method validation include:
3.4K
Multiple Regression01:25

Multiple Regression

4.2K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
4.2K
Sensitivity, Specificity, and Predicted Value01:13

Sensitivity, Specificity, and Predicted Value

1.6K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
1.6K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

9.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
9.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

When to Adjust for Multiple Testing: A Unifying Guiding Principle.

Biometrical journal. Biometrische Zeitschrift·2026
Same author

STrategies for developing REseArch Methods guidance (STREAM): Protocol.

Journal of clinical epidemiology·2026
Same author

Longitudinal associations of long-term exposure to ambient air pollution, residential greenness, and air temperature with type 2 diabetes subphenotypes: Results from the KORA cohort study.

Environment international·2026
Same author

The statistical software revolution in pharmaceutical development: challenges and opportunities in open source.

Drug discovery today·2026
Same author

Integrative Metabolomics of Targeted and Nontargeted Analyses in T2D Progression.

Diabetes care·2025
Same author

Accelerometry-assessed sleep and liver health in adolescents and adults: Links to liver enzymes, MASLD, and MRI-derived liver fat.

Sleep medicine·2025

Related Experiment Video

Updated: Mar 13, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.7K

Assessment of predictive performance in incomplete data by combining internal validation and multiple imputation.

Simone Wahl1,2,3, Anne-Laure Boulesteix4, Astrid Zierer5

  • 1Research Unit of Molecular Epidemiology, Helmholtz Zentrum München - German Research Center for Environmental Health, Ingolstädter Landstrasse, Neuherberg, 1, 85764, Germany. simwahl@googlemail.com.

BMC Medical Research Methodology
|October 27, 2016
PubMed
Summary

For studies with missing data, the Val-MI strategy effectively corrects for optimism in predictive performance estimates. This method, combining internal validation before multiple imputation, provides largely unbiased results for prognostic models.

Keywords:
BootstrapCross-validationIncomplete dataInternal validationMICEMissing valuesMultiple imputationPrediction modelPredictive performanceResampling

More Related Videos

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

8.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

842

Related Experiment Videos

Last Updated: Mar 13, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.7K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

8.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

842

Area of Science:

  • Biostatistics
  • Epidemiology
  • Medical Informatics

Background:

  • Missing values are common in human studies, necessitating robust handling strategies.
  • Multiple imputation (MI) is a standard technique for addressing missing data by imputing values multiple times and pooling results.
  • Estimating predictive performance measures, like the area under the receiver-operating characteristic curve (AUC), requires internal validation to correct for optimism, but its combination with MI is not well-understood.

Purpose of the Study:

  • To compare three strategies for combining internal validation with multiple imputation (MI) for estimating predictive performance measures in studies with missing data.
  • To evaluate the bias and mean squared error of performance estimates obtained from different combination strategies.
  • To provide guidance on the number of resamples and imputations and on constructing confidence intervals for incomplete data.

Main Methods:

  • A comprehensive simulation study and a real-world dataset (blood markers predicting mortality) were used for comparison.
  • Three strategies were evaluated: Val-MI (validation then MI), MI-Val (MI then validation), and MI(-y)-Val (MI omitting outcome, then validation).
  • Various validation techniques (e.g., bootstrap, cross-validation), performance measures, and data characteristics were considered.

Main Results:

  • MI-Val estimates were optimistically biased, while MI(-y)-Val estimates tended to be pessimistic.
  • The Val-MI strategy yielded largely unbiased estimates, with only slight pessimism under certain conditions (large effect size, many covariates, small sample size).
  • Increasing bootstrap draws improved Val-MI accuracy more than increasing imputations; a simple approach for valid confidence intervals was established.

Conclusions:

  • The Val-MI strategy is a valid approach for obtaining reliable estimates of predictive performance measures when developing prognostic models on incomplete data.
  • This strategy effectively corrects for optimism, offering a practical solution for handling missing data in model performance evaluation.
  • The findings support the use of Val-MI for accurate assessment of predictive accuracy in clinical and epidemiological research.