Related Experiment Video
Updated: Feb 7, 2026

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Bridging the Gap Between Data Reproduction and Prediction: The Impact of Feature Selection and Cross-Validation
Milad Saeedi1, Jad Zalzal1, Arman Ganji1
1Department of Civil and Mineral Engineering, University of Toronto, Toronto, ON, M5S 1A4 Canada.
Abstract:
Reliable exposure assessment is vital for epidemiological research, but weaknesses in land-use regression (LUR) models undermine its validity. Using mobile ultrafine particle (UFP) data in Toronto, we compared LUR models trained under random, spatial, temporal, and spatiotemporal cross-validation (CV), with and without forward feature selection (FFS). Model hyperparameters and feature subsets were optimized within each CV scheme. Spatial CV folds were designed at fine scales to reflect UFP autocorrelation. Each approach was evaluated on a hold-out test set, across CV schemes, and against independent stationary backyard measurements. Models based on spatiotemporal CV coupled with FFS were able to reduce overfitting, improve generalization, and produce stable exposure surfaces. These surfaces avoided the spatial artifacts and exaggerated variable effects typically seen in models trained with random CV. Models tuned with random CV overfit, performed poorly on independent samples, and were sensitive to outliers. The average percentage error (APE) decreased from ∼217% for a model with random-CV to ∼79% with spatiotemporal CV and FFS. Our findings demonstrate that proper alignment of model design with the data's spatiotemporal structure and modeling objective ensures reliability, minimizes data reproduction, and enables true prediction.
Related Concept Videos
Predicting Molecular Geometry
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Cross-bridge Cycle
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Data Validation
Key parameters for method validation include:

