Related Experiment Video
Updated: May 16, 2025

Quantification of Orofacial Phenotypes in Xenopus
Published on: November 6, 2014
Collinear datasets augmentation using Procrustes validation sets
Sergey Kucheryavskiy1, Sergei Zhilin2
1Department of Chemistry and Bioscience, Aalborg University, Niels Bohrs vej 8, Esbjerg, 6700, Denmark.
Background:
high complexity models, such as artificial neural networks (ANN), require large datasets for training to avoid overfitting and reproducibility issues. However, experimental datasets, especially those involving spectroscopic or other highly collinear data, often suffer from limited size due to practical constraints. Currently available data augmentation methods, either do not handle collinearity well, or require resource-intensive training. Thus, there is a pressing need for an efficient, scalable method for augmenting collinear datasets to enhance model performance in both regression and classification tasks.
Results:
we propose a novel, efficient data augmentation method tailored for datasets with moderate to high collinearity, particularly spectroscopic data. This method utilizes latent variable modeling combined with cross-validation resampling to generate new data points. The approach has been validated using varios datasets, here we report detailed results for two case studies: fat content prediction in minced meat and discrimination of olives based on near-infrared spectra. In both cases, artificial neural networks were employed, resulting in significant improvements in model performance in prediction and classification. Specifically, for fat content prediction, the method reduced the root mean squared error by up to 3-fold on the independent test set.
Significance:
the proposed method provides a fast and simple solution for augmenting collinear datasets, significantly improving model performance without requiring extensive parameter tuning. It is versatile and can be applied to a range of datasets, offering a practical alternative to more complex augmentation techniques.
Related Concept Videos
Improving Translational Accuracy
Wilcoxon Signed-Ranks Test for Matched Pairs
Kendall's Coefficient of Concordance
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Calibration Curves: Correlation Coefficient

