Related Experiment Video
Updated: Jun 28, 2025

Standardizing a Non-Lethal Method for Characterizing the Reproductive Status and Larval Development of Freshwater Mussels Bivalvia: Unionida
Published on: October 4, 2019
Standardised Versioning of Datasets: a FAIR-compliant Proposal.
Alba González-Cebrián1, Michael Bradford2, Adriana E Chis2
1Cloud Competency Centre, National College of Ireland, Dublin, Ireland. alba.gonzalez-cebrian@ncirl.ie.
A new dataset versioning framework uses software engineering principles and data drift metrics to track changes. The dE,PCA metric efficiently identifies dataset updates, ensuring data usability and reproducibility.
Area of Science:
- Data Science
- Machine Learning
- Software Engineering
Background:
- Effective dataset versioning is crucial for reproducibility and workflow integration.
- Existing methods lack standardized tracking for data evolution and usability.
- Data drift quantification is needed to assess statistical property changes over time.
Purpose of the Study:
- To introduce a standardized dataset versioning framework.
- To develop and evaluate data drift metrics for quantifying dataset changes.
- To identify the most effective metric for tracking dataset updates and ensuring data usability.
Main Methods:
- Implementation of a "major.minor.patch" versioning nomenclature.
- Development of three data drift metrics (dP, dE,PCA, dE,AE) using unsupervised Machine Learning (PCA, Autoencoders).
- Evaluation of metrics for dataset creation, update, and deletion scenarios.
Main Results:
- The dE,PCA metric, combining PCA and splines, demonstrated efficient computation.
- Low metric values (<50) indicated new dataset batches or seasonal variations.
- High metric values (100) signified major updates (e.g., scaling >30% of variables).
- The metric showed robustness against information loss, with values near 0 for minor changes.
Conclusions:
- The proposed framework enhances dataset reusability, recognition, and version tracking.
- The dE,PCA metric offers a favorable balance of interpretability, robustness, and computational efficiency.
- This approach facilitates informed decision-making regarding data usability and workflow integration.
Related Concept Videos
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Data Reporting and Recording
Statistical Software for Data Analysis and Clinical Trials
Estimating Population Standard Deviation
Coefficient of Variation
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...

