数据集的标准化版本:一个符合FAIR的提案
Alba González-Cebrián1, Michael Bradford2, Adriana E Chis2
1Cloud Competency Centre, National College of Ireland, Dublin, Ireland. alba.gonzalez-cebrian@ncirl.ie.
Scientific data
|April 9, 2024
概括
一个新的数据集版本管理框架使用软件工程原理和数据漂移指标来跟踪变化. dE,PCA指标有效地识别数据集更新,确保数据的可用性和可重复性.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 软件工程 软件工程 软件工程
背景情况:
- 有效的数据集版本管理对于可重现性和工作流集成至关重要.
- 现有的方法缺乏对数据演变和可用性的标准化跟踪.
- 需要数据漂移量化来评估随时间变化的统计属性.
研究的目的:
- 引入一个标准化的数据集版本管理框架.
- 开发和评估数据漂移指标,以量化数据集变化.
- 确定跟踪数据集更新和确保数据可用性的最有效指标.
主要方法:
- 实施一个"major.minor.patch"版本命名体系.
- 使用无监督机器学习 (PCA,自动编码器) 开发了三种数据漂移指标 (dP,dE,PCA,dE,AE).
- 评估数据集创建,更新和删除场景的指标.
主要成果:
- 结合PCA和splines的dE,PCA度量,证明了高效的计算.
- 低度数值 (<50) 表示新数据集批次或季节性变化.
- 高度指标值 (100) 表示重大更新 (例如,变量缩放>30%).
- 该指标显示了对信息丢失的稳定性,微小变化的值接近0.
结论:
- 拟议的框架提高了数据集的可重复使用性,识别和版本跟踪.
- 该dE,PCA指标提供了一个有利的平衡的可解释性,稳定性和计算效率.
- 这种方法促进了有关数据可用性和工作流集成的知情决策.
相关概念视频
One-Way ANOVA: Equal Sample Sizes
3.3K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.3K
Testing a Claim about Standard Deviation
2.4K
A complete procedure to test a claim about population standard deviation or population variance is explained here.
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
2.4K
Data Reporting and Recording
4.7K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
4.7K
Statistical Software for Data Analysis and Clinical Trials
546
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
546
Estimating Population Standard Deviation
3.0K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.0K
Coefficient of Variation
3.8K
The coefficient of variation measures the dispersion of the data points or distribution around the mean. Using the coefficient of variation, we can compare two data series with drastically different means or different units of measurement. The coefficient of variation for a sample and a population is expressed as a percentage of the ratio of standard deviation to the mean.
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
The coefficient of variation is a practical statistical tool in finance. It allows investors to assess the volatility or...
3.8K


