Related Experiment Video
Updated: Jun 2, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
On the Readiness of Scientific Data Papers for a Fair and Transparent Use in Machine Learning
Joan Giner-Miguelez1,2, Abel Gómez3, Jordi Cabot4,5
1Internet Interdisciplinary Institute (IN3), Universitat Oberta de Catalunya (UOC), Barcelona, Spain. joan.giner@bsc.es.
This study evaluates scientific data papers for machine learning (ML) readiness. We propose guidelines to improve data documentation for fairer and more transparent ML technologies.
Area of Science:
- Data Science
- Machine Learning
- Scientific Publishing
Background:
- Growing demand for fair and trustworthy machine learning (ML) systems necessitates comprehensive data documentation.
- Scientific communities increasingly adopt data-sharing practices for reproducibility, publishing data and documentation in data papers.
- Existing data documentation practices may not fully align with the needs of ML applications and regulatory requirements.
Purpose of the Study:
- To analyze the extent to which scientific data papers meet the documentation needs of the machine learning community and regulatory bodies.
- To assess the coverage and trends in data documentation within scientific data papers across various domains.
- To compare the documentation standards in general scientific data papers with those from an ML-specific venue.
Main Methods:
- Analysis of a sample of 4041 data papers from diverse scientific domains.
- Assessment of data coverage and documentation trends relevant to ML applications.
- Comparative analysis with datasets published in an ML-focused venue (NeurIPS D&B).
Main Results:
- Identification of gaps and trends in data documentation within scientific data papers.
- Evaluation of the suitability of current data paper standards for ML use cases.
- Benchmarking against documentation practices in ML-specific dataset publications.
Conclusions:
- Scientific data papers show varying degrees of preparedness for ML applications.
- Recommendations are proposed for data creators and publishers to enhance data documentation.
- Guidelines aim to improve data transparency, fairness, and trustworthiness in ML technologies.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
08:05Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Related Concept Videos
Statistical Software for Data Analysis and Clinical Trials
Introduction to R
Data Validation
Key parameters for method validation include:
Ethics in Research
Bias in Epidemiological Studies
Regression Toward the Mean