Related Experiment Video
Updated: Jul 6, 2025

Leveraging CyVerse Resources for De Novo Comparative Transcriptomics of Underserved Non-model Organisms
Published on: May 9, 2017
Open source and reproducible and inexpensive infrastructure for data challenges and education
Peter E DeWitt1, Margaret A Rebull2, Tellen D Bennett3,4,5
1Department of Biomedical Informatics, University of Colorado School of Medicine, University of Colorado, Aurora, CO, USA. peter.dewitt@cuanschutz.edu.
Researchers developed an inexpensive, lightweight workflow for data sharing and hosting research data challenges. This approach leverages public repositories and open-source tools, making advanced data analysis accessible to more scientists.
Area of Science:
- Biomedical research
- Data science
- Computational biology
Background:
- Data sharing is crucial for maximizing research knowledge and enabling secondary analyses.
- Current biomedical data challenges often require expensive cloud computing and industry partnerships, creating financial barriers.
- There is a need for cost-effective and computationally accessible methods for data sharing and hosting data challenges.
Purpose of the Study:
- To develop an inexpensive and computationally lightweight workflow for reproducible model training, testing, and evaluation.
- To demonstrate the utility of this workflow by conducting a data challenge.
- To provide a resource for researchers seeking accessible infrastructure for data sharing and data challenges.
Main Methods:
- Leveraged public GitHub repositories for code sharing and version control.
- Utilized open-source computational languages for model development and analysis.
- Employed Docker technology for containerization, ensuring reproducibility and portability.
- Developed a workflow for reproducible model training, testing, and evaluation.
Main Results:
- Successfully developed and implemented a cost-effective infrastructure for hosting a data challenge.
- The workflow enabled reproducible model training, testing, and evaluation.
- The developed infrastructure and workflow proved effective for the conducted data challenge.
Conclusions:
- The developed workflow and infrastructure offer an accessible alternative to expensive cloud-based solutions.
- This approach facilitates data sharing and the hosting of data challenges in biomedical research.
- The methodology is likely to be valuable for both data challenges and educational purposes in science.
More Related Videos
09:43Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
08:29An Open Source Technology Platform to Manufacture Hydrogel-Based 3D Culture Models in an Automated and Standardized Fashion
Published on: March 31, 2022
Related Concept Videos
GIS Software, Hardware, and Sources of GIS Data
Distribution Reliability and Automation
Introduction to R
Data Reporting and Recording
Data Collection by Experiments
An example of the experimental method is a public...
Archival Research