Related Experiment Video
Updated: Feb 25, 2026

09:43
Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
6.8K
Principles of Dataset Versioning: Exploring the Recreation/Storage Tradeoff
Souvik Bhattacherjee1, Amit Chavan1, Silu Huang2
1University of Maryland, College Park.
Summary
Managing dataset versions is challenging due to the storage-recreation trade-off. This study proposes efficient heuristics for dataset version management, balancing storage use and retrieval speed.
Area of Science:
- Data Science
- Computer Science
- Information Management
Background:
- Collaborative data science generates numerous dataset versions, posing management and storage challenges.
- The core issue is the storage-recreation trade-off: increased storage allows faster retrieval but uses more space.
- Existing research on this fundamental problem is limited.
Purpose of the Study:
- To systematically study the storage-recreation trade-off in dataset versioning.
- To formulate and analyze tractable and intractable problems related to this trade-off.
- To develop efficient heuristics for practical dataset version management.
Main Methods:
- Formulation of six distinct problems addressing the storage-recreation trade-off under various constraints.
- Demonstration of the intractability of most formulated problems.
- Development of heuristics inspired by delay-constrained scheduling and spanning tree algorithms.
Main Results:
- Proposed heuristics offer efficient solutions for dataset versioning scenarios.
- Experimental validation confirms the practical effectiveness of the developed heuristics.
- A prototype version management system was built as a foundation for DataHub.
Conclusions:
- The proposed heuristics effectively address the storage-recreation trade-off in dataset versioning.
- The developed system provides a practical foundation for collaborative data science environments.
- Further research can build upon these heuristics for enhanced data management systems.
Related Concept Videos
Storage
437
A schema is a mental framework that helps individuals organize and interpret information. Schemata, formed from previous experiences, influence how we process new information: how we encode it, the inferences we make, and how we retrieve it. For instance, a schema for what a typical classroom looks like might include desks, a teacher's desk, a whiteboard, and students in such an environment. This expectation helps us quickly understand and navigate new classrooms without needing to analyze...
437
Distribution Reliability and Automation
542
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
542
Data Reporting and Recording
5.5K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.5K
Data: Types and Distribution
2.0K
In biostatistics, data are the observations collected for analysis. There are two main types: parametric and non-parametric. Parametric data, which include continuous (e.g., weight) and discrete numerical data (e.g., number of tablets), assume a particular distribution pattern, often the normal distribution. Non-parametric data do not adhere to a specific distribution and typically comprise nominal (e.g., gender) and ordinal categorical data (e.g., pain scale ratings).
Distributions in...
Distributions in...
2.0K
Archival Research
17.5K
Some researchers gain access to large amounts of data without interacting with a single research participant. Instead, they use existing records to answer various research questions. This type of research approach is known as archival research. Archival research relies on looking at past records or data sets to look for interesting patterns or relationships. For example, a researcher might access the academic records of all individuals who enrolled in college within the past ten years and...
17.5K
Methods of Documentation I: Source-Oriented Records
1.8K
Source-oriented records, or SOR, are medical record-keeping organized by the data source. The SOR system was first developed in the mid-1900s to organize the growing patient data in hospitals and other healthcare facilities.
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following:
In an SOR, each discipline involved in patient care maintains a separate medical record section. This record-keeping method enables easy tracking of patient progress and ensures healthcare staff have access to up-to-date information.
Key Attributes include the following:
1.8K

