Related Experiment Video
Updated: Jan 25, 2026

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
Machine learning in a data-limited regime: Augmenting experiments with synthetic data uncovers order in crumpled
Jordan Hoffmann1, Yohai Bar-Sinai1, Lisa M Lee1
1John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA 02138, USA.
Machine learning can analyze complex data, even when experimental data is limited. By adding simulated data from simpler systems, predictive power is significantly improved for real-world experiments.
Area of Science:
- Physics
- Materials Science
- Data Science
Background:
- Machine learning (ML) excels with complex, high-dimensional data.
- ML application is limited in experimental systems with scarce or costly data.
- Crumpled thin sheets exhibit complex spatial order in crease networks.
Purpose of the Study:
- To develop a strategy for applying machine learning in data-limited experimental settings.
- To investigate the effectiveness of augmenting scarce experimental data with simulated data.
- To improve the predictive capabilities of machine learning models in experimental research.
Main Methods:
- Augmenting scarce experimental data with synthetically generated data from a simpler sister system.
- Simulating rigid flat-folded sheets to generate inexhaustible training data.
- Utilizing machine learning techniques to analyze local order in crease networks of crumpled sheets.
- Testing the improved predictive power on a pattern completion task.
Main Results:
- Machine learning techniques proved effective even in a data-limited regime for crumpled sheet analysis.
- Augmenting experimental data with simulated data significantly improved predictive power.
- The strategy demonstrated the utility of ML in bench-top experiments with limited data.
- Common statistical properties between systems facilitated successful data augmentation.
Conclusions:
- A novel strategy effectively overcomes data scarcity limitations in experimental machine learning.
- Simulated data from simpler systems can successfully augment experimental datasets.
- This approach enhances the applicability of machine learning in resource-constrained experimental research.
- The findings pave the way for broader ML adoption in experimental science.
More Related Videos
07:40A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
07:59Author Spotlight: Alignment of Synchronized Time-Series Data Using the Characterizing Loss of Cell Cycle Synchrony Model for Cross-Experiment Comparisons
Published on: June 9, 2023
Related Concept Videos
Data Collection by Experiments
An example of the experimental method is a public...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data Reporting and Recording
Data Collection I
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...