Related Experiment Video
Updated: Sep 4, 2025

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
Surrogate- and invariance-boosted contrastive learning for data-scarce applications in science
Charlotte Loh1, Thomas Christensen2, Rumen Dangovski3
1Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA, USA. cloh@mit.edu.
This study introduces surrogate- and invariance-boosted contrastive learning (SIB-CL), a deep learning method that significantly reduces the need for labeled data in scientific applications. SIB-CL leverages unlabeled data, symmetries, and surrogate data for efficient model training.
Area of Science:
- Computational Science
- Materials Science
- Quantum Mechanics
- Photonics
Background:
- Deep learning is increasingly used in natural sciences for tasks like property prediction and material discovery.
- Training deep learning models typically requires large amounts of labeled data, which is often scarce and expensive to obtain in scientific research.
- Auxiliary information sources are often available in scientific problems but are underutilized in current deep learning frameworks.
Purpose of the Study:
- To develop a novel deep learning framework to address data scarcity challenges in scientific applications.
- To incorporate readily available auxiliary information—unlabeled data, prior knowledge of symmetries/invariances, and low-cost surrogate data—to enhance model training.
- To demonstrate the framework's effectiveness and generalizability across diverse scientific problems.
Main Methods:
- Introduced surrogate- and invariance-boosted contrastive learning (SIB-CL), a deep learning framework.
- SIB-CL integrates three key auxiliary information sources: abundant unlabeled data, known physical symmetries or invariances, and cost-efficient surrogate data.
- The framework was applied to predict the density-of-states of 2D photonic crystals and solve the 3D time-independent Schrödinger equation.
Main Results:
- SIB-CL demonstrated significant effectiveness and generality across different scientific problems.
- The framework achieved orders of magnitude reduction in the number of required labels while maintaining high network accuracy.
- SIB-CL successfully leveraged auxiliary information to overcome data limitations inherent in scientific deep learning tasks.
Conclusions:
- SIB-CL offers a powerful solution for data-scarce deep learning in the natural sciences.
- The method significantly lowers the barrier to entry for applying advanced machine learning techniques to scientific discovery.
- SIB-CL's ability to utilize inexpensive auxiliary data makes it a practical and efficient approach for scientific research.
Related Concept Videos
Improving Translational Accuracy
Generalization, Discrimination, and Extinction
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
Observational Learning
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...

