Related Experiment Video
Updated: Jan 13, 2026

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
A Combinatorial Approach to Synthetic Data Generation for Machine Learning
Krishna Khadka1, Jaganmohan Chandrasekaran2, Yu Lei1
1Department of Computer Science and Engineering, The University of Texas at Arlington, Arlington, TX 76019 USA.
This study introduces a novel combinatorial sampling method for generating synthetic data, significantly reducing the number of samples needed for comparable machine learning model performance and enhancing privacy protection.
Area of Science:
- Machine Learning
- Data Privacy
- Synthetic Data Generation
Background:
- Machine learning datasets frequently contain sensitive personal health and financial information, posing privacy risks.
- Existing synthetic data generation methods often require numerous samples, impacting downstream task efficiency.
- Current techniques involve encoding data, random sampling in latent space, and decoding to generate synthetic data.
Purpose of the Study:
- To develop an efficient synthetic data generation technique that minimizes sample requirements.
- To enhance the privacy preservation capabilities of synthetic data generation methods.
- To improve the performance of machine learning models using synthetic data.
Main Methods:
- A combinatorial approach to sampling the latent space is proposed, focusing on t-way interactions among latent dimensions.
- This method is motivated by findings that model predictions are often driven by interactions between a limited number of features.
- The approach generates synthetic data samples by utilizing these identified feature interactions.
Main Results:
- The combinatorial sampling approach requires fewer synthetic samples compared to traditional random sampling to achieve similar model performance.
- When combined with differential privacy, this method shows a smaller performance degradation than random sampling.
- Empirical results demonstrate the effectiveness of leveraging feature interactions for efficient synthetic data generation.
Conclusions:
- The proposed combinatorial sampling method offers a more efficient alternative for generating high-quality synthetic data.
- This technique improves the trade-off between data utility and privacy preservation in machine learning.
- The findings suggest that targeted sampling based on feature interactions can significantly enhance synthetic data generation processes.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Combinatorial Gene Control
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
Random Sampling Method
Synthetic Disvision of Polynomials
Mechanistic Models: Compartment Models in Individual and Population Analysis
Bootstrapping

