Related Experiment Video
Updated: Mar 28, 2026

A Web Tool for Generating High Quality Machine-readable Biological Pathways
Published on: February 8, 2017
Reward-Guided Generation Improves the Scientific Utility of Synthetic Biomedical Data
Nicholas J Jackson1, Natalia Espinosa-Dice2, Chao Yan3
1Vanderbilt University, Nashville, TN.
Abstract:
Synthetic data generation is a promising approach for biomedical data sharing and dataset augmentation, yet existing methods lack mechanisms to preserve statistical properties necessary for scientific analysis. To address this, we introduce RLSyn+Reg, a reinforcement learning-driven generative model, which encourages that regression models trained on synthetic data reproduce the coefficients and predictions of their real-data counterparts. We evaluate RLSyn+Reg on MIMIC-III and the American Community Survey (ACS) across regression model reproduction, fidelity to real data, and privacy. Synthetic data from RLSyn+Reg substantially improves upon that of RLSyn, raising correlations between real and synthetic regression coefficients from 0.054 to 0.600 on MIMIC-III and from 0.160 to 0.376 on ACS. Predictive performance also improves, reducing the gap between real-data baselines by 81.4% and 97.6% on MIMIC-III and ACS, respectively. These improvements come with negligible cost to fidelity or privacy and are robust to reductions in training data.
Related Concept Videos
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Improving Translational Accuracy
Improving Translational Accuracy

