一种用于机器学习的合成数据生成的组合方法
Krishna Khadka1, Jaganmohan Chandrasekaran2, Yu Lei1
1Department of Computer Science and Engineering, The University of Texas at Arlington, Arlington, TX 76019 USA.
概括
这项研究引入了一种用于生成合成数据的新型组合采样方法,大大减少了用于可比机器学习模型性能所需的样本数量,并加强了隐私保护.
科学领域:
- 机器学习 机器学习
- 数据 隐私 数据 隐私 数据
- 合成数据生成 合成数据生成
背景情况:
- 机器学习数据集经常包含敏感的个人健康和财务信息,构成隐私风险.
- 现有的合成数据生成方法通常需要大量的样本,影响下游任务效率.
- 目前的技术包括编码数据,在潜空间中随机采样,以及解码以生成合成数据.
研究的目的:
- 开发一种高效的合成数据生成技术,尽量减少样本要求.
- 增强合成数据生成方法的隐私保护能力.
- 用合成数据提高机器学习模型的性能.
主要方法:
- 建议采用组合方法对隐性空间进行采样,重点关注隐性维度之间的t路相互作用.
- 这种方法的动机是发现模型预测通常是由有限数量的特征之间的相互作用驱动的.
- 该方法通过利用这些已识别的特征相互作用来生成合成数据样本.
主要成果:
- 与传统的随机抽样相比,组合抽样方法需要更少的合成样本来实现类似的模型性能.
- 当与差异隐私相结合时,这种方法比随机抽样显示出较小的性能退化.
- 经验结果证明了利用特征交互来有效生成合成数据的有效性.
结论:
- 拟议的组合采样方法为生成高质量的合成数据提供了更有效的替代方案.
- 这种技术在机器学习中改善了数据实用性和隐私保护之间的权衡.
- 这些发现表明,基于特征相互作用的有针对性的采样可以显著提高合成数据生成过程.
相关概念视频
Synthetic Biology
5.5K
Synthetic biology is an interdisciplinary science that involves using principles from disciplines such as engineering, molecular biology, cell biology, and systems biology. It involves remodeling existing organisms from nature or constructing completely new synthetic organisms for applications such as protein or enzyme production, bioremediation, value-added macromolecule production, and the addition of desirable traits to crops, to name a few.
Golden rice
Golden rice is a genetically modified...
Golden rice
Golden rice is a genetically modified...
5.5K
Combinatorial Gene Control
9.5K
Combinatorial gene control is the synergistic action of several transcriptional factors to regulate the expression of a single gene. The absence of one or more of these factors may lead to a significant difference in the level of gene expression or repression.
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
9.5K
Random Sampling Method
14.1K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
14.1K
Synthetic Disvision of Polynomials
139
Synthetic division is an efficient algorithmic approach for dividing a polynomial by a linear binomial of the form x - c, where c is a real number. This method is helpful due to its streamlined process, which avoids the more cumbersome steps involved in the traditional long division of polynomials. It simplifies computation and serves as a practical tool for evaluating polynomials and identifying their factors.To perform synthetic division, one begins by listing the coefficients of the...
139
Mechanistic Models: Compartment Models in Individual and Population Analysis
237
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
237
Bootstrapping
798
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
798


