一个新的生成对抗性网络模型,用于高维欧米克数据中的类不平衡问题
Samuel Cusworth1,2, Georgios V Gkoutos3,4,5,6,7, Animesh Acharjee8,9,10,11
1Institute of Applied Health Research, University of Birmingham, Birmingham, UK.
BMC medical informatics and decision making
|March 29, 2024
概括
在omics数据中的类不平衡阻碍了机器学习. 一种新的生成对抗网络方法创建合成样本,以提高分类器的性能,而不是SMOTE和随机过量采样.
科学领域:
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
- 基因组学就是基因组学.
背景情况:
- 类不平衡是高通量omics分析的一个重大挑战,导致有偏见的机器学习分类器.
- 传统的过量采样方法,如SMOTE和随机过量采样,可能会带来不准确性,特别是在高维和杂的医疗保健数据中.
研究的目的:
- 提出一种基于生成对抗网络 (GAN) 的新方法,用于从小的,高维的奥米克数据集中生成合成样本.
- 改进现有的生成方法来解决机器学习中的阶级不平衡问题.
主要方法:
- 开发了一个生成对抗网络 (GAN) 模型,为少数阶级创建合成数据点.
- 提出的基于GAN的过量采样方法与合成少数群体过量采样技术 (SMOTE) 和随机过量采样 (RO) 相比较.
- 分类人员接受了数据平衡的培训,使用每种过量采样技术来验证生成方法.
主要成果:
- 基于GAN的方法证明了在高维的奥米克数据中生成合成样本的潜力.
- 通过分类员培训进行的绩效评估表明,与天真的过量采样技术相比,有所改善.
- 该研究强调了生成模型在缓解阶级不平衡偏差方面的有效性.
结论:
- 生成对抗性网络提供了一种有希望的方法来解决高通量欧米克数据中的类不平衡.
- 这种方法为传统的过量采样技术提供了更强大的替代方案,特别是对于杂的高维数据集.
- 拟议的基于GAN的策略增强了机器学习分类器在omics研究中的概括性.
更多相关视频
03:37Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
709
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
4.0K
相关概念视频
Aggregates Classification
317
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
317
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
How Data are Classified: Categorical Data
32.7K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
32.7K
Classification of Systems-I
184
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
184
Classification of Systems-II
144
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
144
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
