评估点突变作为深度学习数据增强的可靠性,使用基因组数据进行深度学习
Hyunjung Lee1, Utku Ozbulak2, Homin Park2,3
1Korea University, Seoul, South Korea.
BMC bioinformatics
|April 30, 2024
概括
这项研究引入了一种使用点突变的基因组数据的新型数据增强方法. 这种生物启发的技术提高了深度神经网络在遗传疾病分析中的性能.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 深度神经网络 (DNN) 对遗传疾病研究具有前景,但需要广泛的训练数据.
- 来自其他领域的现有数据增强方法通常在具有独特基因组数据属性的情况下失败.
- 基因组数据增强对于推动DNN在遗传学中的应用至关重要.
研究的目的:
- 为基因组数据开发一种新的数据增强技术.
- 为了解决数据稀缺的局限性,在训练深度神经网络中进行基因分析.
- 提高DNN在基因组任务上的性能.
主要方法:
- 提出了一种由生物点突变启发的基因组数据的新数据增强技术.
- 利用点突变作为基因组序列中的代码子的替代品.
- 评估了该技术对遗传任务DNN性能的影响.
主要成果:
- 拟议的基于点突变的数据增强增强了DNN在基因组任务上的性能.
- 在涉及编码区域的任务中观察到改善,例如翻译启动和拼接位置检测.
- 沉默和错误的突变对模型的有效性产生了积极的影响.
结论:
- 基于点突变的数据增强为改善DNA序列的预测模型提供了有价值的策略.
- 仔细选择突变类型 (例如,沉默,错误) 对于积极的结果很重要.
- 这种方法为提高基因组预测模型的准确性和可靠性提供了机会.
相关概念视频
Mutations
82.0K
Overview
82.0K
Mismatch Repair
4.8K
Organisms are capable of detecting and fixing nucleotide mismatches that occur during DNA replication. This sophisticated process requires identifying the new strand and replacing the erroneous bases with correct nucleotides. Mismatch repair is coordinated by many proteins in both prokaryotes and eukaryotes.
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
4.8K
Genome Copying Errors
4.2K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.2K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
In-vitro Mutagenesis
13.9K
To learn more about the function of a gene, researchers can observe what happens when the gene is inactivated or “knocked out,” by creating genetically engineered knockout animals. Knockout mice have been particularly useful as models for human diseases such as cancer, Parkinson’s disease, and diabetes.
13.9K


