动态伯恩斯坦GCN用于泛癌亚型分类使用RNA-Seq和CNV数据.
IEEE transactions on computational biology and bioinformatics
|January 12, 2026
概括
这项研究引入了动态伯恩斯坦图形卷积网络 (DB-GCN),用于准确的癌症亚型分类. DB-GCN有效地捕捉复杂的多omics相互作用,提高精确的瘤学和生物标志物发现.
科学领域:
- 计算生物学和生物信息学
- 在瘤学中的机器学习
- 系统生物学和网络分析.
背景情况:
- 由于复杂的多omics相互作用,癌症亚型的分类是具有挑战性的.
- 传统的机器学习模型很难有效地代表这些交互.
- 图形卷积网络 (GCNs) 使用生物拓,但具有固定的传播,限制了适应性.
研究的目的:
- 引入一种新的架构,即动态伯恩斯坦图形卷积网络 (DB-GCN),用于癌症亚型分类.
- 为了使拓意识学习使用适应性光谱传播与伯恩斯坦多项式.
- 在图形框架内支持单个和多个学科的数据集成.
主要方法:
- 使用伯恩斯坦多项式开发了具有适应性光谱传播的DB-GCN,避免了自身分解.
- 综合的单基因组 (RNA) 和多基因组 (RNA+CNV) 数据.
- 采用了双流设计,结合了伯恩斯坦图流和omics多层感知器.
主要成果:
- 通过使用多omics数据 (STRING上86.05%±0.83) 在28个TCGA亚型中实现了泛癌亚型分类的高准确性.
- 使用SHAP分析确定了假定生物标记基因 (例如KLK11,OR4F15,UBE2DNL).
- 发现,在前50个已识别的基因中,12个基因映射到KEGG癌症途径中.
结论:
- DB-GCN为胰腺癌亚型分类提供了一个准确和可解释的基于图形的框架.
- 该模型有效地捕捉复杂的基因相互作用,以提高精确的瘤学.
- DB-GCN为癌症研究提供了强大的生物标志物发现.
更多相关视频
相关概念视频
Genomics
35.5K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
35.5K
Comparing Copy Number Variations and SNPs
11.6K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
11.6K
DNA Microarrays
16.8K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
16.8K
Next-generation Sequencing
87.9K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
87.9K
Genome-wide Association Studies-GWAS
12.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.6K


