在基因组分析中整合先前的网络知识的权重重叠群拉索
Dan Huang1, Geunsu Jo1, Kipoong Kim2
1Department of Statistic, Pusan National University, Busan, 46241, Korea.
BMC bioinformatics
|September 1, 2025
概括
这项研究引入了一种用于基因组分析的新计算方法,利用网络结构来识别微妙的基因表达变化. 这种新方法有效地检测出与癌症相关的途径,
科学领域:
- 生物信息学
- 计算生物学
- 基因组学
背景情况:
- 基因组分析在实验条件之间识别差异性表达的基因,通常使用基因调控网络.
- 目前的统计方法忽略了网络结构,无法检测微妙或稀疏的基因表达信号.
- 这种局限性阻碍了由少数关键基因调节的复杂生物通路的识别.
研究的目的:
- 开发一种用于基因组分析的新计算方法,该方法结合了先前的网络知识.
- 通过利用基因网络结构来改善稀疏差异表达信号的基因组的检测.
- 加强像癌症基因组图集这样复杂的数据集中的生物相关途径的识别.
主要方法:
- 建议采用一种新方法,将基于网络的规范化与重叠的群体拉索结合起来.
- 基于网络的规范化增强了链接基因之间的关联信号.
- 叠加组激光器可以更轻松地选择相关的基因组,将网络信息作为权重 (加权叠加组激光器 - wOGL).
主要成果:
- 与现有方法相比,广泛的模拟证明了拟议的方法的优越性能.
- 对癌症基因组图谱的应用 乳腺侵入性癌症 (TCGA- BRCA) 数据成功确定了与癌症相关的重要途径.
- 这些途径以前没有被传统的基因组分析方法检测到.
结论:
- 权重重叠群拉索 (wOGL) 方法有效地利用先前的网络信息进行基因组分析.
- wOGL增强了含有差异表达基因的基因组的识别,特别是那些信号稀疏的基因组.
- 这种方法为发现基因组数据中的复杂监管途径提供了强大的工具.
相关概念视频
Genome-wide Association Studies-GWAS
14.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.1K
Combinatorial Gene Control
8.4K
Combinatorial gene control is the synergistic action of several transcriptional factors to regulate the expression of a single gene. The absence of one or more of these factors may lead to a significant difference in the level of gene expression or repression.
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
8.4K
Comparing the Survival Analysis of Two or More Groups
280
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
280
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
708
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
708
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K


