从较低层次的测量中直接估计和推断更高层次的相关性,用于基因通路和蛋白质组学研究
1Department of Biostatistics and Informatics, Colorado School of Public Health, 13001 E. 17th Place, Aurora, CO 80045, United States.
Biostatistics (Oxford, England)
|July 31, 2024
概括
本研究引入了一种新的潜伏因子模型,可以从低级数据直接估计高级生物变量之间的相关性,避免数据聚合. 这种方法提高了准确性,并使显著的生物相关性可靠地识别.
科学领域:
- 生物信息学是一种生物信息学.
- 统计遗传学 统计遗传学
- 系统生物学 系统生物学
背景情况:
- 估计高层次生物变量 (例如蛋白质,基因通路) 之间的相关性至关重要,但当仅直接观察低层次数据 (例如,基因) 时,这是具有挑战性的.
- 目前的方法汇总低级别数据,从而根据汇总技术进行可变的相关性估计.
- 这种聚合方法可以掩盖真正的生物学关系,并引入偏见.
研究的目的:
- 开发一种方法,从低水平测量中直接估计高水平的生物相关性,而无需数据聚合.
- 提高复杂生物系统中相关性估计的准确性和可靠性.
- 提供一个统计学上可靠的框架,用于识别显著的生物相关性.
主要方法:
- 建议使用潜在因子模型,从较低级别的数据直接估计较高级别的生物变量之间的相关性.
- 整合了一个收缩估计器,以确保正确的确定性,并提高相关性矩阵的准确性.
- 为了高效的P值计算,估计器的异常正常性被确立.
主要成果:
- 隐性因子模型成功地从低级数据直接估计了高水平的相关性,绕过了聚合偏差.
- 收缩估计器可以提高估计相关性矩阵的稳定性和准确性.
- 该方法在模拟和现实世界的蛋白质组学和基因表达数据分析中表现出有效性.
- 该R包"高得分"是为实际实施而开发的.
结论:
- 与基于聚合的方法相比,拟议的潜在因子模型为估计生物相关性提供了一种优越的方法.
- 这种方法提供了一个强大的工具,用于揭示复杂的生物关系在多omics数据.
- 开发的R包有助于在生物研究中应用这种先进的统计技术.
相关概念视频
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Proteomics
7.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.3K
DNA Microarrays
17.3K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.3K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K


