对于没有UMI的单细胞RNA-seq数据,复合模型和皮尔森残留值.
bioRxiv : the preprint server for biology
|August 14, 2023
概括
本研究引入了一种新的化合物分布模型,以使非UMI测序数据正常化. 这种方法通过解决测序数据中的过度分散和零通货膨胀来改善基因选择和数据嵌入.
科学领域:
- 计算生物学 计算生物学
- 基因组学就是基因组学.
- 统计建模 统计建模
背景情况:
- 目前用于规范化单细胞RNA测序 (scRNA-seq) 数据的现有方法通常依赖于独特的分子标识符 (UMI).
- 非UMI数据的规范化,如Smart-seq2,由于放大偏差和复杂的分布模式,提出了挑战.
研究的目的:
- 开发一个新的统计框架来规范非UMI scRNA-seq数据.
- 扩展皮尔森残余的实用性,用于基因选择和维度减小,用于缺乏UMI的数据集.
- 准确地建模非UMI协议中固有的技术噪声.
主要方法:
- 使用负二项式分布对测序RNA分子进行建模.
- 纳入放大分布以考虑非UMI数据中的技术偏差.
- 根据拟议的化合物分布,开发化合物皮尔森残留物.
- 用破碎的功率定律描述放大分布.
主要成果:
- 复合皮尔森残余模型有效地使Smart-seq2数据集正常化.
- 该模型为非UMI数据提供了有意义的基因选择和信息嵌入.
- 一个被打破的功率定律准确地描述了跨各种测序协议的放大分布.
- 复合模型成功地解决了特定于非UMI数据的过度分散和零通胀模式.
结论:
- 拟议的化合物分布模型为规范非UMI scRNA-seq数据提供了一个强大的方法.
- 这种方法提高了像Smart-seq2.2这样的协议中的基因表达数据的解释性和实用性.
- 这些发现提供了一个更准确的RNA测序实验的统计表现,没有UMIs.
更多相关视频
12:54Real-time Analysis of Transcription Factor Binding, Transcription, Translation, and Turnover to Display Global Events During Cellular Activation
Published on: March 7, 2018
13.6K
05:59Author Spotlight: Deciphering the Cellular Mysteries of Intermuscular Adipose Tissue in Humans
Published on: May 3, 2024
735
相关概念视频
RNA-seq
10.1K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis
64
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
64
