SAE-Impute:通过子空间回归和自动编码器对单单元数据进行归算
Liang Bai1, Boya Ji2, Shulin Wang3
1College of Computer Science and Electronic Engineering, Hunan University, Changsha, 410082, China.
BMC bioinformatics
|October 1, 2024
概括
SAE-Impute有效地解决了单细胞RNA测序 (scRNA-seq) 数据中的脱落事件. 这种新方法通过利用子空间回归和自动编码器来赋值缺失值来提高数据的准确性和可解释性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 单细胞RNA测序 (scRNA-seq) 对于研究细胞异质性至关重要.
- 在scRNA-seq数据中的脱落事件对分析提出了重大挑战.
- 现有的归算方法往往忽略了样本间的相关性.
研究的目的:
- 介绍SAE-Impute,一种用于赋值scRNA-seq数据的新型计算方法.
- 为了提高失业数据归算的准确性和可靠性.
- 改进scRNA-seq数据的下游分析.
主要方法:
- SAE-Impute结合了子空间回归和自动编码器.
- 亚空间回归评估样本相关性.
- 自动编码器使用预测数据来插入失效值.
主要成果:
- 在scRNA-seq数据中,SAE-Impute减少了虚假负信号.
- 该方法改善了丢失值和相关性的检索.
- 下游分析,如差异性基因表达和细胞聚类,得到了增强.
结论:
- 在单细胞数据集中,SAE-Impute有效地减少了失业率.
- 归算方法提高了scRNA-seq数据的功能解释性.
相关概念视频
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
411
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
411
Extraction: Partition and Distribution Coefficients
2.2K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
2.2K
Regression Analysis
5.6K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.6K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


