対照的線形回帰
Boyang Zhang1, Sarah Nyquist2, Andrew Jones3
1Department of Genetics, Stanford University.
The annals of applied statistics
|December 12, 2025
まとめ
応答変数を持つ症例対照データを分析するための新しい方法である対照的回帰を導入します。このアプローチは、自閉症の重症度や腫瘍の段階などの結果に関連する主要な生物学的予測因子を特定します。
科学分野:
- 生物統計学
- ゲノミクス
- 計算生物学
背景:
- 症例対照研究は生物医学研究で一般的です。
- 既存の次元削減方法は、症例と対照間のバリエーションを特定します。
- 応答変数を持つ症例対照データの分析にはギャップがあります。
研究 の 目的:
- 応答変数を持つ症例対照データのための対照的回帰を開発すること。
- 症例と対照間の共有バリエーションを捉えること。
- 残りの予測因子バリアンスを使用して症例特異的応答を説明すること。
主な方法:
- 対照的線形回帰モデルを開発しました。
- このモデルを単一細胞RNAシーケンシングデータ(慢性副鼻腔炎)に適用しました。
- このモデルを単一核RNAシーケンシングデータ(自閉症の重症度)に適用しました。
主要な成果:
- 対照的線形回帰モデルは効果的に特徴をランク付けします。
- 応答変数に関連する生物学的に情報量の多い予測因子を特定しました。
- これらの予測因子は他の方法では特定できませんでした。
結論:
- 対照的回帰は、複雑な症例対照データを分析するための強力なツールです。
- この方法は、疾患メカニズムと患者層別化の理解を強化します。
- 自閉症の重症度や細胞分化などのデータセットで新たな洞察を提供します。
関連する概念動画
Correlation and Regression
3.0K
In statistics, correlation describes the degree of association between two variables. In the subfield of linear regression, correlation is mathematically expressed by the correlation coefficient, which describes the strength and direction of the relationship between two variables. The coefficient is symbolically represented by 'r' and ranges from -1 to +1. A positive value indicates a positive correlation where the two variables move in the same direction. A negative value suggests a...
3.0K
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Regression Analysis
7.8K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.8K
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Calibration Curves: Linear Least Squares
4.0K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
4.0K
Calculating and Interpreting the Linear Correlation Coefficient
7.7K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
7.7K


