贝叶斯概况回归用于集群分析,涉及纵向响应和解释变量
Anaïs Rouanet1, Rob Johnson1, Magdalena Strauss1,2
1MRC Biostatistics Unit, School of Clinical Medicine, University of Cambridge, U.K.
概括
本研究介绍了PReMiuMlongi,这是贝叶斯概况回归的R包,可以分析纵向基因表达数据. 它确定了酵母细胞周期中的共同调节的基因组及其调节因子.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 统计建模 统计建模
背景情况:
- 在现代基因组学中,识别具有共同功能的共同调节基因至关重要.
- 贝叶斯概况回归是一种半监督的方法,用于基于响应变量的数据聚类.
- 现有的方法处理单变结果,限制了更广泛的应用.
研究的目的:
- 扩展贝叶斯概况回归用于纵向 (多变量连续) 结果.
- 推出PReMiuMlongi,这是一个更新的R包用于配置回归分析.
- 应用扩展模型来识别酵母中共同调节的基因组.
主要方法:
- 开发了对纵向数据的贝叶斯概况回归的扩展.
- 纳入多变量正常和高斯过程回归响应模型.
- 使用PReMiuMlongi R包进行分析和模拟研究.
主要成果:
- 成功地将该模型应用于Saccharomyces cerevisiae细胞循环中的开花酵母数据.
- 根据它们的表达轨迹,确定了四个不同的协同调节基因组.
- 将这些基因组与参与共同调节的特定转录因子联系起来.
结论:
- 扩展的贝叶斯概况回归模型有效地分析纵向基因表达数据.
- PReMiuMlongi促进了共同调节的基因组及其调节机制的发现.
- 这种方法提高了我们对酵母细胞周期期间基因调节的理解.
相关概念视频
Longitudinal Studies
158
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
158
Parametric Survival Analysis: Weibull and Exponential Methods
424
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
424
Longitudinal Research
12.0K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
12.0K
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K


