通过随机森林增强酒精消费相关基因的选择
Chenglin Lyu1,2, Roby Joehanes3, Tianxiao Huan3
1Department of Biostatistics, Boston University School of Public Health, Boston, MA02118, USA.
The British journal of nutrition
|April 12, 2024
概括
一个受监督的机器学习算法识别了新的与酒精相关的基因转录. 博鲁塔算法选择了25个成绩单,显示了与传统方法在区分酒精消费水平方面的准确性.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 文字转录学 (Transcriptomics) 是一个学科.
背景情况:
- 机器学习 (ML) 越来越多地用于识别omics标记.
- 酒精消费的转录基因标记需要进一步研究.
研究的目的:
- 评估监督的ML算法是否可以增强与酒精相关的转录组标记物的识别.
- 识别与饮酒模式相关的新型基因转录.
主要方法:
- 分析了来自5508名弗雷明汉心脏研究参与者的基于阵列的全血基因表达数据.
- 应用Boruta算法,一个监督随机森林 (RF) 基于特征选择方法.
- 使用接收器操作特征 (ROC) 曲线分析 (曲线下的面积 - AUC) 验证选定的转录.
主要成果:
- 博鲁塔算法确定了25个与酒精相关的转录.
- 选择的25个转录实现了AUC值为0.73 (非与中度饮酒者),0.69 (非与重度饮酒者) 和0.66 (中度与重度饮酒者).
- 这些AUC值与使用常规线性回归模型获得的值可比.
- 在25个选定的成绩单中,发现了13个与肥胖相关的成绩单,3个与2型糖尿病相关的成绩单,1个与高血压相关的成绩单.
- 像DOCK4,IL4R和SORT1这样的特定基因显示出与酒精消费的反向关联,以及与肥胖或高血压的联系.
结论:
- 监督ML,特别是基于RF的Boruta算法,有效地识别了新的酒精相关基因转录.
- 鉴定到的转录显示了酒精消费和相关健康状况 (如肥胖和高血压) 的生物标志物的潜力.
相关概念视频
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Mutation, Gene Flow, and Genetic Drift
58.4K
In a population that is not at Hardy-Weinberg equilibrium, the frequency of alleles changes over time. Therefore, any deviations from the five conditions of Hardy-Weinberg equilibrium can alter the genetic variation of a given population. Conditions that change the genetic variability of a population include mutations, natural selection, non-random mating, gene flow, and genetic drift (small population size).
58.4K


