基于树的机器学习以识别邻居层面的牛皮发病率预测因素:加拿大北克的一项人口研究
Anastasiya Muntyanu1,2, Raymond Milan1, Mohammed Kaouache3
1Department of Experimental Medicine, McGill University, Montreal, Canada.
American journal of clinical dermatology
|March 18, 2024
概括
邻里因素,如气候和社会经济地位显著影响牛皮发病率. 这项研究确定了影响社区层面上的牛皮风险的关键环境和社会预测因素,揭示了疾病发病率的地理差异.
科学领域:
- 流行病学 流行病学
- 环境健康 环境健康
- 机器学习 机器学习
背景情况:
- 牛皮影响全球大约6000万人,造成严重的健康负担.
- 以前的研究主要集中在个人行为上,忽视了健康的更广泛的决定因素.
- 新出现的证据强调了社会,经济和环境因素在健康结果中的作用.
研究的目的:
- 为了确定牛皮发病率的社区级风险因素.
- 利用来自加拿大北克的人口数据和先进的机器学习技术.
- 了解生活环境对牛皮发育的影响.
主要方法:
- 成人牛皮病例是使用来自北克人口数据库 (1997-2015) 的ICD-9/10代码来识别的.
- 牛皮发病前一年的环境和社会经济数据从CANUE和加拿大统计局收集.
- 使用渐变增强机器学习模型来预测牛皮发病率,通过AUC评估性能.
主要成果:
- 牛皮发病率在北克的地理位置上有所不同,从每10万人/年1.6到325.6不等.
- 通过节模型 (AUC 0.77) 确定的前9个预测因素包括紫外线辐射,温度,城市化和土壤水分 (负相关).
- 夜间灯光亮度显示出积极的关联,而社会经济因素表明中产阶级社区的发病率更高.
结论:
- 这项研究揭示了牛皮发病率在司法管辖区之间存在显著的差异.
- 生活环境,包括气候,植被,城市化和社区社会经济特征,与牛皮发病率有关.
- 这些发现强调了在牛皮研究和公共卫生战略中考虑宏观水平决定因素的重要性.
相关概念视频
Survival Tree
84
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
84
Statistical Methods for Analyzing Epidemiological Data
364
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
364
Cancer Survival Analysis
345
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
345
Mechanistic Models: Compartment Models in Individual and Population Analysis
40
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
40
Genome-wide Association Studies-GWAS
13.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.4K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


