使用机器学习识别肥胖人口中的集群:马斯特里赫特研究的二次分析
Maik Jm Beuken1, Melanie Kleynen2, Susy Braun2
1Faculty of Financial Management, Research Center for Statistics & Data Science, Zuyd University of Applied Sciences, Sittard, Netherlands.
JMIR medical informatics
|February 5, 2025
概括
这项研究使用机器学习来识别三个不同的肥胖个体群,揭示了能量摄入,职业,性别,认知功能和教育的关键差异. 这些发现凸显了个性化健康干预的潜力.
科学领域:
- 数据科学数据科学数据科学
- 机器学习 机器学习
- 个性化医疗是个性化的医疗.
背景情况:
- 现代的生活方式与身体不活动和营养不良驱动肥胖和慢性疾病.
- 个性化干预对于长期的行为改变是有效的.
- 机器学习 (ML) 可以在大型数据集中发现复杂的关系和人口集群.
研究的目的:
- 为了识别不同群体的肥胖个体.
- 用数据驱动,无假设的ML方法来发现区分这些集群的相关变量.
主要方法:
- 使用了来自马斯特里赫特研究的横截面数据 (n=4128) 与2971个变量.
- 应用了因子概率距离聚类算法用于高维数据分析.
- 采用统计学相当的签名算法来识别不同的,最小冗余变量.
主要成果:
- 在肥胖群体中确定了3个不同的群体.
- 集群1:以较低的能源消耗和较高的失业率为特征.
- 集群2:以更高的能量摄入量和主要男性参与者为特征.
- 集群3:具有更高的认知功能和更高的教育成就.
结论:
- 一种数据驱动的,无假设的ML方法成功地在一个大而复杂的肥胖数据集中确定了可区分的集群.
- 确定了关键的差异化变量 (能量摄入量,职业,性别,认知功能,教育).
- 这些发现支持针对肥胖管理制定有针对性的,个性化的干预措施.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
26
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
26
Genome-wide Association Studies-GWAS
12.4K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.4K
Obesity
371
The Body Mass Index (BMI) is a numerical value derived from a person's weight and height, used to categorize individuals into weight ranges. It is calculated using the formula: weight in kilograms divided by height in meters squared. Obesity is a health condition characterized by excessive accumulation of adipose tissue that poses health risks, often diagnosed with a BMI ≥ 30. This excess fat storage occurs when surplus dietary calories are converted into triglycerides and stored in...
371
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K


