用于纵向和层次数据建模的创新统计方法:GMEXGBoost方法
Fariba Asadi1,2, Reza Homayounfar3,4, Yaser Mehrali5
1Ferdows Faculty of Medical Sciences, Birjand University of Medical Sciences, Birjand, Iran.
BMC medical research methodology
|January 7, 2026
概括
一个名为GMEXGBoost的新算法通过结合通用混合效应模型和XGBoost.com来改进复杂医疗数据的分析. 它为具有强相关性的层次和纵向数据集提供了卓越的稳定性和准确性.
科学领域:
- * 计算统计和机器学习.
- * 生物统计和健康数据科学.
背景情况:
- * 医疗保健中的指数级数据增长需要超越传统机器学习的先进分析方法.
- * 传统的算法在医疗保健中常见的相关,纵向和层次数据上扎.
- *一般化混合效应模型 (GLMMs) 和XGBoost在独立用于复杂数据结构时存在局限性.
研究的目的:
- * 介绍GMEXGBoost,这是一个新的算法,通过XGBoost的增强框架扩展GLMM.
- * 开发一种明确结合数据相关性的方法,同时保持预测能力.
- * 评估GMEXGBoost的性能与模拟和现实数据中的现有模型相比.
主要方法:
- * 通过将GLMM固定效应估计与XGBoost的提升和随机效应会计结合起来,开发了GMEXGBoost.
- *使用不同效果结构的模拟和现实世界队列研究来评估性能.
- *与GLMM,GLMMTree,GMERF和XGBoost进行基准测试,使用RStudio中的PMAD,PMCR,灵敏度,特异性,精度和AUC等指标.
主要成果:
- * XGBoost 在大多数场景中显示了最低的平均误差.
- *GMEXGBoost表现出卓越的稳定性和准确性,具有很大的随机效应差异或强烈的相关性.
- *GMEXGBoost在关键绩效指标上的真实数据中表现优于其他模型.
结论:
- * GMEXGBoost有效地结合了GLMM和XGBoost的功能,以提高复杂问题的性能.
- * 该算法为分析具有强相关性的层次和纵向数据集提供了明显的优势.
- *GMEXGBoost是医疗保健和其他领域具有结构化数据的决策的宝贵工具.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
548
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
548
Parametric Survival Analysis: Weibull and Exponential Methods
1.0K
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
1.0K
Longitudinal Studies
469
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
469
Statistical Methods for Analyzing Epidemiological Data
889
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
889
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Friedman Two-way Analysis of Variance by Ranks
478
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
478


