与结缔组织疾病相关的间歇性肺病再入院风险和全因死亡率:可解释的机器学习方法
Boyi Chen1, Xixian Hu2,3, Xuefei Shi1,2,3
1Department of Respiratory Medicine, Affiliated Huzhou Hospital, Zhejiang University School of Medicine, Huzhou, People's Republic of China.
Chronic respiratory disease
|December 22, 2025
概括
机器学习识别了血液生物标志物,包括低白蛋白和升高的CA125,CYFRA 21-1和中性粒细胞与淋巴细胞的比率 (NLR),以预测带有间歇性肺病 (CTD-ILD) 患者的结缔组织疾病的医院再入院和死亡率.
科学领域:
- 生物标志物 生物标志物
- 机器学习 机器学习
- 自免疫性疾病 自免疫性疾病
背景情况:
- 连接组织疾病 (CTD) 是一组自身免疫性疾病.
- 间歇性肺病 (ILD) 是CTD中最常见的肺部并发症.
- 预测CTD-ILD患者的结果仍然是一个挑战.
研究的目的:
- 使用机器学习识别CTD-ILD的血液生物标志物.
- 评估这些生物标志物与1年再入院的相关性.
- 评估CTD-ILD患者1年全因死亡率的预测值.
主要方法:
- 在210名CTD-ILD患者的数据集上使用了机器学习模型 (逻辑回归,SVM,XGBoost).
- 使用后勤回归和SHAP分析来识别风险因素和解释模型.
- 使用ROC曲线和决策曲线分析 (DCA) 评估模型性能.
主要成果:
- 低白蛋白,高CA125和高CYFRA 21-1是重新入院的重要预测因素.
- 在XGBoost模型显示高效率 (AUC 0.857训练,0.788测试).
- 低白蛋白显著影响模型预测;中性粒细胞与淋巴细胞比率 (NLR) 的升高预测死亡率.
- 一个组合模型 (白蛋白,CA125,CYFRA 21-1,NLR) 在死亡率预测方面达到0.944的AUC.
结论:
- 血液生物标志物包括白蛋白,CA125,CYFRA 21-1和NLR可以预测CTD-ILD的预后.
- 机器学习有效地识别了再接收和死亡率的预测生物标志物.
- 这些发现为改善CTD-ILD的患者管理和风险分层提供了潜力.
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
Steps in Outbreak Investigation
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:

