使用可解读性的特征识别 机器学习 预测疾病风险因素 佛罗里达州南部COVID-19患者的病情严重性
Debarshi Datta1, Subhosit Ray1, Laurie Martinez1
1Christine E. Lynn College of Nursing, Florida Atlantic University, Boca Raton, FL 33431, USA.
Diagnostics (Basel, Switzerland)
|September 14, 2024
概括
一个人工智能系统确定了预测COVID-19患者严重程度的关键因素. 年龄,糖尿病等并发症和人口统计数据对于预测重症监护室 (ICU) 住院需求至关重要.
科学领域:
- 医疗保健中的人工智能
- 临床信息学 临床信息学
- 流行病学 流行病学
背景情况:
- COVID-19患者需要不同级别的护理,从中等护理单位 (IMCUs) 到重症监护室 (ICU),有些人需要机械通风 (MV).
- 预测患者的发展轨迹对于资源分配和及时治疗至关重要,特别是在激增期间.
研究的目的:
- 开发人工智能驱动的决策支持系统,以确定预测COVID-19患者疾病严重程度的关键特征.
- 预测需要IMCU,ICU或ICU与MV入院的需求.
主要方法:
- 分析了佛罗里达州南部5371名COVID-19患者的电子健康记录 (eHR) 数据.
- 使用随机森林分类器与SMOTE数据增强.
- 采用SHAP,MDI和变换 对于模型可解释性和特征分析的重要性.
主要成果:
- 所有入院级别 (ICU with MV,ICU,IMCU) 的关键预测因素包括年龄,种族,性别,BMI,腹,糖尿病,高血压,早期病和肺炎.
- 老年人,男性,吸烟者和BMI较高的人面临严重性风险增加.
- 伴随性疾病之间的相互作用,如腹和糖尿病,加剧了疾病的严重程度.
结论:
- 社会人口统计学特征和医院前并发症是COVID-19患者结果的重要预测因素.
- 模型的解释性,包括特征相互作用,为患者急剧增长期间的快速治疗计划提供了关键的见解.
相关概念视频
Steps in Outbreak Investigation
114
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
114
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Statistical Methods for Analyzing Epidemiological Data
324
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
324
Sensitivity, Specificity, and Predicted Value
220
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
220
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
124
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
124
Receiver Operating Characteristic Plot
101
A ROC (Receiver Operating Characteristic) plot is a graphical tool used to assess the performance of a binary classification model by illustrating the trade-off between sensitivity (true positive rate) and specificity (false positive rate). By plotting sensitivity against 1 - specificity across various threshold settings, the ROC curve shows how well the model distinguishes between classes, with a curve closer to the top-left corner indicating a more accurate model. The area under the ROC curve...
101


