在机器学习和深度神经网络中整合可解释性:在COVID-19症状和疫苗有效性中的特征重要性和异常检测的新方法
Shadi Jacob Khoury1,2, Yazeed Zoabi1,3, Mickey Scheinowitz1,2
1Faculty of Medical and Health Sciences, Tel Aviv University, Tel Aviv 6997801, Israel.
Viruses
|January 8, 2025
概括
结合机器学习和深度学习可解释性的新方法将喉痛确定为一个关键的COVID-19症状. 这一发现有助于早期检测和了解免疫反应,以改善疫苗策略.
科学领域:
- 医疗信息学 医疗信息学
- 机器学习 机器学习
- 流行病学 流行病学
背景情况:
- 可解释机器学习 (ML) 和深度学习 (DL) 模型对于理解复杂的医疗数据至关重要.
- 检测异常值和确定健康数据中的关键预测特征仍然是一个挑战.
研究的目的:
- 开发和应用一个综合的可解释性方法来量化医疗数据中特征的重要性.
- 分析早期的COVID-19流行病数据 (2020年),以确定感染预测的关键症状.
主要方法:
- 从传统的ML和深度神经网络 (DNN) 集成的解释性技术.
- 应用全球和本地解释方法来量化特征的重要性.
- 分析了2020年对COVID-19进行测试的个人自我报告的症状和测试结果的数据集.
主要成果:
- 喉痛,尽管被报告的不到1%的队列,出现了作为一个重要的预测者COVID-19感染.
- 患有喉疼痛的人患入院的可能性更高 (5%),并且可能在感染后免疫反应减弱.
- 这项研究强调了对COVID-19疫苗有效性的潜在影响,以及在某些人群中需要进行强剂注射的需要.
结论:
- 综合解释性方法增强了对COVID-19症状的理解,并有助于检测医学异常值.
- 研究结果表明,喉疼痛可能是COVID-19的早期症状标志物.
- 该研究支持开发可解释模型,以改善COVID-19管理和疫苗优化.
相关概念视频
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Steps in Outbreak Investigation
105
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
105
What Are Outliers?
3.6K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.6K
Sensitivity, Specificity, and Predicted Value
180
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
180
Detection of Gross Error: The Q Test
5.6K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
5.6K
Statistical Methods for Analyzing Epidemiological Data
299
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
299


