对一个不平衡的学习问题的增强策略在一个新的COVID-19严重程度数据集上
Daniel Schaudt1, Reinhold von Schwerin2, Alexander Hafner2
1Department of Computer Science, Ulm University of Applied Science, Albert-Einstein-Allee 55, 89081, Ulm, Baden-Wurttemberg, Germany. daniel.schaudt@thu.de.
Scientific reports
|October 25, 2023
概括
这项研究介绍了一个大型的COVID-19严重性数据集和深度学习模型. 增强策略改善了严重COVID-19病例的表现,有助于未来的临床研究.
科学领域:
- 医疗成像医学成像
- 人工智能的人工智能
- 数据科学数据科学数据科学
背景情况:
- 从胸部X射线检测COVID-19的机器学习模型是常见的.
- 二元分类模型对治疗的影响有限.
- 预测COVID-19严重程度对于量身定制的医疗干预至关重要.
研究的目的:
- 创建和发布最大的公开可用的COVID-19严重程度数据集之一.
- 建立基于深度学习的COVID-19严重程度分类的基准.
- 调查不平衡严重程度数据的增强策略.
主要方法:
- 编制了2358张COVID-19阳性胸部X射线图像的数据集,并进行了严重程度评分 (COVIDx8B数据集).
- 训练并评估深度学习模型用于严重程度分类.
- 实施和测试了针对多数和少数阶级的数据增强技术.
主要成果:
- 新创建的数据集作为COVID-19严重程度分类的基准.
- 增强策略显著提高了严重的COVID-19病例的精度和召回.
- 模型在罕见的,严重的病例中表现得更好,尽管最初的类不平衡.
结论:
- 开发的数据集和模型为进一步研究提供了有价值的起点.
- 增强技术在解决严重疾病预测的阶级不平衡方面是有效的.
- 未来的工作可以优化模型在资源分配和治疗规划中的临床应用.
相关概念视频
Bias in Epidemiological Studies
314
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
314
Improving Translational Accuracy
11.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.4K
Survival Tree
88
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
88
Strategies for Assessing and Addressing Confounding
104
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
104
Steps in Outbreak Investigation
135
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
135
Statistical Methods for Analyzing Epidemiological Data
385
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
385


