对频率学,贝叶斯学和机器学习模型进行比较分析,以预测SARS-CoV-2 PCR阳性
Francis Chukwuebuka Ihenetu1, Chinyere Ihuarulam Okoro2, Makuochukwu Maryann Ozoude3
1Deapartment of Microbiology, Imo State University, Owerri, Imo, Nigeria.
Frontiers in artificial intelligence
|December 19, 2025
概括
一个随机森林模型使用临床和人口统计数据准确预测了SARS-CoV-2感染状况,超过了后勤回归. 这种方法提供了快速查,特别是当实验室检测有限时.
科学领域:
- 流行病学 流行病学
- 生物统计学 生物统计学
- 机器学习 机器学习
背景情况:
- 准确预测感染状况对于管理SARS-CoV-2等疾病至关重要.
- 传统的SARS-CoV-2诊断方法在敏感性和特异性方面存在局限性.
- 为了改进预测,正在探索先进的统计和机器学习模型.
研究的目的:
- 评估频率主义逻辑回归,贝叶斯逻辑回归和随机森林分类器的预测性能.
- 确定SARS-CoV-2 PCR阳性的关键临床和人口预测因素.
- 为了比较不同建模方法的区分精度.
主要方法:
- 对950名参与者的分析使用频率主义后勤回归,贝叶斯主义后勤回归和随机森林分类器.
- 在随机森林模型中应用合成少数群体过量采样技术 (SMOTE) 对类不平衡.
- 预测因素的评估包括IgG血清状况,旅行史,症状,性别和年龄;通过AUC评估的表现.
主要成果:
- 随机森林分类器实现了最高的区分性能,其AUC为0.947-0.963.
- 频率主义后勤回归确定了国际旅行,嗅觉丧失和国内旅行作为重要的预测因素.
- 在随机森林模型中,年龄和性别是显著的,这表明了潜在的非线性影响.
结论:
- 机器学习,特别是随机森林方法,与物流回归模型相比,显示出更高的预测准确性.
- 贝叶斯回归为关键预测因素提供了可靠的估计和量化的不确定性.
- 常规收集的症状和暴露数据可以促进快速,资源高效的SARS-CoV-2查.
相关概念视频
Steps in Outbreak Investigation
468
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
468
Sensitivity, Specificity, and Predicted Value
1.2K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.2K
Statistical Methods for Analyzing Epidemiological Data
864
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
864
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
425
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
425
Comparing the Survival Analysis of Two or More Groups
533
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
533
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
461
Drug disposition in the body is a complex process and can be studied using two major approaches: the model and the model-independent approaches.
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
461


