开发和验证一个可解释的预测模型,用于表菌病的血清阳性:一个基于人口的查研究在湖南省,中国
Yu Zhou1, Ling Tang2, Mao Zheng2
1Fudan University School of Public Health, Building 8, 130 Dong'an Road, Shanghai 200032, China; Key Laboratory of Public Health Safety, Fudan University, Ministry of Education, Building 8, 130 Dong'an Road, Shanghai 200032, China; Fudan University Center for Tropical Disease Research, Building 8, 130 Dong'an Road, Shanghai 200032, China.
International journal for parasitology
|December 29, 2025
概括
早期检测杆菌病风险是阻止传播的关键. 一个机器学习模型通过症状,水暴露和村庄因素准确识别有风险的个体,帮助公共卫生干预.
科学领域:
- 流行病学 流行病学
- 机器学习 机器学习
- 公共卫生 公共卫生
背景情况:
- 顺病的传播需要早期识别有风险的人群.
- 预测模型可以帮助针对性干预.
研究的目的:
- 开发和验证一种可解释的机器学习模型,用于早期识别杆菌病风险.
- 将人口,行为和环境因素整合到预测模型中.
主要方法:
- 一个随机森林 (RF) 模型被开发并使用大量队列 (103,707人参加培训, 16,574人参加外部验证) 进行验证.
- 可解释的人工智能技术,包括夏普利添加式扩展 (SHAP),用于解释模型预测.
- 模型性能使用AUC和F1分数进行评估.
主要成果:
- 射频模型在内部验证 (AUC=0.943,F1=0.809) 和外部验证 (AUC=0.897,F1=0.770) 中都取得了高的区分性能.
- 关键预测因素包括杆菌病症状,水暴露史,村庄特有病,性别和村庄风险类别.
- 该模型被翻译成一个实用的工具,用于现实世界的应用.
结论:
- 一种可解释的机器学习模型有效地识别出患杆菌病的高风险个体.
- 该模型的可解释性促进了对风险因素的理解,并支持了实际的公共卫生应用.
- 这种方法有助于将理论建模转化为可行的疾病控制干预措施.
相关概念视频
Sensitivity, Specificity, and Predicted Value
1.2K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.2K
Steps in Outbreak Investigation
466
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
466
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K


