预测美国各县的麻疹病例:一种机器学习方法
Stephanie A Kujawski1, Boshu Ru1, Nelson Lee Afanador1
1Merck & Co., Inc. Rahway, NJ, USA.
Vaccine
|September 7, 2024
概括
一个机器学习模型准确地预测了美国大多数县的麻疹疫情. 这种工具可以帮助卫生机构有效预防和控制未来的麻疹病例.
科学领域:
- 流行病学 流行病学
- 公共卫生 公共卫生
- 医疗保健中的机器学习
背景情况:
- 麻疹在2000年被宣布在美国被消灭,但疫情的爆发有所增加.
- 预测麻疹病例的位置对于有效的预防和制战略至关重要.
研究的目的:
- 开发和验证用于预测美国县级麻疹风险的机器学习模型.
- 帮助公共卫生机构识别麻疹预防的高风险地区.
主要方法:
- 使用17个预测变量开发了一个机器学习模型.
- 该模型是根据2014年和2018年的美国县级麻疹病例数据进行训练的.
- 通过将预测与实际2019年麻疹病例数据进行比较来评估模型性能.
主要成果:
- 该模型在预测2019年有或没有麻疹病例的县里实现了95%的特异性和72%的灵敏性.
- 它占2019年所有报告的麻疹病例的94%.
- 该模型正确识别了前30个最高风险县的73%,占所有病例的72%.
结论:
- 开发的机器学习模型准确地预测了大多数面临麻疹高风险的美国县.
- 这种预测框架可以显著支持州和国家卫生机构在预防和制麻疹的努力.
相关概念视频
Steps in Outbreak Investigation
114
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
114
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Statistical Methods for Analyzing Epidemiological Data
330
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
330
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Classification of Leukocytes
1.8K
Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
1.8K
End Point Prediction: Gran Plot
300
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
300


