利用机器学习方法来预测潜在的莱姆病病例和美国的发病率,使用Twitter
Srikanth Boligarla1, Elda Kokoè Elolo Laison2, Jiaxin Li1
1Harvard Extension School, Harvard University, Cambridge, USA.
BMC medical informatics and decision making
|October 16, 2023
概括
这项研究使用了Twitter数据来追踪美国的莱姆病,发现推特和病例数量之间存在很强的相关性. 通过社交媒体监控,BERTweet证明有效地识别了潜在的莱姆病病例.
科学领域:
- 计算流行病学计算流行病学
- 公共卫生监督是对公共卫生的监督.
- 自然语言处理 (NLP) 是一种自然语言处理.
背景情况:
- 莱姆氏病是美国流行的载体传播疾病之一.
- 准确的发病率对于公共卫生管理至关重要.
研究的目的:
- 使用Twitter数据预测潜在的莱姆病病例.
- 评估美国莱姆病发病率.
- 评估社交媒体作为公共卫生监测工具.
主要方法:
- 收集并预处理了130万条推文,策划了77,500个标记数据集.
- 训练并测试NLP和机器学习模型 (TF-IDF,Word2vec,BERTweet) 用于推文分类.
- 分析了10年来与莱姆病相关的推特的时空模式.
主要成果:
- 在识别莱姆病推特时,BERTweet获得了最高的准确性和F1得分.
- 一个一致的模式显示,美国西部和东北部的推特率更高.
- 机密推特数量与官方莱姆病数量有很强的相关性.
结论:
- 推特数据可以成为美国莱姆病的宝贵监测工具.
- BERTweet是一个可靠的NLP分类器,用于检测相关的莱姆病推文.
- 社交媒体在提高莱姆病的认识方面发挥了作用.
相关概念视频
Steps in Outbreak Investigation
135
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
135
Statistical Methods for Analyzing Epidemiological Data
385
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
385
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Principles of Disease Surveillance
105
Disease surveillance is the systematic collection, analysis, and interpretation of health data essential to the planning, implementation, and evaluation of public health practice. This process integrates data dissemination to entities responsible for preventing and controlling disease, injury, and disability. Surveillance systems provide crucial information for action, helping public health authorities make informed decisions to manage and prevent outbreaks, ensure public safety, optimize...
105


