SmokeBERT:基于BERT的模型,从临床叙述中提取定量吸烟史,以改善肺癌查
Yiming Xue1, Yunzheng Zhu2, Luoting Zhuang2
1Department of Statistics & Data Science, University of California, Los Angeles, CA, USA.
一个新的AI模型,SmokeBERT,准确地从临床笔记中提取详细的吸烟史,改善肺癌查资格识别. 这促进了用于提取关键患者数据的自然语言处理.
科学领域:
- 计算语言学计算语言学
- 医疗信息学医学信息学
- 公共卫生 公共卫生
背景情况:
- 吸烟是癌症和心血管疾病的主要危险因素.
- 在结构化领域中,电子健康记录往往缺乏详细的定量吸烟数据 (例如,包装年).
- 准确的吸烟史对于疾病风险评估和肺癌查 (LCS) 资格至关重要.
研究的目的:
- 开发和评估基于BERT的自然语言处理 (NLP) 模型SmokeBERT,用于从临床叙述中提取详细的定量吸烟史.
- 提高识别符合肺癌查条件的患者的准确性.
主要方法:
- 在临床笔记上微调基于BERT的模型 (SmokeBERT),以提取细粒度的吸烟数据.
- 将SmokeBERT的性能与基于最先进规则的NLP模型进行比较.
- 使用F1分数评估提取精度,并识别符合LCS条件的患者.
主要成果:
- 与基于规则的模型 (0.88) 相比,SmokeBERT在持久测试组中获得了更高的F1得分 (0.97),而不是基于规则的模型 (0.88).
- SmokeBERT显著改善了对符合LCS条件的患者的识别 (例如,在≥20包年内98%与60%相比).
结论:
- SmokeBERT在从临床文本中提取详细的吸烟史方面表现出卓越的性能.
- 这种NLP的进步可以提高肺癌查资格决定的准确性.
- 未来的工作旨在将SmokeBERT的功能扩展到多语言和更大规模的应用程序.
更多相关视频
09:50Impact Assessment of Repeated Exposure of Organotypic 3D Bronchial and Nasal Tissue Culture Models to Whole Cigarette Smoke
Published on: February 12, 2015
10:37Automated Measurement of Pulmonary Emphysema and Small Airway Remodeling in Cigarette Smoke-exposed Mice
Published on: January 16, 2015
相关概念视频
Statistical Methods for Analyzing Epidemiological Data
Observational Studies
There are three types of observational studies – Prospective, retrospective, and cross-sectional.
Prospective Study
Prospective studies, also known as longitudinal or cohort studies, are carried out by collecting future data from groups sharing similar characteristics. One...
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Criteria for Causality: Bradford Hill Criteria - II
Physical Assessment of the Respiratory Tract I: Health History
Subjective Data
Subjective data provides vital information about the patient's health history and symptoms. This data is typically collected through interviews in which patients describe their experiences, symptoms, and concerns.
Health history and...
Longitudinal Research
