在药理流行病学中,自然语言处理衍生的边际混影响与结构化特征的比较排名
Joseph M Plasek1, Richard D Wyss2, Janick G Weberpals2
1Division of General Internal Medicine, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
在药理流行病学研究中,自然语言处理 (NLP) 方法可以识别比单独的索赔代码更多的混信息. NLP工具有效地产生了许多用于高维调整的功能,增强了因果效应估计.
科学领域:
- 药学流行病学 药学流行病学
- 医疗信息学 医疗信息学
- 计算语言学 计算语言学
背景情况:
- 传统的药物流行病学依赖于索赔代码,这可能会错过关键的混信息.
- 电子健康记录 (EHR) 包含丰富的临床笔记,这些笔记往往未被充分用于混者识别.
研究的目的:
- 评估自然语言处理 (NLP) 方法在识别混信息方面的能力,超出了单独从索赔代码中获得的信息.
- 评估各种NLP技术的实用性,以提高药理学流行病学中的混杂物识别.
主要方法:
- 一项回顾性队列研究将Medicare索赔和EHR数据与患有胃潰瘍疾病或骨关节炎的患者联系起来.
- 现成的NLP工具 (例如,袋子-n-grams,LDA,BERT,GloVe) 处理了临床笔记以提取特征.
- 候选混因子特征使用Bross公式来排名,以评估它们的边际混影响.
主要成果:
- 从NLP衍生的特征构成了排名最高的混因子的很大一部分,在胃潰瘍疾病的前5000个特征中,从39%到93%不等.
- 在不同结局 (胃肠道出血,急性损伤) 的骨关节炎患者中观察到类似的NLP衍生的特征的高百分比.
- 以语言学为重点的工具和嵌入在识别中等级特征排名中的有影响力的混因素方面占据了突出地位.
结论:
- 通过识别大量额外的混特征,NLP方法显著补充了索赔数据.
- 无监督的,现成的NLP工具是可扩展的,用于生成适合在药理流行病学中进行高维代理调整的特性.
- 这些发现支持将NLP纳入药理流行病学,以提高因果效应估计的准确性.
更多相关视频
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
09:35A Protocol for Using Gene Set Enrichment Analysis to Identify the Appropriate Animal Model for Translational Research
Published on: August 16, 2017
相关概念视频
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Analysis of Population Pharmacokinetic Data
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Mechanistic Models: Compartment Models in Individual and Population Analysis
