使用BERT语言模型,对欧洲公开评估报告和EMA指南之间的语义关系进行完整的文档分析
Erik Bergman1, Anna Maria Gerdina Pasmooij2, Peter G M Mol2,3
1Swedish Medical Products Agency, Uppsala, Sweden.
PloS one
|December 15, 2023
概括
这项研究使用文本嵌入来比较欧洲公共评估报告 (EPAR) 与欧洲药物局 (EMA) 的指导方针. 抗病毒和抗出血药物与指南具有更大的语义相似性,这表明了加强监管支持的潜在领域.
科学领域:
- 药物监督和监管科学 药物监督和监管科学
- 计算语言学 计算语言学
- 药物开发 药物开发
背景情况:
- 欧洲药物管理局 (EMA) 提供科学指导方针,以确保药物的有效性和安全性.
- 欧洲公开评估报告 (EPAR) 记录了在欧盟申请营销许可的药物的评估.
- 了解指导方针与评估报告之间的协调对于监管审查至关重要.
研究的目的:
- 通过使用文本嵌入来调查EMA指南和EPAR之间的语义相似性.
- 为了确定治疗领域与指导方针和批准的药物之间显著的语义对齐或分歧.
- 探索计算方法在分析监管文件中的潜力.
主要方法:
- 一组1024份EPAR和669份EMA指南 (2008-2022) 被处理成文本块.
- 句子 BERT语言模型为文本块生成了嵌入.
- 一个零碎匹配算法计算了文档之间的语义距离.
- 线性回归分析了与产品特征相比的文档距离得分.
主要成果:
- 在EPAR与抗病毒药物指南 (J05) 和抗出血药物指南 (B02) 之间观察到统计学上显著较低的语义距离.
- 这一发现在调整产品年龄和EPAR长度后仍然存在.
- 结果表明,在特定的治疗领域,监管评估与现有科学指南密切一致.
结论:
- 文本嵌入方法有效量化了监管文件之间的语义关系.
- 这些发现提供了关于EMA指导方针和监管审查流程之间的相互作用的见解.
- 这种方法可以帮助确定可能受益于更新或附加监管指导的治疗领域.
更多相关视频
相关概念视频
Stereotype Content Model
14.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.7K
Protein Folding Quality Check in the RER
3.7K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
3.7K


