从变压器 (BERT) 来进行双向编码器表示预训练的超样本效应,以定位医疗BERT并增强生物医学BERT
Shoya Wada1, Toshihiro Takeda1, Katsuki Okada1
1Department of Medical Informatics, Osaka University Graduate School of Medicine, Japan.
Artificial intelligence in medicine
|May 10, 2024
概括
过度采样特定领域的 corpora 并以平衡的方式预训练来自变压器 (BERT) 模型的双向编码器表示,可显著提高专门语言任务的性能. 这种方法增强了在具有有限高质量数据的领域的信息提取.
科学领域:
- 自然语言处理自然语言处理.
- 机器学习 机器学习
- 生物信息学是一种生物信息学.
背景情况:
- 大规模的神经语言模型增强了NLP中的转移学习.
- 像BERT这样的基于变压器的模型可以改善信息提取,但在数据稀缺的领域却存在困难.
- 培训特定领域的BERT模型是具有挑战性的,因为有限的高质量,大规模的数据库.
研究的目的:
- 为了应对在数据有限的领域培训高性能BERT模型的挑战.
- 为了研究过量采样域特定体的有效性进行预训练.
- 为特定领域的BERT模型开发和评估一种新的预培训方法.
主要方法:
- 同时预训练模型,使用来自不同领域的过量采样知识.
- 英语生物医学BERT从一个小的语料库的开发.
- 从一个小的语料库中创建日本医学BERT.
- 使用平衡的PubMed摘要进行增强的生物医学BERT预训.
主要成果:
- 英语BERT在生物医学语言理解评估 (BLUE) 基准上取得了实际表现.
- 拟议的方法在相似大小的生物医学体中表现优于常规方法.
- 在大多数医疗任务上,日本医疗BERT超越了传统模型.
- 增强的生物医学BERT在BLUE基准上显示出优异的临床和生物医学分数.
结论:
- 均衡的预训练与过量采样,适合任务的 corpora 能够实现高性能 BERT 模型构建.
- 提出的方法为开发有效的特定领域语言模型提供了可行的解决方案.
- 这种方法对于数据有限的专业领域尤其有利.
更多相关视频
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


