通过对大型语言模型的指令调整,在生物医学中推进实体识别
Vipina K Keloth1, Yan Hu2, Qianqian Xie1
1Section of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT-06510, United States.
Bioinformatics (Oxford, England)
|March 21, 2024
概括
我们开发了一种新方法,使用大型语言模型 (LLM) 改进生物医学命名实体识别 (NER). 我们的方法将NER转化为生成任务,超越现有模型并展示LLMs.
科学领域:
- 计算语言学计算语言学
- 生物信息学是一种生物信息学.
- 医疗保健中的人工智能
背景情况:
- 大型语言模型 (LLM) 在医疗保健方面表现有前途,但在生物医学命名实体识别 (NER) 上扎.
- 生物医学NER通常是一个序列标记任务,这种格式对于LLMs来说不是最佳的.
- 与生物医学NER专业化,微调模型相比,现有的LLM表现不佳.
研究的目的:
- 开发一个基于指令的学习范式,以适应生物医学NER的LLMs.
- 将生物医学NER任务从序列标记转变为生成任务.
- 评估LLM在生物医学NER任务上的表现,使用重用数据集.
主要方法:
- 为生物医学NER开发了一个端到端的基于指令的学习范式.
- 重新利用现有的生物医学NER数据集进行培训和评估.
- 使用LLaMA-7B作为基础的LLM,创建了BioNER-LLaMA.
主要成果:
- 在各种生物医学NER数据集上,BioNER-LLaMA获得了比GPT-4更高的F1分数 (530%).
- 证明通用领域的LLM可以匹配微调的,域特定的模型,如PubMedBERT.
- 与生物医学特定的PMC-LLaMA模型相比,展示了竞争性表现.
结论:
- 拟议的范式有效地将通用领域的LLM适应于高性能生物医学NER.
- 这种方法使LLM能够在多任务,多领域的生物医学应用中与最先进的性能竞争.
- 该研究强调了生成LLM在复杂的生物医学NLP任务中的潜力.
相关概念视频
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K


