Taiyi:一个双语微调的大型语言模型,用于各种生物医学任务
Ling Luo1, Jinzhong Ning1, Yingwen Zhao1
1School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China.
Journal of the American Medical Informatics Association : JAMIA
|February 29, 2024
概括
Taiyi是一个双语大语言模型 (LLM),在各种生物医学自然语言处理 (NLP) 任务上表现出卓越的性能. 这种精心调整的LLM显示了双语生物医学多任务处理的潜力,优于一般的LLM.
科学领域:
- 生物医学自然语言处理 (NLP)
- 医疗保健中的人工智能
- 计算语言学 计算语言学
背景情况:
- 现有的生物医学大语言模型 (LLM) 主要侧重于单语言的问答和对话.
- 需要在各种生物医学NLP任务和语言中评估LLM绩效.
研究的目的:
- 开发和评估Taiyi,一个双语微调的LLM,用于广泛的生物医学NLP任务.
- 对各种生物医学数据集进行监督微调策略的有效性进行评估.
主要方法:
- 在10多种任务类型中策划了140个生物医学文本挖掘数据集 (102个英语,38个中文).
- 将 corpora转换为指令数据,用于对一般LLM进行监督微调.
- 采用了两阶段的微调策略来优化性能.
主要成果:
- 在13个测试组中,Taiyi取得了卓越的表现,包括命名实体识别,关系提取,文本分类和问题解答.
- 在一个案例研究中证明了双语生物医学多任务处理的巨大潜力.
- 在各种生物医学NLP基准上表现优于一般LLM.
结论:
- 高质量的生物医学机构和有效的微调策略显著提高了生物医学领域的LLM绩效.
- 泰通过监督微调表现出双语多任务能力.
- 生成型LLM在信息提取等非生成任务中仍然面临挑战,而歧视型模型仍然占优势.
相关概念视频
Improving Translational Accuracy
10.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.4K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K


