BioInstruct:用于生物医学自然语言处理的大型语言模型的指令调整
Hieu Tran1, Zhichao Yang1, Zonghai Yao1
1Manning College of Information and Computer Sciences, University of Massachusetts Amherst, Amherst, MA 01003, United States.
概括
使用BioInstruct数据集调整大型语言模型 (LLM) 的指令显著提高了生物医学自然语言处理 (BioNLP) 的性能. 这种特定领域的方法增强了回答问题,提取信息和生成文本的任务.
科学领域:
- 生物医学自然语言处理 (BioNLP)
- 人工智能在医学中的应用
背景情况:
- 大型语言模型 (LLM) 在生物医学自然语言处理 (BioNLP) 中表现有前途.
- 特定领域的微调对于优化医学等专业领域的LLM绩效至关重要.
研究的目的:
- 引入BioInstruct数据集,用于调整指令的LLMs.
- 评估特定领域指令调整对BioNLP任务的影响.
- 探索指令调整和多任务学习原则之间的协同作用.
主要方法:
- 开发了BioInstruct,一个包含25005个LLM指令的数据集 (LLaMA 1和2).
- 利用GPT-4生成基于人类精选样本的指令.
- 员工低级调整 (LoRA) 对于高效的微调.
- 评估了针对问题的答案 (QA),信息提取 (IE) 和文本生成 (GEN) 任务的指令调整的LLM.
主要成果:
- 调整指令的LLM实现了显著的性能增长:质量保证准确率为17.3%,IE F1得分为5.7%,GEN GPT-4得分为96%.
- 与其他特定领域的LLMs相比,7B参数指令调整的LLaMA 1模型表现出具有竞争力或优越的性能.
- 当指令微调涉及密切相关的任务时,性能改善显著更高,这表明多任务学习协同作用.
结论:
- 生物Instruct数据集是推动生物NLP发展的宝贵资源.
- 调整指令的LLM代表了高性能BioNLP应用的最新技术.
- 指令调整和多任务学习之间的协同作用增强了生物医学领域的LLM能力.
相关概念视频
Improving Translational Accuracy
10.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.1K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Nonsense-mediated mRNA Decay
10.6K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.6K


