对大型语言模型进行全面评估,对生物医学文本处理任务进行基准测试
Israt Jahan1, Md Tahmid Rahman Laskar2, Chun Peng3
1Department of Biology, York University, Canada; Information Retrieval and Knowledge Management Research Lab, York University, Canada.
Computers in biology and medicine
|March 6, 2024
概括
大型语言模型 (LLM) 在生物医学任务中表现有前途,在没有特定任务培训的情况下,在小数据集上表现优于微调模型. 性能因任务而异,但在大型注释数据稀缺的情况下,LLM提供了潜力.
科学领域:
- 生物医学信息学 生物医学信息学
- 人工智能的人工智能
背景情况:
- 大型语言模型 (LLM) 已经显示出广泛的任务适用性.
- 它们在专门的生物医学领域的表现仍然在很大程度上未被探索.
研究的目的:
- 综合评估和比较流行的LLM在各种生物医学任务上的能力.
- 根据既有模型评估LLM的表现,特别是在数据稀缺的情况下.
主要方法:
- 在6个生物医学任务和26个数据集中对4个LLM进行了全面评估.
- 对比零射击LLM性能与微调的最先进模型在较小的数据集上.
主要成果:
- 在小型生物医学数据集的零射击设置中,LLM表现出强的性能,有时超过了微调模型.
- 没有一个LLM在所有任务中始终表现优于其他人;绩效取决于任务.
- 总体而言,LLM的表现落后于在广泛的数据集上微调的模型.
结论:
- 在大型机构的预培训为LLM提供了重要的生物医学领域专业化.
- 在有限的注释数据的情况下,LLM为生物医学应用提供了有价值的潜在工具.
- 需要进一步的研究来优化LLM的性能,以应对复杂的生物医学挑战.
相关概念视频
Improving Translational Accuracy
10.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.4K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K


