ZeroTuneBio NER:使用大型语言模型和提示工程来实现零射击和零调节生物医学实体提取的三阶段框架
Mingyuan Qin1, Lei Feng2, Jing Lu3
1Department of Dermatology, Huashan Hospital, Shanghai Institute of Dermatology, Fudan University, Shanghai, China; Greater Bay Area Institute of Precision Medicine, School of Life Sciences, Fudan University, Shanghai, China.
Computer methods and programs in biomedicine
|September 13, 2025
概括
本研究介绍了ZeroTuneBio NER,一个框架,使大语言模型 (LLM) 能够执行高质量的生物医学命名实体识别 (NER) 无需微调. 这种方法提高了LLM的性能,并减少了对手册注释的依赖.
科学领域:
- 生物医学信息学 生物医学信息学
- 自然语言处理自然语言处理.
- 人工智能的人工智能
背景情况:
- 生物医学实体提取对于知识发现至关重要.
- 大型语言模型 (LLM) 是有前途的,但往往需要大量的微调.
- 生物医学等专业领域的LLM中零射击能力尚未得到充分探索.
研究的目的:
- 为了提高生物医学实体提取的LLM绩效.
- 为了调查零射击命名实体识别 (NER) 没有LLM微调.
- 将拟议的框架与现有模型和人类注释进行比较.
主要方法:
- 开发了一个三阶段的NER框架,ZeroTuneBio NER.
- 该框架集成了思维链推理和快速工程.
- 对疾病,化学和基因数据集进行了评估,没有特定任务的例子或LLM微调.
主要成果:
- 与直接的LLM查询相比,ZeroTuneBio NER实现了0.28的平均F1得分改善.
- 该框架显示F1部分匹配得分约为88%.
- 性能与微调模型相美,在排除严格匹配错误时超越其他模型,同时优化手动注释速度和成本.
结论:
- 在没有微调的情况下,LLM可以实现高质量的NER,减少手动注释需求.
- ZeroTuneBio NER框架扩展了生物医学NER中的LLM应用.
- 该研究强调了可扩展性,并建议了未来的研究方向.
相关概念视频
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K

