通过快速工程改进miRNA信息提取的大型语言模型
Rongrong Wu1, Hui Zong2, Erman Wu3
1Department of Urology and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital, Sichuan University, Chengdu, China; Operation Management Department, The First Affiliated Hospital of Soochow University, Suzhou, China.
Computer methods and programs in biomedicine
|August 28, 2025
概括
大型语言模型 (LLM) 显示了有限的miRNA提取能力,但即时工程提高了性能. 生物医学发现和生物标志物识别需要进一步的LLM改进.
科学领域:
- 生物医学信息学
- 计算生物学
- 医疗保健中的人工智能
背景情况:
- 大型语言模型 (LLM) 为生物医学知识的发现提供了潜力.
- 提取细粒度的生物信息,如microRNAs (miRNAs),对于了解疾病机制和识别生物标志物至关重要.
- 在miRNA信息提取中LLM的性能需要全面评估.
研究的目的:
- 评估LLM在miRNA信息提取方面的能力.
- 评估各种快速学习策略对LLM绩效的影响.
- 将LLM的性能与miRNA数据提取中的传统方法进行比较.
主要方法:
- 构建三个高质量的miRNA信息提取数据集 (Re-Tex,Re-miR,miR-Cancer) 用于基准测试和培训.
- 使用基线,五射线思维链和生成知识提示的三个LLM (GPT-4o,Gemini,Claude) 的评估.
- 与传统计算模型的LLM性能比较.
主要成果:
- 优化的提示策略显著改善了实体提取性能.
- 生成的知识提示产生了最高的F1分数 (实体为76.6%,关系提取为54.8%).
- GPT-4o的表现优于Gemini和Claude;miRNA实体识别最高,基因/蛋白质最低;LLM没有超过传统方法.
结论:
- 为信息提取和知识发现建立了高质量的miRNA数据集.
- 在miRNA提取中LLM的性能仍然有限,但快速优化增强了功能.
- 为了加快诊断和治疗目标的发现,需要进一步完善LLM.
相关概念视频
MicroRNAs
3.1K
MicroRNA (miRNA) are short, regulatory RNA transcribed from introns (non-coding regions of a gene) or intergenic regions (stretches of DNA present between genes). Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself, forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA...
3.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K


