迅速なエンジニアリングによるmiRNA情報抽出のための大規模な言語モデルの改善
Rongrong Wu1, Hui Zong2, Erman Wu3
1Department of Urology and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital, Sichuan University, Chengdu, China; Operation Management Department, The First Affiliated Hospital of Soochow University, Suzhou, China.
Computer methods and programs in biomedicine
|August 28, 2025
まとめ
大型言語モデル (LLM) は限られたmiRNA抽出能力を示しますが,プロンプトエンジニアリングはパフォーマンスを向上させます. バイオメディカル発見とバイオマーカーの特定のために,さらにLLMの精錬が必要である.
科学分野:
- 生物医学情報学
- 計算生物学
- 医療における人工知能
背景:
- 大規模な言語モデル (LLM) は,生物医学知識発見の可能性を秘めています.
- マイクロRNA (miRNA) などの微細な生物学的情報を抽出することは,病気のメカニズムを理解し,バイオマーカーを特定するために不可欠です.
- miRNA情報抽出におけるLLMの性能は,包括的な評価を必要とする.
研究 の 目的:
- miRNA情報抽出におけるLLMの能力を評価する.
- LLMの成績に対する様々な迅速な学習戦略の影響を評価する.
- miRNAデータ抽出における従来の方法と比較してLLMのパフォーマンスを比較する.
主な方法:
- 3つの高品質のmiRNA情報抽出データセット (Re-Tex,Re-miR,miR-Cancer) をベンチマークとトレーニングのために構築.
- 3つのLLM (GPT-4o,Gemini,Claude) の評価は,ベースライン,5ショットチェーンの思考と生成された知識プロンプトを使用しています.
- 伝統的な計算モデルとLLMのパフォーマンスの比較.
主要な成果:
- エンティティ抽出のパフォーマンスを大幅に改善しました.
- 生成された知識の誘導は,最も高いF1スコア (76.6%のエンティティ, 54.8%の関係抽出) をもたらしました.
- GPT-4oはジェミニとクロードを上回り,miRNAエンティティの認識は最高で,遺伝子/タンパク質の認識は最低であり,LLMは従来の方法を上回らなかった.
結論:
- 情報抽出と知識発見のために,高品質のmiRNAデータセットが確立されました.
- miRNA抽出におけるLLMの性能は限られているが,迅速な最適化は能力を高める.
- 診断と治療の標的の発見を加速するには,LLMのさらなる精錬が必要である.
関連する概念動画
MicroRNAs
3.1K
MicroRNA (miRNA) are short, regulatory RNA transcribed from introns (non-coding regions of a gene) or intergenic regions (stretches of DNA present between genes). Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself, forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA...
3.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K


