生物医学研究における予測モデリングのための大規模な言語モデルをベンチマークし,生殖健康に焦点を当てます
Reuben Sarwal1, Victor Tarca2, Claire A Dubin1
1Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA.
Cell reports. Medicine
|February 18, 2026
まとめ
大型言語モデル (LLM) は,オミックスのデータ分析において有望な結果を示しています. LLMで生成されたコードは,複雑なモデリングを民主化し,予測タスクにおける人間のパフォーマンスに匹敵し,またはそれを上回りました.
科学分野:
- 計算生物学とは,計算生物学である.
- バイオインフォマティックス
- ゲノミクスゲノミクスとは
背景:
- 大型言語モデル (LLM) は,コード生成とデータ分析のための新興ツールです.
- オミックスのデータにおける予測モデリングは,複雑な課題を提示します.
研究 の 目的:
- オミックスのデータを用いて予測タスクの完了における様々なLLMのパフォーマンスを評価する.
- 生物学的予測のためのLLMで生成されたコードの精度を評価する.
主な方法:
- LLMは4つのDREAMチャレンジタスクのタスク説明,データ位置,およびターゲットアウトカムを提示されました.
- LLMで生成されたRとPythonのコードは,予測モデルに適合するように実行されました.
- モデルの精度は,妊娠年齢回帰と早産分類のためのテストセットで決定されました.
主要な成果:
- テストされた8人のLLMのうち4人 (o3-mini-high,4o,DeepseekR1,Gemini 2.0) が少なくとも1つのタスクを成功裏に完了しました.
- Rコード生成 (14/16タスク) は,Python (7/16タスク) よりも成功しました.
- OpenAIのo3-mini-highは,7/8のタスクを完了して最高のパフォーマンスを示しました.
- トップのLLMで生成されたモデルは,中間の人間のチームに匹敵する,またはそれを上回るパフォーマンスを達成し,1つのタスクでトップの人間のチームを上回りました.
結論:
- LLMは,オミックスの研究における予測モデリングを民主化するための大きな可能性を示しています.
- LLMで生成されたコードは,人間が開発したモデルと比較して,競争力のあるまたは優れたパフォーマンスを達成することができます.
- これらの発見は,LLMがバイオインフォマティクスにおける研究成果とアクセシビリティを向上させることを示唆しています.
関連する概念動画
Mechanistic Models: Compartment Models in Individual and Population Analysis
288
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
288
Regression Toward the Mean
7.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K
Analysis of Population Pharmacokinetic Data
829
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
829
Pharmacokinetic Models: Comparison and Selection Criterion
392
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
392
Overview of Biostatistics in Health Sciences
5.4K
Biostatistics involves the application of statistical techniques to scientific research in health-related fields, including biology and public health. These techniques are essential for designing studies, collecting data, and analyzing it to draw meaningful conclusions. Given the complexity of biological processes, particularly in studies involving human subjects, biostatistical methods are crucial for effectively organizing and interpreting data that might otherwise obscure underlying patterns...
5.4K
Improving Translational Accuracy
3.7K
3.7K


