对生物医学研究中的预测建模进行大型语言模型的基准测试,重点是生殖健康
Reuben Sarwal1, Victor Tarca2, Claire A Dubin1
1Bakar Computational Health Sciences Institute, University of California, San Francisco, San Francisco, CA 94158, USA.
Cell reports. Medicine
|February 18, 2026
概括
大型语言模型 (LLM) 在omics数据分析中表现有前途. 在预测任务中,LLM生成的代码与人类的性能相匹配或超过,使复杂的建模民主化.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 大型语言模型 (LLM) 是代码生成和数据分析的新兴工具.
- 在omics数据中的预测建模提出了复杂的挑战.
研究的目的:
- 评估各种LLM在使用omics数据完成预测任务方面的表现.
- 评估LLM生成的生物预测代码的准确性.
主要方法:
- 在四个DREAM挑战任务中,LLM被提示提供任务描述,数据位置和目标结果.
- 由LLM生成的R和Python代码被执行以适应预测模型.
- 模型的准确性是通过测试组来确定妊娠年龄回归和早产分类的模型准确性.
主要成果:
- 在8个测试的LLM中,有4个 (o3-mini-high,4o,DeepseekR1,Gemini 2.0) 成功完成了至少一个任务.
- R代码生成 (14/16任务) 比Python (7/16任务) 更成功.
- 开放AI的o3-mini-high表现最好,完成了7/8任务.
- 顶级LLM生成的模型实现了与中位数的人类团队相匹配或超过的性能,并在一个任务中超越了顶级的人类团队.
结论:
- 在欧米克学研究中,LLM显示了民主化预测建模的巨大潜力.
- 与人类开发的模型相比,LLM生成的代码可以实现竞争性或优异的性能.
- 这些发现表明,LLM可以提高生物信息学研究成果和可访问性.
相关概念视频
Mechanistic Models: Compartment Models in Individual and Population Analysis
288
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
288
Regression Toward the Mean
7.2K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
7.2K
Analysis of Population Pharmacokinetic Data
829
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
829
Pharmacokinetic Models: Comparison and Selection Criterion
392
Physiological and compartmental models are valuable tools used in studying biological systems. These models rely on differential equations to maintain mass balance within the system, ensuring an accurate representation of the dynamic processes at play.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
392
Overview of Biostatistics in Health Sciences
5.4K
Biostatistics involves the application of statistical techniques to scientific research in health-related fields, including biology and public health. These techniques are essential for designing studies, collecting data, and analyzing it to draw meaningful conclusions. Given the complexity of biological processes, particularly in studies involving human subjects, biostatistical methods are crucial for effectively organizing and interpreting data that might otherwise obscure underlying patterns...
5.4K
Improving Translational Accuracy
3.7K
3.7K


