基准测试大型语言模型用于生物医学研究中的预测建模,重点是生殖健康
bioRxiv : the preprint server for biology
|August 1, 2025
概括
生成型人工智能和大型语言模型 (LLM) 在计算生物学中显示出对自动化代码生成在omics数据分析中的承诺. 开放AI是一个开放的AI.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 人工智能在基因组学中的应用
背景情况:
- 大型语言模型 (LLM) 在计算生物学中越来越多地用于自动化数据分析代码生成.
- 使用了来自对话反向工程评估和方法 (DREAM) 生殖健康挑战的标准化分子数据集.
- 预测建模任务包括从基因表达,DNA甲基化和微生物组数据的妊娠年龄回归和早产分类.
研究的目的:
- 评估LLM在生成功能性R和Python代码以进行生殖健康omics数据中的预测建模方面的能力.
- 评估跨多种预测任务的LLM绩效,并确定表现最佳的模型.
- 为了比较R与Python代码生成的成功率,并分析影响性能的因素.
主要方法:
- 八个不同的LLM被提示任务描述,数据位置和四个预测任务的目标结果.
- 由LLM生成的R和Python代码被执行以适应模型,应用预测并生成图形.
- 绩效是根据任务完成成功和测试数据集的预测准确度来排名的.
主要成果:
- 四个LLM (o3-mini-high,4o,DeepseekR1,Gemini 2.0) 成功生成了至少一个任务的无错误代码.
- 通过生物导体包的帮助,R代码生成比Python (7/16) 更成功 (14/16任务).
- OpenAI的o3-mini-high表现出卓越的性能,完成7/8任务,测试组的性能与原来的DREAM挑战顶级团队相匹配或超过.
结论:
- 法律学具有很大的潜力,可以增强探索性数据分析和民主化OMICS研究中的预测建模.
- 使用LLM自动化分析管道的关键组件可以显著增加研究成果,特别是标准化公共数据集.
- 针对生物信息学任务量身定制的LLM的进一步开发可以加速生殖健康和其他OMIC领域的发现.
相关概念视频
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Mechanistic Models: Compartment Models in Individual and Population Analysis
87
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
87
Analysis of Population Pharmacokinetic Data
387
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
387
Statistical Methods for Analyzing Epidemiological Data
537
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
537
Comparing the Survival Analysis of Two or More Groups
289
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
289
Steps in Outbreak Investigation
207
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
207


