使用大型语言模型在贝叶斯回归分析中建议信息性先前分布.
Michael A Riegler1, Kristoffer H Hellton2, Vajira Thambawita2,3
1Simula Research Laboratory, Oslo, Norway.
Scientific reports
|September 29, 2025
概括
大型语言模型 (LLM) 可以建议贝叶斯回归的信息先前分布,帮助客观分析. 虽然能够识别正确的关联,但校准先前的分布宽度仍然是LLMs的一个挑战.
科学领域:
- 贝叶斯统计学 贝叶斯统计学
- 机器学习 机器学习
- 统计建模 统计建模
背景情况:
- 在贝叶斯回归中选择先前分布是复杂和主观的.
- 现有的引出信息先验的方法资源密集,难以客观地执行.
研究的目的:
- 研究大语言模型 (LLM) 在建议适合贝叶斯回归分析的先前分布方面的潜力.
- 评估不同LLM在产生基于知识和客观信息的先验的表现.
主要方法:
- 开发了一个广泛的提示,让LLMs建议,验证和反思之前的分发.
- 在两个真实世界数据集 (心脏病风险,混凝土强度) 上评估了三个LLM (Claude Opus,Gemini 2.5 pro,ChatGPT 4o-mini).
- 使用Kullback-Leibler差异对最大概率估计器的分布进行评估.
主要成果:
- 在这两个数据集中,LLM成功地建议了两个数据集中的变量之间的关联的正确方向.
- 克劳德和双子座在建议以前的发行版方面总体上表现优于ChatGPT.
- 由LLM建议的中度信息先验通常过于自信,显示与数据的有限一致.
- 克劳德通过不默认为弱信息先验的平均值为0来证明了一个优势,与ChatGPT和Gemini不同.
结论:
- 在贝叶斯回归中,LLM显示出开发高效和客观的信息先前分布的巨大潜力.
- 一个关键的挑战在于校准LLM建议的先验的宽度,因为它们表现出过度自信和缺乏自信的趋势.
- 与双子座和ChatGPT相比,Claude Opus在建议先的方法中表现出了显著的优势.
相关概念视频
Distributions to Estimate Population Parameter
5.0K
The accurate values of population parameters such as population proportion, population mean, and population standard deviation (or variance) are usually unknown. These are fixed values that can only be estimated from the data collected from the samples. The estimates of each of these parameters are sample proportion, the sample mean, and sample standard deviation (or variance). To obtain the values of these sample statistics, data are required that have particular distribution and central...
5.0K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
242
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
242
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
290
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
290
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K


