线下大型语言模型评估的不足:需要考虑模型行为中的个性化
Angelina Wang1, Daniel E Ho2, Sanmi Koyejo2
1Cornell Tech, New York, NY, USA.
Patterns (New York, N.Y.)
|December 31, 2025
概括
离线评估并不反映大型语言模型 (LLM) 的现实性能. 个性化显著影响LLM行为,通过比较标准测试与实际用户的现场评估来证明这一点.
科学领域:
- 人工智能的人工智能
- 自然语言处理自然语言处理.
- 人与计算机的交互
背景情况:
- 大型语言模型 (LLM) 的标准评估是在线进行的.
- 这些离线方法不考虑现实世界的使用,个性化显著影响模型行为.
- 了解线下和在线LLM绩效之间的差异对于准确的评估至关重要.
研究的目的:
- 实证地证明,离线评估如何未能捕捉个性化的法学士的实际行为.
- 将标准线下评估的结果与使用真实用户的现场评估进行比较.
- 为了突出个性化对LLM在现场环境中的表现的影响.
主要方法:
- 与800名真实用户进行现场评估,与ChatGPT和Gemini互动.
- 用户向聊天界面提出了基准和其他问题.
- 将这些现场评估的结果与标准的线下评估指标进行比较.
主要成果:
- 离线评估指标不能准确地预测在现实世界中,个性化的场景LLM的表现.
- 用户互动显示,与线下评估相比,模型行为存在显著差异.
- 个性化被确定为改变LLM响应和有效性的关键因素.
结论:
- 离线评估不足以评估个性化的法学士的真实表现.
- 与真实用户的现场评估提供了一个更准确的衡量LLM在实践中的行为.
- 未来的LLM评估应纳入现实世界的使用模式和个性化效应.
相关概念视频
Self-Evaluation Maintenance Model
248
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...
248
Mechanistic Models: Compartment Models in Individual and Population Analysis
225
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
225
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
255
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
255
Modeling in Therapy
360
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
360
Survival Tree
369
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
369
Stereotype Content Model
15.3K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.3K

