Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The inadequacy of offline large language model evaluations: A need to account for personalization in model behavior
Angelina Wang1, Daniel E Ho2, Sanmi Koyejo2
1Cornell Tech, New York, NY, USA.
Abstract:
Standard offline evaluations for language models fail to capture how these models actually behave in practice, where personalization fundamentally alters model behavior. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other questions to their chat interfaces.
Related Concept Videos
Self-Evaluation Maintenance Model
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Survival Tree
Building a Survival Tree
Constructing a...
Stereotype Content Model
