Related Experiment Video
Updated: Feb 4, 2026

07:13
A Two-interval Forced-choice Task for Multisensory Comparisons
Published on: November 9, 2018
11.5K
Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge
1SUNY Empire State University, New York, United States of America.
Plos One
|February 2, 2026
Summary
Tracking large language models (LLMs) reveals divergent evolution. Some models remained stable, others improved, and one degraded, highlighting the challenge of reproducible LLM evaluation.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Machine Learning Evaluation
Background:
- Service large language models (LLMs) evolve rapidly without transparent changelogs, hindering reproducible research and evaluation.
- Assessing LLM performance and stability over time is crucial due to their increasing integration into various applications.
Purpose of the Study:
- To longitudinally track the performance and stability of major service LLM families over time.
- To develop and validate a robust methodology for reproducible LLM evaluation using human and LLM-as-judge assessments.
- To identify patterns of model drift, including stability, improvement, and degradation.
Main Methods:
- A preregistered, human-anchored longitudinal study design with ten weekly evaluation waves.
- Utilized a fixed prompt bank (N=240) across six diverse domains for consistent assessment.
- Employed blinded human raters for correctness judgments and a bias-calibrated LLM-as-judge for pairwise preferences, corrected weekly using a Bradley-Terry model.
- Applied mixed-effects modeling and change-point detection (PELT with MBIC penalty) to identify significant drift patterns.
Main Results:
- Observed divergent stability trajectories across three major LLM families: one consistently stable, one showing improvement, and one exhibiting degradation mid-study.
- Judge calibration enhanced agreement with human judgments (Kendall's τ = 0.59-0.68) and reduced evaluation volatility.
- Safety metrics demonstrated co-variation with detected drift events, suggesting behavioral shifts rather than confirmed causal changes.
Conclusions:
- Service LLMs exhibit distinct and dynamic stability patterns, necessitating continuous and reproducible evaluation frameworks.
- The developed human-anchored, LLM-as-judge methodology offers a scalable and reliable approach for tracking LLM evolution.
- Transparency in LLM development and deployment is essential for fostering trust and enabling accurate performance assessment.
Related Concept Videos
The Anchoring-and-Adjustment Heuristic
7.8K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.8K
Longitudinal Research
13.3K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.3K
Confirmation Biases
8.2K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
8.2K
Hindsight Biases
4.3K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
4.3K
Bias
7.4K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
7.4K
Anchoring Junctions
5.0K
Anchoring junctions are multiprotein complexes that help cells connect to other cells and the extracellular matrix. Anchoring junctions are present on the lateral and basal surfaces of cells, providing strong and flexible connections. Focal adhesions are often formed due to cell interactions with the ECM substrata, which initiate signal transduction via kinase cascades and other mechanisms. Together, they provide stability and tissue integrity. There are three types of anchoring junctions:...
5.0K

