Related Experiment Video
Updated: Sep 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination
1, Berlin, Germany. wilde.k.vera@gmail.com.
Abstract:
The growing use of large language models (LLMs) by participants in online studies threatens the validity of behavioral research data. I propose that Bayesian reasoning problems, which have decades of well-replicated human performance benchmarks, can serve as calibrated detectors of LLM contamination in online samples - an approach I call a capability-gap test. In two preregistered pilots of a Bayesian reasoning training tool (N = 148), participants' accuracy on positive predictive value (PPV) calculation problems reached roughly three times established human performance ceilings. Specifically, 57% of Pilot 2 participants achieved perfect 5/5 scores, against a meta-analytic ceiling of approximately 24% for a single problem presented in natural frequency format, a level at which perfect scores are effectively unattainable. Including two outcome measures - accuracy (correct numerical answer) and Bayesian algorithm use (evidence of the reasoning process) - allows researchers to distinguish LLM contamination from authentic learning effects: accuracy detects contamination, while algorithm use preserves treatment effect signals. This is itself a signal detection problem, structurally analogous to the mass screenings for low-prevalence problems around which the Bayesian reasoning literature was developed. Bayesian reasoning problems offer four advantages as contamination detectors: well-established human performance ceilings from meta-analyses, a large human-LLM performance gap, known reference distributions enabling nuanced assessment, and ease of embedding in existing surveys. I provide practical recommendations for using these problems as data quality diagnostics in online research.
Related Concept Videos
Difference from Background: Limit of Detection
The LOD indicates the presence or absence...
Inductive Reasoning
Quantifying and Rejecting Outliers: The Grubbs Test
Goodness-of-Fit Test
Contaminants and Errors
Another key consideration is determining the appropriate number of samples required to...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...