Related Experiment Video
Updated: Aug 13, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious
Thomas Renault1, Antonin Bergeaud2, Clément Bosquet3
1Université Paris-Saclay, Faculté Jean Monnet, Sceaux 92330, France.
Abstract:
Large language models (LLMs) are increasingly used in scientific writing, but their effect on individual productivity is difficult to identify because adoption is rarely directly observed. [K. Kusumegi et al., Science 390, 1240-1243 (2025)] infer adoption from the first paper detected as LLM-assisted and report large productivity gains after adoption. We show that this treatment-timing rule mechanically generates positive event-study dynamics even in the absence of any causal effect. Because high-output months are more likely to produce a detected paper, treatment assignment becomes intrinsically linked to productivity. Using reconstructed arXiv data, we show that random treatment assignments, neutral-keyword triggers, inverted treatment, and pre-ChatGPT placebo periods all generate similar dynamics. Simulations with no treatment effect also reproduce the same posttreatment patterns reported in K. Kusumegi et al., Science 390, 1240-1243 (2025). Our results demonstrate that first-detection timing alone creates spurious evidence of productivity gains from LLM adoption.
Related Concept Videos
Steps in Outbreak Investigation
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...
Mechanistic Models: Compartment Models in Individual and Population Analysis
