Related Experiment Videos
Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment
Francesco Veri1, Gustavo Kreia Umbelino1
1Centre for Democracy Studies Aarau, University of Zurich, Aarau 5000, Switzerland.
Abstract:
Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.