Related Experiment Videos
Beyond Accuracy: A Mixed-Methods Audit of Chain-of-Thought Failures in LLM-Based COVID-19 Vaccine Stance Detection
Andreas Praschk1, Valentin Fischill-Neudeck2, Thomas Caspari3
1Master Programme Public Health, Center for Public Health and Healthcare Research, Paracelsus Medical Private University, Strubergasse 21, Salzburg, 5020, Austria.
Abstract:
This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring's content analysis guided by the FUTURE-AI framework. At each model's best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008-0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label-rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label-rationale coherence). CoT availability, label-rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.
Related Concept Videos
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Random and Systematic Errors
Random and Systematic Errors
Propagation of Uncertainty from Systematic Error
Bias in Epidemiological Studies