Related Experiment Video
Updated: Jun 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Lexical hints of accuracy in LLM reasoning chains
Arne Vanhoyweghen1, Brecht Verbeken2, Andres Algaba2
1Data Analytics Lab, Vrije Universiteit Brussel, Brussels, 1050, Belgium. arne.vanhoyweghen@vub.be.
Researchers found that lexical uncertainty cues within a Large Language Model's (LLM) Chain-of-Thought (CoT) reasoning are reliable indicators of confidence. These signals help detect incorrect answers, improving model calibration.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Machine Learning
Background:
- Large Language Models (LLMs) fine-tuned with reinforcement learning and Chain-of-Thought (CoT) reasoning show improved performance.
- However, LLMs exhibit poor calibration, often expressing high confidence in incorrect answers on challenging tasks like Humanity's Last Exam (HLE).
Purpose of the Study:
- To investigate if measurable properties of CoT reasoning can serve as reliable, model-internal confidence signals.
- To identify which CoT features best predict model accuracy and calibration.
Main Methods:
- Analysis of three feature classes in CoT: length, sentiment volatility, and lexicographic markers (including hedging terms).
- Evaluation on diverse benchmarks: Humanity's Last Exam (HLE), Omni-MATH, and GPQA-diamond.
- Utilized models: DeepSeek-R1, Claude 3.7 Sonnet, and Qwen-235B-Think.
Main Results:
- Lexical uncertainty cues (e.g., "guess", "stuck", "hard") are highly informative confidence indicators across benchmarks.
- Sentiment shifts offer a weaker, complementary signal.
- CoT length predicts correctness only on intermediate-difficulty benchmarks (Omni-MATH, GPQA), not on harder tasks (HLE).
- Uncertainty signals are more salient than high-confidence markers for detecting errors.
Conclusions:
- Measurable properties of CoT reasoning, particularly lexical uncertainty, provide reliable internal confidence signals for LLMs.
- These findings support a lightweight post-hoc calibration method to enhance LLM reliability by complementing self-reported probabilities.
Related Concept Videos
Deductive Reasoning
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Inductive Reasoning
The Anchoring-and-Adjustment Heuristic
Formal Charges
Hindsight Biases