Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Deductive Reasoning01:16

Deductive Reasoning

Deductive reasoning, or deduction, is the type of logic used in hypothesis-based science. In deductive reasoning, the pattern of thinking moves in the opposite direction from inductive reasoning. It uses a general principle or law to predict specific results. From these general principles, a scientist can predict specific results that remain valid as long as the general principles are correct.For example, a researcher can make specific predictions from the hypothesis "butterflies are attracted...
Reasoning01:30

Reasoning

Reasoning is the action of thinking about something in a logical, sensible way. It is integral to problem-solving, decision-making, and critical thinking. Reasoning can be inductive or deductive. Reasoning involves transforming information into conclusions, which is essential for problem-solving, decision-making, and critical thinking.
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Inductive Reasoning00:59

Inductive Reasoning

Inductive reasoning is a form of logical thinking that uses related observations to arrive at a general conclusion. It is uncertain and operates in degrees to which the conclusions are credible. As such, inductive arguments can be weak or strong, rather than valid or invalid, and conclusions can be used to formulate testable, falsifiable hypotheses.Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
The Anchoring-and-Adjustment Heuristic01:25

The Anchoring-and-Adjustment Heuristic

In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the $2,000...
Formal Charges02:42

Formal Charges

In some cases, there are seemingly more than one valid Lewis structures for molecules and polyatomic ions. The concept of formal charges can be used to help predict the most appropriate Lewis structure when more than one reasonable structure exists.
Hindsight Biases01:12

Hindsight Biases

Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The relationship between reasoning and performance in large language models-o3 (mini) thinks harder, not longer.

Scientific reports·2026
Same author

Multimode Single-Ring Photonic Molecule.

Physical review letters·2026
Same author

Deep neural networks for inverse design of multimode integrated gratings with simultaneous amplitude and phase control.

Nanophotonics (Berlin, Germany)·2025
Same author

Cascaded-mode interferometers: Spectral shape and linewidth engineering.

Science advances·2025
Same author

Helicity and Polarization Gradient Optical Trapping in Evanescent Fields.

Physical review letters·2023
Same author

Polarization-Dependent Forces and Torques at Resonance in a Microfiber-Microcavity System.

Physical review letters·2023

Related Experiment Video

Updated: Jun 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Lexical hints of accuracy in LLM reasoning chains.

Arne Vanhoyweghen1, Brecht Verbeken2, Andres Algaba2

  • 1Data Analytics Lab, Vrije Universiteit Brussel, Brussels, 1050, Belgium. arne.vanhoyweghen@vub.be.

Scientific Reports
|June 4, 2026
PubMed
Summary

Researchers found that lexical uncertainty cues within a Large Language Model's (LLM) Chain-of-Thought (CoT) reasoning are reliable indicators of confidence. These signals help detect incorrect answers, improving model calibration.

Related Experiment Videos

Last Updated: Jun 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

Area of Science:

  • Artificial Intelligence
  • Natural Language Processing
  • Machine Learning

Background:

  • Large Language Models (LLMs) fine-tuned with reinforcement learning and Chain-of-Thought (CoT) reasoning show improved performance.
  • However, LLMs exhibit poor calibration, often expressing high confidence in incorrect answers on challenging tasks like Humanity's Last Exam (HLE).

Purpose of the Study:

  • To investigate if measurable properties of CoT reasoning can serve as reliable, model-internal confidence signals.
  • To identify which CoT features best predict model accuracy and calibration.

Main Methods:

  • Analysis of three feature classes in CoT: length, sentiment volatility, and lexicographic markers (including hedging terms).
  • Evaluation on diverse benchmarks: Humanity's Last Exam (HLE), Omni-MATH, and GPQA-diamond.
  • Utilized models: DeepSeek-R1, Claude 3.7 Sonnet, and Qwen-235B-Think.

Main Results:

  • Lexical uncertainty cues (e.g., "guess", "stuck", "hard") are highly informative confidence indicators across benchmarks.
  • Sentiment shifts offer a weaker, complementary signal.
  • CoT length predicts correctness only on intermediate-difficulty benchmarks (Omni-MATH, GPQA), not on harder tasks (HLE).
  • Uncertainty signals are more salient than high-confidence markers for detecting errors.

Conclusions:

  • Measurable properties of CoT reasoning, particularly lexical uncertainty, provide reliable internal confidence signals for LLMs.
  • These findings support a lightweight post-hoc calibration method to enhance LLM reliability by complementing self-reported probabilities.