Related Experiment Video
Updated: Sep 14, 2025

An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Quantifying uncert-AI-nty: Testing the accuracy of LLMs' confidence judgments
Trent N Cash1,2, Daniel M Oppenheimer3,4, Sara Christie4
1Department of Social and Decision Sciences, Carnegie Mellon University, 5000 Forbes Ave., 224 Porter Hall, Pittsburgh, PA, 15213, USA. trentncash@gmail.com.
Large Language Model (LLM) chatbots show strong metacognitive accuracy in confidence judgments, comparable to humans. However, LLMs, particularly ChatGPT and Gemini, struggle to adjust confidence based on past performance, revealing a key limitation.
Area of Science:
- Artificial Intelligence
- Cognitive Science
- Human-Computer Interaction
Background:
- Large Language Models (LLMs) like ChatGPT and Gemini are transforming information access.
- Metacognitive confidence judgments are crucial for human uncertainty quantification.
- The accuracy of LLM confidence judgments remains largely unexplored.
Purpose of the Study:
- To investigate the capability of LLMs to quantify uncertainty through confidence judgments.
- To compare the metacognitive accuracy of LLMs and humans across various tasks.
- To identify similarities and differences in confidence judgment strategies between LLMs and humans.
Main Methods:
- Four LLMs (ChatGPT, Bard/Gemini, Sonnet, Haiku) and human participants evaluated their confidence in predictions and answers.
- Studies covered aleatory uncertainty (NFL, Oscar predictions) and epistemic uncertainty (Pictionary, Trivia, university life questions).
- Absolute and relative accuracy of confidence judgments were analyzed.
Main Results:
- LLMs demonstrated comparable, and sometimes superior, absolute and relative metacognitive accuracy to humans.
- Both LLMs and humans exhibited overconfidence in their judgments.
- LLMs, especially ChatGPT and Gemini, often failed to adjust confidence based on prior performance, unlike humans.
Conclusions:
- LLMs possess significant capabilities in metacognitive confidence judgments, approaching human levels of accuracy.
- Overconfidence is a shared trait between LLMs and humans.
- A key limitation for LLMs is their reduced ability to dynamically adjust confidence based on experience, unlike humans.
Related Concept Videos
Uncertainty: Confidence Intervals
Confidence Intervals
A...
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Uncertainty in Measurement: Accuracy and Precision
Confidence Coefficient
Uncertainty: Overview

