Related Experiment Video
Updated: Jul 8, 2025

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Quantifying confidence shifts in a BERT-based question answering system evaluated on perturbed instances
1Information Sciences Institute, University of Southern California, Marina del Rey, California, United States of America.
Transformer models excel at multiple-choice natural language processing (NLP) tasks. However, their confidence in ambiguous situations, like incorrect answer choices, differs significantly from expected behavior, necessitating improved testing.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Machine Learning
Background:
- Transformer-based neural networks show significant progress in multiple-choice natural language processing (NLP) tasks, including Question Answering (QA).
- Systematic evaluation of these models in ambiguous scenarios, where no correct answer may be present, remains underexplored despite real-world relevance.
Purpose of the Study:
- To experimentally evaluate transformer-based QA models in ambiguous situations.
- To investigate how models perform when presented with prompts lacking correct answer choices.
Main Methods:
- Designed three probes to systematically 'confuse' QA instances by introducing perturbations.
- Conducted experiments using an established transformer-based multiple-choice QA system.
- Evaluated performance on two benchmark datasets.
Main Results:
- Transformer models exhibit confidence levels that diverge from expected behavior in ambiguous QA scenarios.
- Model confidence in incorrect choices suggests a lack of certainty in distinguishing correct from incorrect options.
- High performance on idealized QA tasks does not reliably predict performance on ambiguous instances.
Conclusions:
- Current transformer-based QA models may struggle with ambiguous situations, performing differently than expected.
- The models' inability to reliably distinguish ambiguous from unambiguous contexts highlights limitations.
- Enhanced testing protocols and benchmarking are crucial before deploying these models in critical applications with human impact.
More Related Videos
Related Concept Videos
Uncertainty: Confidence Intervals
Confidence Coefficient
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Uncertainty in Measurement: Accuracy and Precision
Confidence Intervals
A...
Propagation of Uncertainty from Systematic Error

