Related Experiment Video
Updated: Sep 10, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
CeRTS: certainty retrieval token search in large language model clinical information extraction.
Lars E Schimmelpfennig1, Kriti Bhattarai2, Inez Y Oh1
1Institute for Informatics, Data Science & Biostatistics, Washington University School of Medicine, St. Louis, MO, United States.
Certainty Retrieval Token Search (CeRTS) improves large language model (LLM) uncertainty estimation for clinical data extraction. This new method offers more reliable confidence scores for extracted information, enhancing LLM viability in healthcare.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Clinical Informatics
Background:
- Large language models (LLMs) require robust uncertainty estimation for clinical applications, particularly in extracting information from electronic health records.
- Existing token-level uncertainty estimators are limited as they only consider probabilities within a single output sequence.
Purpose of the Study:
- To develop and evaluate Certainty Retrieval Token Search (CeRTS), a novel uncertainty estimator for structured information extraction using LLMs.
- To leverage JSON output structure constraints to consider all likely sequences and their probabilities for a more accurate measure of model confidence.
Main Methods:
- CeRTS was evaluated against a gold-standard uncertainty estimator on eight open-source LLMs.
- Clinical features were extracted from lung cancer discharge summaries.
- Performance was quantified using calibration (Brier score) and discrimination (AUROC).
Main Results:
- CeRTS demonstrated superior discriminatory power compared to the previous estimator across all evaluated LLMs.
- CeRTS achieved better calibration in most cases, indicating improved reliability.
- The strongest agreement between model confidence and accuracy was observed with the Qwen-2.5 model.
Conclusions:
- CeRTS provides well-calibrated confidence scores for LLM-based information extraction from clinical text, offering researchers a quantitative reliability measure.
- While generally robust, CeRTS faced challenges with the DeepSeek-R1 model, potentially due to its Chain-of-Thought reasoning.
- The CeRTS method is applicable to any domain necessitating reliable uncertainty estimation beyond clinical data.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Improving Translational Accuracy
Clinical Trials: Overview
Uncertainty: Confidence Intervals
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...