Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Validating Radiology Artificial Intelligence Model Performance on Photon-Counting CT Images Using Large Language
Yee Seng Ng1, Mohammed M Kanani2, William E King1
1Department of Radiology, University of Washington, Seattle, Washington.
Large language models (LLMs) automate ground truth label extraction from radiology reports, enabling scalable assessment of artificial intelligence (AI) tools. This method reliably validates AI performance, even with new imaging hardware like photon-counting CT scanners.
Area of Science:
- Radiology
- Artificial Intelligence
- Medical Informatics
Background:
- Radiologic artificial intelligence (AI) tools require continuous monitoring and validation.
- Manual ground truth label extraction from radiology reports is time-consuming and resource-intensive.
- New imaging hardware, such as photon-counting CT (PCCT) scanners, can introduce input drift affecting AI performance.
Purpose of the Study:
- To evaluate the feasibility of using large language models (LLMs) for automated ground truth label extraction from radiology reports.
- To enable scalable assessment and monitoring of radiologic AI tools.
- To validate AI model performance on a new PCCT scanner.
Main Methods:
- Retrospective analysis of four FDA-cleared AI tools for pulmonary embolism, intracranial hemorrhage, cervical spinal fractures, and vertebral compression fractures.
- LLM (Llama 3.3) used to extract binary ground truth labels from radiology reports of PCCT and conventional scanner data.
- Comparison of AI outputs with LLM-extracted labels, with discrepant cases adjudicated by human annotators.
- Interrater reliability measured using Fleiss's κ test; performance metrics recalculated after LLM error correction.
Main Results:
- LLM-extracted labels facilitated rapid AI performance assessment across all four diagnostic tasks.
- No statistically significant performance differences were observed between PCCT and non-PCCT cohorts.
- LLM labels showed strong agreement with final human annotations (κ = 0.731), comparable to interreader agreement (κ = 0.720), confirming LLM labeling reliability.
Conclusions:
- Large language models offer a scalable and efficient automated solution for ground truth label extraction from radiology reports.
- This LLM-based approach supports rapid local validation of AI tools, effectively addressing challenges posed by new imaging hardware and input drift.
More Related Videos
05:49Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
Published on: February 23, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025