Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Computed Tomography01:10

Computed Tomography

7.9K
Tomography refers to imaging by sections. Computed tomography (CT) is a non-invasive imaging technique that uses computers to analyze several cross-sectional X-rays to reveal minute details about structures in the body.
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
7.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Radiology report quality improvement through structured provider feedback: a data-driven initiative.

Abdominal radiology (New York)·2026
Same author

Reply to "Artificial Intelligence in Radiology: The Paradox Between Efficiency and Burnout".

AJR. American journal of roentgenology·2026
Same author

Iterative immunoprecipitation and phage pre-wash dramatically improve epitope-resolved serology by VirScan.

Frontiers in virology (Lausanne, Switzerland)·2026
Same author

Artificial Intelligence Sees the Image, Radiologists See the Patient.

AJR. American journal of roentgenology·2026
Same author

Proximity proteomics reveals a role for IFI16 during human coronavirus infection.

bioRxiv : the preprint server for biology·2026
Same author

Modeling heart rate patterns to quantify neonatal opioid withdrawal syndrome.

Pediatric research·2026

Related Experiment Video

Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000

Validating Radiology Artificial Intelligence Model Performance on Photon-Counting CT Images Using Large Language

Yee Seng Ng1, Mohammed M Kanani2, William E King1

  • 1Department of Radiology, University of Washington, Seattle, Washington.

Journal of the American College of Radiology : JACR
|November 20, 2025
PubMed
Summary

Large language models (LLMs) automate ground truth label extraction from radiology reports, enabling scalable assessment of artificial intelligence (AI) tools. This method reliably validates AI performance, even with new imaging hardware like photon-counting CT scanners.

Keywords:
Artificial intelligenceMLOpsground truth extractionlarge language modelsmodel validationphoton-counting CTradiology

More Related Videos

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
05:49

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization

Published on: February 23, 2024

1.3K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K

Related Experiment Videos

Last Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1000
Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization
05:49

Author Spotlight: Advancing CBCT and Digital Dental Image Integration with AI-Assisted Digitization

Published on: February 23, 2024

1.3K
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
05:47

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems

Published on: June 13, 2025

1.3K

Area of Science:

  • Radiology
  • Artificial Intelligence
  • Medical Informatics

Background:

  • Radiologic artificial intelligence (AI) tools require continuous monitoring and validation.
  • Manual ground truth label extraction from radiology reports is time-consuming and resource-intensive.
  • New imaging hardware, such as photon-counting CT (PCCT) scanners, can introduce input drift affecting AI performance.

Purpose of the Study:

  • To evaluate the feasibility of using large language models (LLMs) for automated ground truth label extraction from radiology reports.
  • To enable scalable assessment and monitoring of radiologic AI tools.
  • To validate AI model performance on a new PCCT scanner.

Main Methods:

  • Retrospective analysis of four FDA-cleared AI tools for pulmonary embolism, intracranial hemorrhage, cervical spinal fractures, and vertebral compression fractures.
  • LLM (Llama 3.3) used to extract binary ground truth labels from radiology reports of PCCT and conventional scanner data.
  • Comparison of AI outputs with LLM-extracted labels, with discrepant cases adjudicated by human annotators.
  • Interrater reliability measured using Fleiss's κ test; performance metrics recalculated after LLM error correction.

Main Results:

  • LLM-extracted labels facilitated rapid AI performance assessment across all four diagnostic tasks.
  • No statistically significant performance differences were observed between PCCT and non-PCCT cohorts.
  • LLM labels showed strong agreement with final human annotations (κ = 0.731), comparable to interreader agreement (κ = 0.720), confirming LLM labeling reliability.

Conclusions:

  • Large language models offer a scalable and efficient automated solution for ground truth label extraction from radiology reports.
  • This LLM-based approach supports rapid local validation of AI tools, effectively addressing challenges posed by new imaging hardware and input drift.