Related Experiment Video
Updated: Apr 15, 2026

Author Spotlight: Standardizing Mouse In Vivo PET Imaging with Body Conforming Molds and Automated Analysis
Published on: October 25, 2024
ConTEXTual Net 3D: Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
Zachary Huemann1, Samuel Church1, Joshua D Warner1
1Department of Radiology, University of Wisconsin-Madison, 1111 Highland Ave, Madison, WI, 53705, USA.
A new automated pipeline generates weak labels for PET/CT reports, creating image-text pairs for visual grounding models. ConTEXTual Net 3D, trained on this data, shows promise in lesion detection but requires larger datasets for human-level performance.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Radiology
Background:
- Vision-language models (VLMs) enable visual grounding, linking text descriptions to image locations.
- VLMs have potential in radiology reporting, but lack sufficient annotated PET/CT datasets.
- Developing effective VLMs for PET/CT requires novel approaches to data generation and model training.
Purpose of the Study:
- To develop an automated pipeline for generating weak image-text labels from PET/CT reports.
- To train a 3D visual grounding model (ConTEXTual Net 3D) using the generated weak labels.
- To evaluate the performance of ConTEXTual Net 3D against existing models and radiologists.
Main Methods:
- An automated pipeline identified sentences with SUVmax and slice numbers in PET/CT reports to create weak labels.
- Lesion masks were automatically generated and paired with corresponding text descriptions.
- A 3D visual grounding model, ConTEXTual Net 3D, was trained on 11,356 PET/CT image-text pairs.
Main Results:
- The weak-labeling pipeline achieved 98% accuracy in identifying lesion locations.
- ConTEXTual Net 3D achieved an F1 score of 0.80, significantly outperforming LLMSeg (0.22) and a 2.5D model (0.53).
- Performance varied by radiotracer, with higher scores for 18F-FDG and DCFPyL compared to DOTATATE and 18F-fluciclovine.
Conclusions:
- The novel weak-labeling pipeline successfully created an annotated PET/CT dataset.
- ConTEXTual Net 3D demonstrates strong performance, surpassing other automated methods but not yet matching expert radiologists.
- Larger datasets are likely necessary to bridge the performance gap between AI models and human experts in PET/CT interpretation.
More Related Videos
10:23Author Spotlight: Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Imaging Studies III: Computed Tomography
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...