Related Experiment Video
Updated: Apr 15, 2026

Author Spotlight: Standardizing Mouse In Vivo PET Imaging with Body Conforming Molds and Automated Analysis
Published on: October 25, 2024
ConTEXTual Net 3D: Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
Zachary Huemann1, Samuel Church1, Joshua D Warner1
1Department of Radiology, University of Wisconsin-Madison, 1111 Highland Ave, Madison, WI, 53705, USA.
Abstract:
Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radiology reporting. However, these models require large annotated image-text datasets, which are lacking for PET/CT. We developed an automated pipeline to generate weak image-text labels and used it to train a 3D visual grounding model. Our weak-labeling pipeline identified sentences describing positive findings in PET/CT reports by searching for mentions of standardized uptake values (SUVmax) and axial slice numbers. These were used to automatically generate lesion masks, which were paired with the corresponding text descriptions. From 25,578 PET/CT exams, we extracted 11,356 sentence-label pairs. Using this data, we trained ConTEXTual Net 3D, which takes as input a description of a lesion and generates a corresponding segmentation mask. The model's performance was evaluated on 251 radiologist-reviewed cases and compared against LLMSeg, a 2.5D version of ConTEXTual Net, and two radiologists. We evaluated detection performance using F1 score. The weak-labeling pipeline accurately identified lesion locations in 98% of cases (246/251). ConTEXTual Net 3D achieved an F1 score of 0.80, outperforming LLMSeg (F1 = 0.22) and the 2.5D model (F1 = 0.53), though it underperformed both radiologists (F1 = 0.94 and 0.91). The model achieved better performance on 18F-fluorodeoxyglucose (F1 = 0.78) and DCFPyL (F1 = 0.75) exams than on DOTATATE (F1 = 0.58) and 18F-fluciclovine (F1 = 0.66) exams. In conclusion, our novel weak labeling pipeline accurately produced an annotated dataset of PET/CT image-text pairs. ConTEXTual Net 3D significantly outperformed other models but fell short of the performance of nuclear medicine physicians. Our study suggests that even larger datasets may be needed to close this performance gap.
More Related Videos
10:23Author Spotlight: Three-Dimensional Cephalometric Landmark Annotation Demonstration on Human Cone Beam Computed Tomography Scans
Published on: September 8, 2023
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Imaging Studies III: Computed Tomography
Imaging Studies I: CT and MRI
Description of the Procedures
Computed Tomography (CT) scan:
Computed Tomography (CT) scans use X-ray technology to generate detailed images of bones, organs, and tissues. During the scan, the patient lies on a moving table...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...