Related Experiment Videos
An Automated, Contamination-Controlled VQA Benchmark for Evaluating Vision-Language Models on 3D Oncology Imaging
Abstract:
Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that builds multiple-choice benchmarks directly from paired private radiology reports and three-dimensional oncology imaging. It generates two complementary question types: schema-driven items populated deterministically from established reporting frameworks, and report-derived items verified against the source text. Because the source reports are private and single-institution, the resulting questions cannot have entered any model's pretraining data, controlling instance-level contamination by construction. We applied the pipeline to four oncology cohorts (liver CT, liver MRI, lung CT and brain MRI; 2,509 cases and 33,870 questions) and evaluated five contemporary VLMs zero-shot. Accuracies ranged from 0.28 to 0.81 and no model was reliable across cohorts. Replacing the image with a blank input changed accuracy little for the highest-scoring models, and on schema-driven brain questions frontier models scored 0.17-0.19 higher on public than on matched private imaging. The benchmark thus separates image-dependent performance from question-text and dataset-familiarity effects, and the released pipeline lets institutions regenerate it on their own reports and images.