Related Experiment Video
Updated: Oct 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Enhancing Cancer Outcomes Research Through Vision Language Models for Cause of Death Ascertainment
Enshuo Hsu1,2, Giordana De Las Pozas3, Trey Kell1
1Enterprise Development and Integration, University of Texas MD Anderson Cancer Center, Houston, TX.
Background And Scope:
Cause of death is an important variable in cancer registries that supports downstream analyses, including cause-specific survival analysis. Unfortunately, in current practice, it is only available on government-issued death certificates, which are often received as scanned images. Current computer vision-based optical character recognition (OCR) methods struggle with handwriting, background noise, and blurriness, resulting in poor performance. A robust method for the cause of death ascertainment is urgently needed.
Solution:
Latest vision language models (VLMs), such as Qwen3-VL and GPT-4.1, have shown promising OCR performance on poor-quality scanned documents in the general domain. We proposed two prompting methods: (1) the two-step approach, which prompts VLMs for a full-text OCR followed by a text-based information extraction and (2) end-to-end approach, which prompts VLMs with a JavaScript Object Notation schema for image-based information extraction.
Evaluation:
We evaluated VLMs for the cause of death extraction from scanned government-issued death certificates in a major cancer center. Of 22,733 death records, 500 were manually annotated as ground truth. For the immediate cause of death, Qwen3-VL-30B-A3B-Thinking with the end-to-end approach achieved the best accuracy of 0.8625, significantly outperforming the Tesseract baseline (0.2396). For the underlying causes of death, the two-step method achieved the best F1 score of 0.7637, outperforming the baseline (0.2451).
Relevance:
The growing emphasis on precision medicine underscores the importance of accurate cause of death ascertainment for survival analyses. Our method demonstrated significant improvements over the current OCR approaches and provided a reliable alternative to reliance on the National Death Index.
How To Use:
Our VLM OCR computational framework is publicly available on GitHub. The open-weight VLMs are available for download on the Hugging Face repository.
