Related Experiment Videos

Do Multimodal Vision-Language Models Enhance the Medical Diagnostic Process? A Systematic Review

Lattawat Eauchai1, Laura Otálora González1, Yifan Shi1

  • 1Department of Anesthesiology and Perioperative Medicine, Division of Critical Care, Mayo Clinic, Rochester, MN 55905, USA.

Summary

Multimodal vision-language models (VLMs) outperform unimodal models for medical diagnosis. While standalone VLMs show inconclusive results against physicians, copilot models enhance diagnostic accuracy, though further research is needed.