Enhancing vision-language model with pretraining for reasoning medical applications.

Yu Zhang1, Shuihua Wang1, Jia Meng1

  • 1Department of Biosciences and Bioinformatics, Suzhou Municipal key Lab AI4Health, School of Science, Xi'an Jiaotong-Liverpool University, Suzhou 215000, China; Department of Mathematical Sciences, University of Liverpool, Liverpool, UK.

Summary

This study introduces a Multi-modal Medical Reasoning Model (MMRM) that enhances vision-language models (VLMs) with chain-of-thought reasoning for medical diagnosis. The MMRM improves diagnostic accuracy and provides interpretable explanations, boosting clinical trust.