Compact Vision-Language Models Enable Efficient and Interpretable Automated OCT Analysis Through Layer Specific

Tania Haghighi1, Sina Gholami1, Jared Todd Sokol2

  • 1Department of Electrical and Computer Engineering, University of North Carolina at Charlotte, Charlotte, NC 28223, USA.

Summary

We developed LO-VLM, an efficient AI model for interpreting OCT B-scans, to generate accurate clinical narratives and classify retinal diseases. This vision-language model significantly outperforms existing methods in both summary generation and diagnostic accuracy.