Related Experiment Video
Updated: May 2, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
M3: multimodal artificial intelligence for medical report generation and visual question answering from 3D abdominal
Abdullah Hosseini1, Ahmed Ibrahim1, Ahmed Serag1
1AI Innovation Lab, Weill Cornell Medicine-Qatar, Doha, 24144, Qatar.
Objectives:
Medical imaging is indispensable for diagnosis, with abdominal imaging playing a pivotal role in generating medical reports and informing clinical decision-making. Recent works in artificial intelligence (AI), particularly in multimodal approaches such as vision-language models, have demonstrated significant potential to enhance medical image analysis by seamlessly integrating visual and textual data. While 2D imaging has been the main focus of many studies, the enhanced spatial detail and volumetric consistency offered by 3D images, such as CT scans, remain relatively underexplored. This gap underscores the need for innovative approaches to unlock the potential of 3D imaging in clinical workflows.
Methods:
In this study, we utilized a multimodal AI pipeline, Phi3-V, to address 2 key challenges in abdominal imaging: generating clinically coherent medical reports from 3D CT images and performing visual question answering based on these images.
Results:
Our optimized model attained an average GREEN score of 0.409 for medical report generation and an accuracy of 79% for multiple-choice visual question answering on the validation cases.
Conclusions:
These findings demonstrate the potential of multimodal AI in advancing the analysis of 3D medical imaging, paving the way for more robust and efficient applications in healthcare.
Advances In Knowledge:
This study advances the use of multimodal AI for 3D CT imaging, achieving improvements in medical report generation and visual question answering.

