Related Experiment Videos
Synthesized clinical notes enable training robust multimodal AI models from unimodal dermatology datasets
Niccolo Marini1, Zhaohui Liang2, Sivaramakrishnan Rajaraman2
1National Library of Medicine, National Institutes of Health, Bethesda, MD, USA. niccolo.marini@nih.gov.
None:
Multimodal (MM) algorithms and foundation models have shown strong potential for automated skin lesion analysis in dermatology, however, their translation to clinical practice is hindered by the limited availability of large, high-quality image-text datasets. Most public dermatology datasets are small, unimodal, and paired with heterogeneous labels and metadata, restricting effective model development. Large Language Models (LLMs) provide an opportunity to synthesize clinical notes from existing unimodal datasets, but their application is limited by hallucinations that can hinder MM training. This paper investigates and evaluates strategies to generate reliable clinical notes using existing LLMs and metadata paired with real images, aiming to mitigate hallucinations and improve performance on downstream tasks. We train a MM architecture on real dermatology images paired with synthesized notes and evaluate it across 15 datasets (6 internal, 9 external) on cross-modal retrieval and zero-shot learning tasks. Results demonstrate improved robustness and generalization compared to state-of-the-art medical foundation models, under specific synthesis conditions.