Related Experiment Video
Updated: Oct 10, 2026

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
Improving Clinical Reliability of Diffusion-Based Dermatology Image Synthesis Through Synthetic Note Alignment and
Niccolo Marini1, Zhaohui Liang1, Sivaramakrishnan Rajaraman1
1Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD, 20894, USA.
Abstract:
The development of dermatology image synthesis methods is gaining increasing attention as a solution to address persistent data scarcity and class imbalance, which constrain the development of reliable foundation models for clinical decision support. Although recent dermatology foundation models achieve dermatologist-level performance, their effectiveness critically depends on large-scale multimodal datasets. Publicly available dermatology datasets are usually limited in size, imbalanced, and often lack structured textual descriptions of lesions, limiting robust training and generalization. In addition, diffusion-based generative models require image-text pairs and may produce hallucinated or diagnostically inconsistent outputs, limiting their safe integration into clinical pipelines. This paper proposes a framework to improve dermatology image synthesis, leveraging synthetic clinical notes as structured conditioning prompts within a Stable Diffusion model, combined with a dedicated semantic filtering strategy to enhance reliability. First, multiple Large Language Models (LLMs) generate structured clinical descriptions to pair with real dermatology images. Second, a dermatology-adapted diffusion model is fine-tuned on these pairs to enable controllable text-conditioned image generation and generates 160,000 synthetic images. Third, an auxiliary foundation model evaluates the semantic consistency of synthetic images and their prompts, filtering hallucinated or diagnostically inconsistent samples. The filtered synthetic images are combined with 16,000 real images from six public repositories to train a foundation model. The evaluation involves fifteen datasets (37,000 images) and two downstream tasks (cross-modal retrieval and zero-shot learning), demonstrating improved performance, robustness, and cross-dataset generalization compared to a baseline model trained without synthetic note conditioning.