Related Experiment Video
Updated: Aug 28, 2026

Computed Tomography-guided Time-domain Diffuse Fluorescence Tomography in Small Animals for Localization of Cancer Biomarkers
Published on: July 17, 2012
Text-SpecDiffusion: A text-guided multimodal spectral diffusion model for few-shot renal cell carcinoma staging
Zehai Hou1, Shengkun Yan1, Keyao Li1
1School of Optical and Electronic Information, School of Software Engineering, Huazhong University of Science and Technology, Wuhan, Hubei, 430074, China.
Abstract:
Despite rapid progress in multimodal spectroscopy for cancer diagnosis, its practical application remains limited by the scarcity of annotated data, especially for rare cancers. The scarcity of such samples often hampers the effective training of deep neural networks. Moreover, clinical text is rarely used to guide spectroscopic data augmentation, and a unified framework for text-guided spectral augmentation is still lacking. To address these challenges, we introduce a clinical text-guided generative framework based on diffusion models for cancer spectral analysis (Text-SpecDiffusion). Specifically, we first pretrain a vector-quantized variational autoencoder on a large urinary stone spectral dataset, enabling the model to learn shared spectral priors of the urinary system and thereby reduce the domain gap under few-shot conditions. We then incorporate clinical text into spectral generation through a cross-attention mechanism to guide the formation of microscopic spectral features and preserve pathophysiological consistency. Finally, we introduce a physics-informed loss based on the relative intensities of characteristic peaks to preserve the physical fidelity of the generated spectra. Experiments on a real-world renal cell carcinoma (RCC) dataset demonstrate that, by combining explicit knowledge (physical constraints and clinical semantics) with implicit knowledge from the data distribution, the proposed method achieves high-quality spectral generation, with Pearson correlation coefficient (PCC) of 0.970 and 0.935 for surface-enhanced Raman scattering (SERS) and laser-induced breakdown spectroscopy (LIBS), respectively. In addition, relative to the original dataset, the generated data substantially improves few-shot RCC staging performance, increasing accuracy, sensitivity, and specificity by 34.37%, 35.62%, and 16.33%, respectively.