Related Experiment Video
Updated: Jan 6, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
SegORG: Report Generation of Oral Potentially Malignant Disorders Image Based on Lesion Segmentation-Enhanced LLM
Rui Zhang1,2,3, Peng Huang1, Tingting Ding1
1Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang Provincial Clinical Research Center for Oral Diseases, Zhejiang Key Laboratory of Oral Biomedical, Hangzhou, China.
Objectives:
To develop an automated system for generating standardized reports for oral potentially malignant disorders (OPMDs) from white-light images, aiming to reduce documentation workload, facilitate early intervention, and enable longitudinal lesion monitoring.
Methods:
We proposed the SegORG model using 441 oral mucosa images and 1323 corresponding reports. It employed SegFormer for lesion segmentation; a visual encoder extracted global and local visual embeddings, which were projected into a pre-trained large language model (LLM feature) space via a lightweight visual mapper. The Qwen2.5-7B model then generated structured diagnostic reports, enhanced by text augmentation techniques to improve diversity and professionalism.
Results:
SegORG achieved BLEU-4, ROUGE-L, and CIDEr scores of 0.291, 0.517, and 0.578, respectively. Additionally, the model obtained a clinical diagnostic F1-score of 0.695 and a median expert rating of 4 on the Likert scale (p < 0.001). It significantly outperformed conventional baseline models (R2Gen, METransformer, and SwinB+BERT9k) and contemporary general-purpose multimodal LLMs (GPT-4, Gemini 2.5 Pro, and Qwen2.5-VL), despite its lean architecture (90.4 M trainable parameters).
Conclusions:
By enhancing visual feature extraction and achieving efficient text alignment, SegORG offers an effective technical pathway for OPMDs reports automation. While single-center validation shows promise, multicenter trials are needed to assess generalizability.

