Related Experiment Video
Updated: Aug 6, 2026

A New Technique for Treating Low-risk Prostate Cancer—Super Active Surveillance
Published on: November 7, 2025
A Supervised Fine-Tuned Large Language Model for Lifestyle Management in Patients With Prostate Cancer: Development
Fangyuan Jiang1, Qiuwen Yang1, Xin Zheng1
1Department of Medical Informatics, School of Medicine, Nantong University, Qixiu Road 19#, Nantong, Jiangsu, 226001, China, +86 0513 8505 1891.
Background:
Lifestyle interventions for patients with prostate cancer have been shown to improve treatment adherence and quality of life. However, there remains a lack of large language models (LLMs) capable of delivering individualized and professional lifestyle recommendations under clearly defined medical safety boundaries and controlled evidence sources.
Objective:
This study aimed to develop and evaluate a supervised fine-tuned LLM-PCaPLMM_SFT (Prostate Cancer Patient Lifestyle Management Model via Supervised Fine-Tuning)-to support health literacy improvement and lifestyle self-management among patients with prostate cancer.
Methods:
We searched English-language literature primarily from PubMed (February 2015 to February 2025) to build a structured lifestyle management knowledge base covering diet, physical activity, weight management, medication adherence, and psychological support. We used a retrieval-augmented generation pipeline to generate patient-style question-answer (QA) pairs from retrieved knowledge slices. Bilingual English-Chinese QA data were generated from English-language source evidence through patient-oriented reformulation and retrieval-augmented generation-based answer generation, and independent English and Chinese test sets were constructed to assess bilingual QA performance. We trained Baichuan2-7B-Chat using a 2-stage strategy, consisting of continued pretraining, followed by supervised fine-tuning with low-rank adaptation. Model outputs were evaluated in 2 double-blind rounds by referee LLMs (Qwen3-Max and DeepSeek-R1) and compared with GPT-3.5-Turbo and the base Baichuan2-7B-Chat using 2500 queries across 5 lifestyle scenarios. Additionally, 3 domain experts conducted a blinded review of 50 QA samples (10 per scenario). We used the Mann-Whitney U test with effect size r, and Benjamini-Hochberg false discovery rate correction, and examined consistency using intraclass correlation coefficients.
Results:
Based on 2211 included publications, we constructed the PCaPLMM_SFT-Train dataset. The knowledge base yielded >150,000 structured knowledge slices. After 2 rounds of review, we obtained 42,330 single-turn QA pairs and 3008 multiturn dialogues, and the supervised fine-tuning phase used 45,338 structured QA samples. In the dual-round referee LLM assessment, PCaPLMM_SFT consistently outperformed Baichuan2-7B-Chat across dimensions and showed comparable or superior performance to GPT-3.5-Turbo across 5 lifestyle scenarios. Consistency analyses indicated moderate to good agreement between referee models across rounds, supporting the robustness of the comparative evaluation.
Conclusions:
PCaPLMM_SFT demonstrates the feasibility of constructing a medical lifestyle-focused LLM by integrating structured medical knowledge, QA-style training data, and a multilayer evaluation system. This framework provides a reproducible methodological foundation for evidence-based health education and lifestyle management and establishes groundwork for future evaluation in real-world health management settings.