Related Experiment Video
Updated: Mar 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Clinical utility of large language models in metastatic prostate cancer: A multicenter expert validation for decision
Yinsong Chen1, Kangwen He1, Weinuo Qu1
1Department of Radiology, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, Hubei, China.
Background:
Initial systemic treatment planning for metastatic prostate cancer (mPCa) requires rapid synthesis of heterogeneous clinical and biomarker information. Large language models (LLMs) could assist clinicians, but their safety and acceptability in this high-stakes setting remain uncertain.
Methods:
We conducted a multicenter retrospective evaluation of 238 consecutive mPCa cases from three tertiary centers (2018-2025). Five contemporary LLMs were tested via publicly available web interfaces under a locked, zero-shot prompting protocol to generate a clinical summary, a first-line systemic treatment recommendation, and a rationale. Outputs underwent two-stage assessment: (1) multidisciplinary team (MDT) binary safety adjudication using a one-strike gate with a prespecified taxonomy of critical errors; unsafe outputs were assigned a Likert score of 1 for all domains; (2) three senior medical oncologists independently rated safety-passed outputs on 5-point Likert scales for summary accuracy, guideline-concordant and patient-tailored recommendations, and rationale quality. Paired ordinal outcomes were analyzed with Friedman tests and Holm-adjusted post hoc comparisons, and binary safety outcomes with Cochran's Q and McNemar tests.
Results:
Safety rates ranged from 79.0% to 84.9%. Among safety-passed outputs, mean utility scores (5-point Likert) were in the low-to-mid 4 range. Between-model differences were most apparent for summarization, whereas treatment recommendations and rationales showed modest separation after multiplicity adjustment. Failures clustered in hard cases with incomplete documentation and were dominated by missingness-related extraction errors, disease-state/pathway errors, guideline logic deviations, and safety-check omissions.
Conclusions:
Current LLMs can support mPCa care as drafting assistants, but ∼15%-21% of outputs breached safety thresholds under a strict gate, precluding unsupervised use at initial treatment-planning encounters.

