Related Experiment Video
Updated: May 16, 2026

08:43
An Orthotopic Bladder Tumor Model and the Evaluation of Intravesical saRNA Treatment
Published on: July 28, 2012
14.7K
Enhanced Artificial Intelligence in Bladder Cancer Management: A Comparative Analysis and Optimization Study of
Kun-Peng Li1,2, Li Wang1,2, Shun Wan1,2
1Department of Urology, The Second Hospital of Lanzhou University, Lanzhou, China.
Journal of Endourology
|March 18, 2025
Summary
Large language models (LLMs) show promise in healthcare but struggle with specialized oncology. Strategic optimization significantly improved GPT-3.5-Turbo’s accuracy for bladder cancer (BLCA) clinical questions to 100%.
Area of Science:
- Artificial Intelligence in Medicine
- Oncology
- Medical Informatics
Background:
- Large language models (LLMs) are increasingly utilized in healthcare, yet their efficacy in specialized fields like oncology is limited.
- Performance variations exist among leading LLMs when addressing complex clinical inquiries.
Purpose of the Study:
- To evaluate the performance of multiple leading LLMs in answering clinical questions related to bladder cancer (BLCA).
- To demonstrate the impact of strategic optimization on enhancing LLM accuracy for specialized oncology applications.
Main Methods:
- A set of 100 clinical questions covering epidemiology, diagnosis, treatment, prognosis, and follow-up for BLCA was developed based on guidelines.
- Six LLMs (Claude-3.5-Sonnet, ChatGPT-4.0, Grok-beta, Gemini-1.5-Pro, Mistral-Large-2, GPT-3.5-Turbo) were tested in three trials.
- GPT-3.5-Turbo underwent a two-phase training optimization process.
Main Results:
- Claude-3.5-Sonnet achieved the highest initial accuracy (89.33%), while GPT-3.5-Turbo had the lowest (74.33%).
- After two optimization phases, GPT-3.5-Turbo's accuracy improved to 86.67% and subsequently reached 100%.
- Comparative analysis revealed significant performance differences among tested LLMs.
Conclusions:
- LLM performance varies in specialized oncology domains like bladder cancer.
- Targeted training optimization can substantially enhance LLM accuracy for clinical decision support.
- GPT-3.5-Turbo's successful refinement to 100% accuracy highlights the potential of strategic model improvement.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Mouse Models of Cancer Study
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...

