在膀癌管理中增强人工智能:对多种大型语言模型进行比较分析和优化研究
Kun-Peng Li1,2, Li Wang1,2, Shun Wan1,2
1Department of Urology, The Second Hospital of Lanzhou University, Lanzhou, China.
Journal of endourology
|March 18, 2025
概括
大型语言模型 (LLM) 在医疗保健方面表现有前途,但与专业瘤学相斗争. 战略优化显著提高了GPT-3.5-Turbo对于膀癌 (BLCA) 临床问题的准确性,达到100%.
科学领域:
- 人工智能在医学中的应用
- 在瘤学瘤学.
- 医疗信息学 医疗信息学
背景情况:
- 大型语言模型 (LLM) 在医疗保健中越来越多地使用,但它们在瘤学等专业领域的有效性是有限的.
- 在解决复杂的临床调查时,领先的LLM之间存在性能差异.
研究的目的:
- 评估多个领先的LLM在回答与膀癌 (BLCA) 相关的临床问题的表现.
- 证明战略优化对提高专业瘤学应用的LLM准确性的影响.
主要方法:
- 根据指导方针制定了一套100个临床问题,涵盖BLCA的流行病学,诊断,治疗,预后和随访.
- 在三个试验中测试了六个LLM (克劳德-3.5-索内特,ChatGPT-4.0,Grok-beta,Gemini-1.5-Pro,Mistral-Large-2,GPT-3.5-Turbo) 在三个试验中进行了测试.
- GPT-3.5-Turbo经历了两阶段的训练优化过程.
主要成果:
- 克劳德3.5-索内特获得了最高的初始精度 (89.33%),而GPT-3.5-Turbo的初始精度最低 (74.33%).
- 经过两个优化阶段,GPT-3.5-Turbo的精度提高到86.67%,随后达到100%.
- 对比分析显示,在测试的LLM中,表现有显著差异.
结论:
- 在膀癌等专门的瘤学领域,LLM的表现各不相同.
- 有针对性的培训优化可以大大提高临床决策支持的LLM准确性.
- 将GPT-3.5-Turbo成功改进到100%的准确度突出了战略模型改进的潜力.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Mouse Models of Cancer Study
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


