Related Experiment Video
Updated: May 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance Evaluation of Large Language Models in Cervical Cancer Management Based on a Standardized Questionnaire:
Warisijiang Kuerbanjiang1, Shengzhe Peng1, Yiershatijiang Jiamaliding1
1Department of Gynecology, Zhongnan Hospital of Wuhan University, Wuhan, Hubei Province, China.
Large language models (LLMs) show potential in cervical cancer management, with proprietary models like ChatGPT-4.0 Turbo outperforming others. Prompts enhance accuracy, but medical-specialized models need further development for clinical use.
Area of Science:
- Artificial Intelligence in Medicine
- Computational Health
- Digital Health
Background:
- Cervical cancer is a significant global health challenge, especially in low-resource settings.
- Comprehensive management requires integrated screening, diagnosis, and treatment strategies.
- Large language models (LLMs) offer potential but are underexplored in cervical cancer care.
Purpose of the Study:
- To systematically evaluate the performance and interpretability of LLMs in cervical cancer management.
- To compare the effectiveness of different LLMs based on accuracy and guideline compliance.
- To assess the impact of prompts on LLM performance in a medical context.
Main Methods:
- Selected LLMs from AlpacaEval leaderboard v2.0.
- Developed 100 standardized questions covering cervical cancer screening, diagnosis, and treatment.
- Utilized the CO-STAR framework for prompt engineering.
- Evaluated responses for accuracy, guideline adherence, clarity, and practicality (A-D scoring).
- Employed LIME for interpretability analysis.
Main Results:
- ChatGPT-4.0 Turbo ranked highest (effective rate 94% with prompts) among nine evaluated LLMs.
- Seven models demonstrated stable responses; QiZhenGPT consistently underperformed.
- Prompts improved alignment with human annotations for proprietary models.
- Medical-specialized models showed limited improvement with prompts.
Conclusions:
- Proprietary LLMs (ChatGPT-4.0 Turbo, Claude 2) show promise for clinical decision support in cervical cancer.
- Prompt engineering can enhance LLM accuracy in this domain.
- Further research is needed to integrate LLMs into practical medical applications.
More Related Videos
06:46Competing-Risk Nomogram for Predicting Cancer-Specific Survival in Multiple Primary Colorectal Cancer Patients after Surgery
Published on: September 27, 2024
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025