Related Experiment Video
Updated: Jul 10, 2026

Using Computer-based Image Analysis to Improve Quantification of Lung Metastasis in the 4T1 Breast Cancer Model
Published on: October 2, 2020
Evaluation of large language models in female malignancy Q&A
Kai Xin1, Sen Hong2, Xiahui Wu3
1Department of Oncology, Nanjing Drum Tower Hospital, Affiliated Hospital of Medical School, Nanjing University, Nanjing, China.
Large language models (LLMs) show potential in answering questions about female malignancy, with Llama-3.1-405B, OpenAI o1, and DeepSeek-R1 performing best. Challenges remain with complex medical terminology and flexible scenarios.
Area of Science:
- Artificial Intelligence in Medicine
- Oncology
- Natural Language Processing
Background:
- Foundational large language models (LLMs) are crucial for AI in medicine.
- Limited evaluation studies exist on LLM Q&A performance in female malignancy.
Purpose of the Study:
- To evaluate and compare the Q&A capabilities of popular LLMs in female malignancy.
- To establish a benchmark dataset for assessing LLM performance in this specialized domain.
Main Methods:
- A cross-sectional study was conducted using a benchmark dataset of 205 female malignancy Q&A.
- Seven popular LLMs were evaluated against three human oncologists.
- Performance metrics included accuracy, consistency, latency, and inter-model agreement.
Main Results:
- Llama-3.1-405B, OpenAI o1, and DeepSeek-R1 exceeded 80% accuracy and consistency.
- Llama-3.1-405B showed higher accuracy and lower latency than OpenAI o1.
- Removing token limits improved OpenAI o1 and DeepSeek-R1 accuracy significantly, nearing Llama-3.1-405B's performance.
Conclusions:
- Llama-3.1-405B is a strong baseline for female malignancy Q&A.
- OpenAI o1 and DeepSeek-R1 show promise, especially without token limits, but require consistency improvements.
- LLMs need further refinement with clinical data to overcome challenges in complex medical terminology and scenarios.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025