Related Experiment Video
Updated: May 23, 2026

Implementation of In Vitro Drug Resistance Assays: Maximizing the Potential for Uncovering Clinically Relevant Resistance Mechanisms
Published on: December 9, 2015
Benchmarking Large Language Models and Prompt Engineering Strategies in Microsatellite Instability Cancers:
Yuxin Zhang1, Jie Song1, Cheng Bi1
1Department of Medical Oncology, Institutes for Systems Genetics, Frontiers Science Center for Disease-Related Molecular Network, West China Hospital, Sichuan University, No 2222 Xinchuan Road, Gaoxin District, Chengdu, Sichuan, 610000, China, 86 15995854635, 86 28 61528682.
Large language models (LLMs) struggle with microsatellite instability (MSI) cancer tasks. Retrieval-augmented generation (RAG) significantly improves accuracy and safety, but requires optimized retrieval and knowledge bases for reliable clinical AI.
Area of Science:
- Artificial Intelligence in Oncology
- Clinical Decision Support Systems
- Biomedical Natural Language Processing
Background:
- General-purpose large language models (LLMs) have uncharacterized reliability for complex clinical tasks in specialized domains like microsatellite instability (MSI) cancers.
- The lack of a domain-specific benchmark for evaluating LLM capabilities in MSI oncology poses risks to patient safety.
Purpose of the Study:
- To develop and validate the Microsatellite Instability Cancer Benchmark (MSIC-Bench) for evaluating LLMs in MSI oncology.
- To systematically assess LLM performance across prompting strategies and identify areas for improvement.
Main Methods:
- Developed MSIC-Bench, a 511-question benchmark from clinical guidelines and curated knowledge.
- Evaluated three state-of-the-art LLMs (GPT-4o, Gemini 2.5 Pro, Claude Opus 4) using four prompting strategies (vanilla, chain-of-thought, reflection of thoughts, RAG).
- Assessed performance based on accuracy, safety, error composition, and token usage across multiple-choice and open-ended modalities.
Main Results:
- LLMs exhibited a 'scaffolding effect,' with accuracy decreasing in open-ended scenarios.
- Retrieval-augmented generation (RAG) was the most effective intervention, shifting bottlenecks from knowledge deficits to retrieval failures.
- RAG improved accuracy and safety by reducing fabrications, though it introduced a trade-off with false refusals. Hybrid-RAG showed robust performance.
Conclusions:
- Current LLMs lack specialized knowledge for MSI oncology; RAG is crucial for addressing this gap.
- Optimizing RAG requires focusing on retrieval precision and high-quality knowledge bases for trustworthy clinical AI.
- MSIC-Bench provides a framework to guide future development of clinical AI in MSI oncology.

