Related Experiment Video
Updated: May 17, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Benchmarking large language models GPT-4o, llama 3.1, and qwen 2.5 for cancer genetic variant classification
Kuan-Hsun Lin1,2, Tzu-Hang Kao3, Lei-Chi Wang3,4
1Department of Information Management, Taipei Veterans General Hospital, Taipei, Taiwan, ROC.
Large language models (LLMs) show promise in classifying cancer genetic variants for precision oncology. GPT-4o achieved the highest accuracy, but further optimization is needed for clinical use.
Area of Science:
- Genomics
- Computational Biology
- Precision Oncology
Background:
- Classifying cancer genetic variants by clinical actionability is essential for personalized cancer treatment.
- Large language models (LLMs) present a novel approach to this classification challenge, but their efficacy requires thorough evaluation.
Purpose of the Study:
- To assess the performance of leading LLMs (GPT-4o, Llama 3.1, Qwen 2.5) in classifying cancer genetic variants.
- To compare LLM accuracy against established databases (OncoKB, CIViC) and real-world clinical data.
- To investigate the impact of prompt engineering and retrieval-augmented generation (RAG) on LLM performance.
Main Methods:
- Evaluated GPT-4o, Llama 3.1, and Qwen 2.5 on variant classification tasks.
- Utilized OncoKB, CIViC databases, and FoundationOne CDx reports for dataset creation.
- Performed prompt engineering and RAG to optimize LLM performance.
- Conducted stability analysis across 100 iterations.
Main Results:
- GPT-4o demonstrated superior accuracy (0.7318) in classifying clinically relevant variants versus variants of unknown significance (VUS).
- LLMs showed higher agreement with expert annotations for strongly supported variants, with increased variability for weaker evidence.
- A tendency towards overclassification (assigning higher evidence levels) was observed across all models.
- Prompt engineering and RAG significantly improved classification accuracy and consistency.
Conclusions:
- LLMs, particularly GPT-4o, show significant potential for automating cancer genetic variant classification.
- Further refinement is necessary to enhance accuracy, reduce overclassification bias, and ensure consistent clinical applicability.
- The study highlights the need for robust validation and optimization strategies for LLMs in precision oncology workflows.
More Related Videos
09:33Author Spotlight: Finding New Therapeutic Targets for Malignant Peripheral Nerve Sheath Tumor Through Genome-Scale shRNA Screens
Published on: August 25, 2023
11:02Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
Published on: October 18, 2013
Related Concept Videos
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Pleiotropy