Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Evaluation of Proprietary and Open-Source Large Language Models for Systematic Multi-source Information
Elif Can1, Wibke Uller2, Elmar Kotter2
1Department of Diagnostic and Interventional Radiology, Medical Center - University of Freiburg, Faculty of Medicine, University of Freiburg, Hugstetter Str. 55, 79106, Freiburg, Germany. elif.can@uniklinik-freiburg.de.
Proprietary large language models (LLMs) like GPT-4o and Gemini demonstrate superior accuracy in extracting clinical data from transarterial chemoembolization (TACE) reports for hepatocellular carcinoma (HCC) patients compared to open-source models. However, expert human oversight is still crucial for certain variables.
Area of Science:
- Artificial Intelligence in Medicine
- Oncology Informatics
- Natural Language Processing in Healthcare
Background:
- Hepatocellular carcinoma (HCC) management relies on accurate interpretation of complex transarterial chemoembolization (TACE) reports.
- Large Language Models (LLMs) show potential for automating clinical data extraction from unstructured text.
- Evaluating LLM performance in extracting specific variables from TACE reports is crucial for clinical adoption.
Purpose of the Study:
- To compare the performance of proprietary (GPT-4o, Gemini 1.5 Pro) and open-source (Llama 3.1 70B, Llama 3.1 405B) LLMs.
- To assess the accuracy of LLMs in extracting clinically relevant variables from TACE reports for HCC patients.
- To evaluate the robustness of LLMs in longitudinal data extraction.
Main Methods:
- Retrospective analysis of 556 TACE-related reports from 50 HCC patients.
- Extraction of predefined binary and ordinal variables using standardized prompts.
- Performance assessment based on accuracy, ordinal scores, and longitudinal error rates via mixed-effects regression.
Main Results:
- Proprietary LLMs (GPT-4o, Gemini) significantly outperformed open-source LLMs (Llama 3.1) in accuracy for both binary and ordinal variables.
- GPT-4o demonstrated the lowest longitudinal error rate, indicating greater robustness over time.
- All evaluated LLMs exhibited poor performance in detecting vascular invasion and follow-up assessments.
Conclusions:
- Proprietary LLMs can accurately extract most TACE-related variables, potentially aiding decision-making in interventional oncology.
- Despite high accuracy for many variables, limitations in detecting vascular invasion and follow-up necessitate continued expert human oversight.
- LLM integration into clinical workflows requires careful validation and consideration of model-specific strengths and weaknesses.
Related Concept Videos
Targeted Cancer Therapies
There are several types of targeted therapies against...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Genomics

