Related Experiment Video
Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model
Soheil Hashtarkhani1, Rezaur Rashid1, Christopher L Brett2
1Center for Biomedical Informatics, Department of Pediatrics, College of Medicine, University of Tennessee Health Science Center, 50 N Dunlap Street, Memphis, TN, 38103, United States, 1 9012875836.
BioBERT and GPT-4o show promise for classifying cancer diagnoses from electronic health records. While BioBERT excels with structured data, GPT-4o performs better with free text, indicating potential for AI in healthcare administration and research.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Electronic health records (EHRs) contain varied data formats, necessitating efficient preprocessing for predictive healthcare models.
- Artificial intelligence (AI) and natural language processing (NLP) tools offer potential for automating diagnosis classification but require rigorous evaluation for clinical reliability.
Purpose of the Study:
- To evaluate the performance of large language models (LLMs) including GPT-3.5, GPT-4o, Llama 3.2, Gemini 1.5, and BioBERT.
- To assess their accuracy in classifying cancer diagnoses from both structured (International Classification of Diseases [ICD] codes) and unstructured (free-text) EHR data.
Main Methods:
- Analysis of 762 unique cancer diagnoses from 3456 patient records.
- Models were tested on classifying diagnoses into 14 predefined categories.
- Classifications were validated by two oncology experts.
Main Results:
- BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in accuracy (90.8).
- GPT-4o outperformed BioBERT in weighted macro F1-score for free-text diagnoses (71.8 vs 61.5) and accuracy (81.9 vs 81.6).
- GPT-3.5, Gemini, and Llama demonstrated lower overall performance; common errors involved metastasis, CNS tumors, and ambiguous terminology.
Conclusions:
- Current AI model performance is adequate for administrative and research purposes in cancer diagnosis classification.
- Clinical applications necessitate standardized documentation and strong human oversight for critical decision-making.
More Related Videos
Related Concept Videos
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Cancer Survival Analysis

