Related Experiment Video
Updated: Jun 11, 2025

07:32
Author Spotlight: Investigating Immune Cell Dynamics in the Tumor Microenvironment — Challenges and Innovations in Cancer Prognosis
Published on: April 12, 2024
1.2K
Validation of large language models for detecting pathologic complete response in breast cancer using
Ken Cheligeer1,2, Guosong Wu1,3, Alison Laws4,5
1The Centre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary, Canada.
BMC Medical Informatics and Decision Making
|October 4, 2024
Summary
Large Language Models (LLMs) show high accuracy in identifying pathologic complete response (pCR) from breast cancer pathology reports. This advancement in medical documentation analysis promises improved patient care and research capabilities.
Area of Science:
- Artificial Intelligence in Medicine
- Computational Pathology
- Natural Language Processing for Healthcare
Background:
- Accurate identification of pathologic complete response (pCR) is crucial for breast cancer management and treatment efficacy assessment.
- Manual review of complex pathology reports is time-consuming and prone to variability.
- Large Language Models (LLMs) offer potential for automated analysis of unstructured clinical data.
Purpose of the Study:
- To evaluate the capability of LLMs in understanding and processing narrative pathology reports for pCR identification.
- To compare the performance of LLM-based pipelines against traditional machine learning models.
- To enhance clinical data extraction for improved health research and patient care.
Main Methods:
- Developed two analytical pipelines using open-source LLMs within a healthcare computing environment.
- Extracted embeddings from 351 breast cancer pathology reports using 15 transformer-based models, followed by logistic regression for pCR classification.
- Fine-tuned the Generative Pre-trained Transformer-2 (GPT-2) model with a feed-forward neural network (FFNN) layer for enhanced pCR detection.
Main Results:
- The optimized LLM pipeline achieved a sensitivity of 95.3%, a positive predictive value of 90.9%, and an F1 score of 93.0% for pCR identification.
- LLM-based methods demonstrated superior performance compared to traditional machine learning models.
- The study highlights the potential of LLMs in extracting critical clinical information from narrative pathology reports.
Conclusions:
- LLMs are effective in interpreting digital pathology data for determining pCR in breast cancer patients post-neoadjuvant chemotherapy (NAC).
- LLM-based pipelines show significant promise for clinical data extraction and analysis, outperforming conventional methods.
- External validation is recommended to confirm the reliability and broad applicability of these LLM approaches in clinical settings.

