Related Experiment Video
Updated: Jan 11, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Development of a machine learning model for automatic data extraction from breast cancer pathology reports
Christy Oi Ting Kwok1, Gregory Arbour2, Annah Zhang1
1Department of Surgery, Faculty of Medicine, University of British Columbia, 2221 Wesbrook Mall, Vancouver, BC, V6T 2B5, Canada.
Automating breast cancer pathology report analysis using Natural Language Processing (NLP) significantly improves data extraction accuracy. A PubMedBERT model achieved 97.4% accuracy, outperforming previous methods for clinical research.
Area of Science:
- Medical Informatics
- Computational Biology
- Oncology Research
Background:
- Manual data extraction from medical records is time-consuming and prone to errors.
- Natural Language Processing (NLP) and Machine Learning (ML) offer potential for automating clinical data extraction.
- Accurate extraction of clinicopathologic features from breast cancer pathology reports is vital for research and patient outcomes.
Purpose of the Study:
- To develop and evaluate an NLP pipeline for automated data extraction from breast cancer pathology reports.
- To compare the performance of different NLP models, including ClinicalBERT, PubMedBERT, BioMedRoBERTa, and Mistral-Nemo LLM.
- To assess the impact of pre-training on the SQuAD question-answering dataset for improved model accuracy.
Main Methods:
- An NLP pipeline was constructed by selecting the best performing model from four candidates.
- Models were evaluated on 1,795 breast cancer pathology reports.
- BERT-based models underwent further pre-training using the SQuAD dataset.
- Model performance was quantified by comparing extracted data against human annotations.
Main Results:
- PubMedBERT, pre-trained on SQuAD, achieved the highest overall accuracy of 97.4%.
- 30 out of 32 Fields of Interest (FOIs) demonstrated accuracy exceeding 95.0%.
- The developed NLP model outperformed a previous rule-based algorithm (95.6% accuracy).
Conclusions:
- A high-performing question-answering NLP pipeline can automate the extraction of clinicopathologic features from breast cancer pathology reports.
- This automated approach offers a scalable solution for high-fidelity data extraction, enhancing research efficiency.
- Improved data extraction can contribute to better clinical outcomes in breast cancer care.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025