Related Experiment Video
Updated: Jan 11, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Development of a machine learning model for automatic data extraction from breast cancer pathology reports
Christy Oi Ting Kwok1, Gregory Arbour2, Annah Zhang1
1Department of Surgery, Faculty of Medicine, University of British Columbia, 2221 Wesbrook Mall, Vancouver, BC, V6T 2B5, Canada.
None:
Data extraction from medical records is crucial for clinical research, with current methods relying on human annotation. Natural Language Processing (NLP) and Machine Learning-based approaches show promise. We develop and evaluate an NLP pipeline constructed by selecting among four candidate models; ClinicalBERT, PubMedBERT, BioMedRoBERTa and Mistral-Nemo LLM to automate data extraction of 1,795 breast cancer pathology reports obtained from the Providence Health Services Authority in British Columbia. We also explore the effect of further pre-training the BERT-based models using the SQuAD question-answering dataset. Accuracy was evaluated by comparing model output and human annotation. PubMedBERT pre-trained on SQuAD proved to be the best performing model, achieving an overall accuracy of 97.4%. 30 of the 32 FOIs had an accuracy greater than 95.0%. Our model outperformed a previous rule-based algorithm (95.6%). Our findings demonstrate how a high-performing question-answering NLP pipeline for breast cancer pathology can provide a scalable approach to high-fidelity extraction of clinicopathologic features, thereby enhancing research efficiency and improving clinical outcomes.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025