Related Experiment Video
Updated: Jun 17, 2026

07:13
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Cancer staging data collection using rules-based natural language processing for entity extraction from pathology
Jade Newton1,2, Jamiu Ekundayo3, Nancy Tippaya3
1School of Population Health, Curtin University, Kent St, Bentley, WA, 6102, Australia.
BMC Medical Informatics and Decision Making
|June 16, 2026
Summary
Automated cancer staging using rules-based natural language processing (NLP) systems accurately extracts tumour, node, and metastasis (TNM) data. This approach supports partial automation of cancer staging workflows, improving efficiency in clinical research and resource allocation.
Area of Science:
- Computational pathology
- Medical informatics
- Natural Language Processing (NLP)
Background:
- Population-level cancer staging data is crucial for treatment planning, outcome prediction, and resource allocation.
- Manual cancer staging is resource-intensive, highlighting the need for efficient automated methods.
- Existing manual methods for cancer data collection are resource-intensive.
Purpose of the Study:
- To develop rules-based NLP systems for extracting explicit and implicit tumour, node, and metastasis (TNM) entities.
- To translate extracted TNM values into cancer stages for melanoma, breast, and colorectal cancers.
- To align cancer staging with the American Joint Committee on Cancer (AJCC) 8th edition TNM staging system.
Main Methods:
- Developed rules-based NLP systems in consultation with domain experts.
- Extracted cancer staging information from pathology reports and hospital inpatient morbidity data.
- Evaluated system performance against manual collections using precision, recall, and F1-scores.
Main Results:
- Rules-based NLP systems achieved 87%-90% accuracy in staging cases compared to manual collections.
- Melanoma NLP system achieved weighted average precision of 0.96, recall of 0.94, and F1-score of 0.94.
- Colorectal and breast cancer NLP models demonstrated high performance with weighted average F1-scores of 0.89 and 0.89, respectively.
Conclusions:
- A rules-based NLP architecture can accurately derive TNM components and cancer stage from clinical text.
- The developed methods support partial automation of cancer staging workflows, especially where training datasets are limited.
- The approach relies on domain-expert input rather than data-driven training.

