Related Experiment Video
Updated: Jul 1, 2025

07:59
Author Spotlight: Advancements in Molecular Biomarker Testing for Non-Squamous Non-Small Cell Lung Cancer
Published on: September 8, 2023
1.1K
Text analysis framework for identifying mutations among non-small cell lung cancer patients from laboratory data
Amman Yusuf1, Devon J Boyne1,2, Dylan E O'Sullivan1,2
1Department of Oncology, University of Calgary, Calgary, AB, T2N 4N2, Canada.
BMC Medical Research Methodology
|March 12, 2024
Summary
This study developed a text analysis framework to extract valuable information from unstructured biomarker laboratory data, specifically Epidermal Growth Factor Receptor (EGFR) test results, achieving high accuracy for cancer research.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Cancer Research
Background:
- Biomarker laboratory data is crucial for cancer research but often exists as unstructured, free-text data, limiting its utility.
- Extracting information from this unstructured data is essential for advancing cancer research and improving patient outcomes.
- Previous methods for information extraction include Natural Language Processing (NLP), Machine Learning (ML), and rule-based Information Extraction (IE).
Purpose of the Study:
- To develop and evaluate a novel text analysis framework for extracting specific information from unstructured biomarker laboratory data.
- To improve the accessibility and usability of clinical data for cancer research.
- To enable more comprehensive analyses by converting free-text data into structured, actionable information.
Main Methods:
- A rule-based Information Extraction (IE) framework was developed, incorporating Lexical and Syntax analyses.
- Data preprocessing included cleaning, Rich Text Format conversion, normalization, and tokenization.
- Context Free Grammar was used to generate rules for deterministic data extraction, specifically for Epidermal Growth Factor Receptor (EGFR) test results.
Main Results:
- The framework successfully extracted 5129 EGFR test results from the Southern Alberta Dataset (SAD) and 3388 from the Northern Alberta Dataset (NAD).
- The system achieved a high accuracy of 97.5% on a random sample of extracted EGFR test results.
- Extracted data was linked with the Alberta Cancer Registry to support real-world cancer research.
Conclusions:
- A robust text analysis framework was presented for extracting specific information from unstructured clinical data.
- The framework demonstrates significant success in extracting relevant information from EGFR test results.
- This approach enhances the value of laboratory data for cancer research and clinical decision-making.
Keywords:
Context free grammarEGFRInformation extractionLaboratory dataLexical analysisNSCLCSyntax analysisText analysis
