Validation of a Text-Mining Tool for Extracting Routine Clinical Care Data in Early-Stage Resectable Non-Small Cell
Hanieh Abedian Kalkhoran1,2, Tobias Martinot3, Lydia H N Schonewille4
1Department of Clinical Pharmacy and Toxicology, Leiden University Medical Center, Leiden, the Netherlands.
JCO Clinical Cancer Informatics
|July 29, 2026
Summary
Natural language processing (NLP) software CTcue shows promise for extracting electronic health record data in early-stage non-small cell lung cancer research. While manual validation is still needed for some variables, CTcue offers accurate and efficient data extraction for real-world data studies.
Area of Science:
- Medical Informatics
- Oncology
- Natural Language Processing (NLP)
Background:
- Manual chart review (MR) of electronic health records (EHRs) is inefficient and hinders real-world data (RWD) research reproducibility.
- Natural language processing (NLP) offers potential for automating and standardizing EHR data extraction.
- CTcue is an NLP platform designed for structured and unstructured EHR data extraction.
Purpose of the Study:
- To evaluate the accuracy and efficiency of the CTcue NLP platform compared to manual chart review (MR).
- To assess CTcue's performance in extracting data for patients with early-stage resectable non-small cell lung cancer (NSCLC).
- To determine the suitability of CTcue for improving the scalability of RWD research.
Main Methods:
- Retrospective study of stage I-III NSCLC patients undergoing lung resection (January 2018 - December 2021) at Leiden University Medical Center.
- Comparison of CTcue data extraction against manual chart review (MR) for demographics, tumor characteristics, treatment, and outcomes.
- Performance metrics included weighted F1-scores, accuracy, precision, recall for categorical variables, and Bland-Altman analysis for continuous variables.
Main Results:
- Eighty-five patients were included in the comparison.
- CTcue achieved weighted F1-scores >0.85 for 7 of 15 categorical variables, with lower performance for variables with inconsistent documentation (e.g., ECOG status, N-stage).
- Continuous variables showed good agreement with negligible mean differences, and survival outcomes were identical between CTcue and MR.
Conclusions:
- CTcue demonstrates accurate and efficient extraction of structured and unstructured EHR data for early-stage NSCLC.
- Manual validation is still required for variables with variable terminology in clinical documentation.
- Further AI development is crucial for enhancing accuracy and scalability in RWD research, especially for free-text data.

