Related Experiment Video
Updated: Sep 22, 2026

Integration of Wet and Dry Bench Processes Optimizes Targeted Next-generation Sequencing of Low-quality and Low-quantity Tumor Biopsies
Published on: April 11, 2016
Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to
Qianyun Luo1, Rui Zhang2,3,4, Nikitha Vobugari5,4
1University of Minnesota Medical School, Minneapolis, MN, USA.
Purpose:
Real-world oncology data are essential for clinical research and precision cancer care. However, genomic biomarkers are often embedded in scanned, unstructured clinical documents requiring manual abstraction before becoming available in cancer registries, delaying real-world evidence generation. This study evaluated and compared three open-source optical character recognition (OCR) approaches, Tesseract, EasyOCR, and a hybrid implementation, to determine which best enables automated extraction of Oncotype DX recurrence scores and improves the timeliness and quality of real-world oncology data.
Methods:
We evaluated the feasibility of automated genomic data extraction using 675 Oncotype DX reports from a Midwestern U.S. health system. EasyOCR, Tesseract, and a hybrid OCR approach were used to extract recurrence scores from scanned reports. OCR-derived values were compared with manually abstracted scores and local cancer registry data. Performance was assessed using agreement, precision, recall, F1 score, and processing time. Multivariable logistic regression was performed to identify factors associated with discordance between registry-reported and manually abstracted scores.
Results:
The hybrid OCR approach demonstrated the highest performance, achieving 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. Registry abstraction demonstrated comparable performance but required greater manual effort. Automated extraction substantially reduced processing time while maintaining high accuracy. Logistic regression showed registry discordance was largely independent of patient and tumor characteristics, with unknown progesterone receptor (PR) status as the only significant predictor.
Conclusion:
Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019