Related Experiment Video
Updated: Sep 22, 2025

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
16.0K
Facilitating clinical research through automation: Combining optical character recognition with natural language
Julie Hom1, Janet Nikowitz1, Rebecca Ottesen1
1Department of Diabetes & Cancer Discovery Science, City of Hope, Duarte, CA, USA.
Clinical Trials (London, England)
|May 24, 2022
Summary
This study developed an optical character recognition (OCR) and natural language processing (NLP) pipeline to efficiently extract crucial patient performance status data from scanned electronic medical records, significantly reducing data abstraction time and improving research accuracy.
Area of Science:
- Medical Informatics
- Clinical Research Data Management
- Natural Language Processing in Healthcare
Background:
- Performance status is vital for clinical research but often unstructured in electronic medical records.
- Scanned documents and image formats hinder data extraction and NLP application.
- Efficient retrieval of performance status data is challenging for research purposes.
Purpose of the Study:
- To develop and evaluate an optical character recognition (OCR) and natural language processing (NLP) pipeline for extracting performance status data.
- To improve the efficiency and accuracy of data abstraction from scanned electronic medical records.
- To reduce the time and effort required for clinical research data retrieval.
Main Methods:
- Utilized optical character recognition (OCR) software (ABBYY FineReader) to convert scanned medical record images into a searchable format.
- Applied natural language processing (NLP) software (Linguamatics i2e) to identify and extract performance status data elements.
- Evaluated the pipeline's accuracy and efficiency using a cohort of diffuse large B-cell lymphoma patients, comparing against manual abstraction.
Main Results:
- The OCR/NLP pipeline demonstrated high accuracy with less than 1% incorrect results (excellent precision, recall, and F score).
- Median review time for documents containing performance status data was reduced by one-third.
- Manual review time for documents without performance status information was reduced by 83% (18 minutes vs. 108 minutes).
Conclusions:
- The OCR/NLP pipeline significantly improves operational efficiency and reduces information retrieval time for clinical research.
- OCR effectively transforms scanned medical record images for NLP application, enabling highly accurate data abstraction.
- This pipeline facilitates research by enabling focused data review, eliminating unnecessary chart review, and freeing up time for other critical data elements.
Keywords:
Eastern Cooperative Oncology GroupKarnofsky performance statusScanned medical recordsnatural language processingoptical character recognitionperformance status
