Related Experiment Video
Updated: Feb 9, 2026

Automated Cell Enrichment of Cytomegalovirus-specific T cells for Clinical Applications using the Cytokine-capture System
Published on: October 5, 2015
Automating Electronic Clinical Data Capture for Quality Improvement and Research: The CERTAIN Validation Project of
Emily Beth Devine1, Erik Van Eaton1, Megan E Zadworny1
1University of Washington, US.
Background:
The availability of high fidelity electronic health record (EHR) data is a hallmark of the learning health care system. Washington State's Surgical Care Outcomes and Assessment Program (SCOAP) is a network of hospitals participating in quality improvement (QI) registries wherein data are manually abstracted from EHRs. To create the Comparative Effectiveness Research and Translation Network (CERTAIN), we semi-automated SCOAP data abstraction using a centralized federated data model, created a central data repository (CDR), and assessed whether these data could be used as real world evidence for QI and research.
Objectives:
Describe the validation processes and complexities involved and lessons learned.
Methods:
Investigators installed a commercial CDR to retrieve and store data from disparate EHRs. Manual and automated abstraction systems were conducted in parallel (10/2012-7/2013) and validated in three phases using the EHR as the gold standard: 1) ingestion, 2) standardization, and 3) concordance of automated versus manually abstracted cases. Information retrieval statistics were calculated.
Results:
Four unaffiliated health systems provided data. Between 6 and 15 percent of data elements were abstracted: 51 to 86 percent from structured data; the remainder using natural language processing (NLP). In phase 1, data ingestion from 12 out of 20 feeds reached 95 percent accuracy. In phase 2, 55 percent of structured data elements performed with 96 to 100 percent accuracy; NLP with 89 to 91 percent accuracy. In phase 3, concordance ranged from 69 to 89 percent. Information retrieval statistics were consistently above 90 percent.
Conclusions:
Semi-automated data abstraction may be useful, although raw data collected as a byproduct of health care delivery is not immediately available for use as real world evidence. New approaches to gathering and analyzing extant data are required.
More Related Videos
06:05The Participant-Reported Implementation Update and Score PRIUS: A Novel Method for Capturing Implementation-Related Data Over Time
Published on: February 19, 2021
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Related Concept Videos
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
The Evidence for Evolution
Statistical Software for Data Analysis and Clinical Trials
Reliability and Validity
Fischer Projections