Related Experiment Video
Updated: Apr 8, 2026

Reduced Procedure Time and Variability with Active Esophageal Cooling During Radiofrequency Ablation for Atrial Fibrillation
Published on: August 25, 2022
An Expedited Chart Review Process for Large Database Studies Using Natural Language Processing and Multiwave Adaptive
Shirley V Wang1, Georg Hahn1, Sushama Kattinakere Sreedhara1
1From the Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA.
Background:
One of the ways to enhance analyses conducted with large claims databases is by validating the measurement characteristics of the code-based algorithms used to identify health outcomes or other key study parameters of interest. These metrics can be used in quantitative bias analyses to assess the robustness of results for an inferential study, given potential bias from outcome misclassification. However, performing this validation through manual chart review of free-text notes from linked electronic health records requires extensive time and resource allocation.
Methods:
We describe an expedited process for validating code-based algorithms that introduces efficiency using two distinct mechanisms: (1) use of natural language processing to reduce the time spent by human reviewers to review each chart, and (2) a multiwave adaptive sampling approach with predefined criteria to stop the validation study once performance characteristics are identified with sufficient precision. We illustrate this process in a case study that validates the performance of a claims-based outcome algorithm for intentional self-harm in patients with obesity.
Results:
We empirically demonstrate that the natural language processing-assisted annotation process reduced the time spent on review per chart by 40%, and the use of the predefined stopping rule with multiwave samples would have prevented review of 77% of patient charts with limited compromise to the precision of performance characteristics.
Conclusion:
This approach could facilitate more routine validation of code-based algorithms used to define key study parameters, ultimately enhancing understanding of the reliability of findings derived from database studies.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019