Text mining applied to electronic cardiovascular procedure reports to identify patients with trileaflet aortic
Aeron M Small1, Daniel H Kiss1, Yevgeny Zlatsin2
1Department of Medicine and Cardiovascular Institute, University of Pennsylvania Perelman School of Medicine, PA, USA.
Insights
Text mining of cardiovascular procedure reports significantly improves patient identification for research compared to traditional billing codes. This method offers higher accuracy in detecting conditions like aortic stenosis and coronary artery disease.
Area of Science:
- Cardiology
- Medical Informatics
- Computational Biology
Background:
- Electronic health records (EHRs) are crucial for clinical research, but using billing codes for diagnoses has variable accuracy.
- Text mining of EHRs has shown mixed success for identifying cardiovascular phenotypes.
- Cardiovascular procedure reports may offer a more accurate data source for patient identification.
Purpose of the Study:
- To evaluate the effectiveness of text mining algorithms applied to cardiovascular procedure reports for identifying patients with specific cardiovascular conditions.
- To compare the accuracy of text mining with traditional billing codes (ICD-9) for identifying patients with trileaflet aortic stenosis (TAS) and coronary artery disease (CAD).
Main Methods:
- Adapted text mining tool (PennSeek) to search cardiovascular procedure reports (echocardiography and cardiac catheterization).
- Imported 282,569 echocardiography and 27,205 cardiac catheterization reports.
- Applied clinical criteria to identify patients with TAS and CAD, comparing text mining results with ICD-9 billing codes.
Main Results:
- Text mining identified 7115 patients with TAS and 9247 with CAD.
- ICD-9 codes identified 8272 patients with TAS and 6913 with CAD.
- Text mining demonstrated superior positive predictive values: 0.95 for TAS (vs. 0.53 for ICD-9) and 0.97 for CAD (vs. 0.86 for ICD-9).
Conclusions:
- Text mining applied to electronic cardiovascular procedure reports is a superior method for identifying patient phenotypes for cardiovascular research.
- This approach offers significantly higher accuracy than using billing codes alone.
- Enhances the reliability of patient cohort selection in cardiovascular research.
Background:
Interrogation of the electronic health record (EHR) using billing codes as a surrogate for diagnoses of interest has been widely used for clinical research. However, the accuracy of this methodology is variable, as it reflects billing codes rather than severity of disease, and depends on the disease and the accuracy of the coding practitioner. Systematic application of text mining to the EHR has had variable success for the detection of cardiovascular phenotypes. We hypothesize that the application of text mining algorithms to cardiovascular procedure reports may be a superior method to identify patients with cardiovascular conditions of interest.
Methods:
We adapted the Oracle product Endeca, which utilizes text mining to identify terms of interest from a NoSQL-like database, for purposes of searching cardiovascular procedure reports and termed the tool "PennSeek". We imported 282,569 echocardiography reports representing 81,164 individuals and 27,205 cardiac catheterization reports representing 14,567 individuals from non-searchable databases into PennSeek. We then applied clinical criteria to these reports in PennSeek to identify patients with trileaflet aortic stenosis (TAS) and coronary artery disease (CAD). Accuracy of patient identification by text mining through PennSeek was compared with ICD-9 billing codes.
Results:
Text mining identified 7115 patients with TAS and 9247 patients with CAD. ICD-9 codes identified 8272 patients with TAS and 6913 patients with CAD. 4346 patients with AS and 6024 patients with CAD were identified by both approaches. A randomly selected sample of 200-250 patients uniquely identified by text mining was compared with 200-250 patients uniquely identified by billing codes for both diseases. We demonstrate that text mining was superior, with a positive predictive value (PPV) of 0.95 compared to 0.53 by ICD-9 for TAS, and a PPV of 0.97 compared to 0.86 for CAD.
Conclusion:
These results highlight the superiority of text mining algorithms applied to electronic cardiovascular procedure reports in the identification of phenotypes of interest for cardiovascular research.
More Related Videos
06:16Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
08:10Estimating Bilateral Atrial Function by Cardiovascular Magnetic Resonance Feature Tracking in Patients with Paroxysmal Atrial Fibrillation
Published on: July 20, 2022
Related Concept Videos
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Imaging Studies for Cardiovascular System V: CT
Acute Coronary Syndrome III: Diagnostic Studies
