Related Experiment Videos
Validation of natural language processing for surgical complication surveillance: detecting 11 postoperative
Emilie Even Dencker1, Alexander Bonde1, Anders Troelsen2
1Department of Organ Surgery and Transplantation, Rigshospitalet, Copenhagen, Denmark.
Abstract:
INTRODUCTION Postoperative complications (PCs) rates are crucial quality metrics in surgery, as they reflect both patient outcomes, perioperative care effectiveness and healthcare resource strain. Despite their importance, efficient, accurate and affordable methods for tracking PCs are lacking. This study aimed to evaluate whether natural language processing (NLP) models could detect 11 PCs from surgical electronic health records at a level comparable to human curation. RESEARCH AND DESIGN METHODS Retrospective study in 18 hospitals across two regions in Denmark. A total of 17 486 surgical cases spanning 6 years were included. The dataset was divided into training, validation and test sets for NLP-model development and evaluation (50.2%/33.6%/16.2%). Model performance was compared against the current method of PC monitoring (International Classification of Diseases, 10th Revision (ICD-10) codes) and manual curation, the latter serving as the gold standard. 17 486 surgical cases from spanning 6 years were included. The dataset was divided into training, validation and test sets for NLP-model development and evaluation (50.2%/33.6%/16.2%). Model performance was compared against the current method of PC monitoring (International Classification of Diseases, 10th Revision (ICD-10) codes) and manual curation, the latter serving as the gold standard. RESULTS The NLP-models had a receiver operating characteristic area under the curve between 0.901 and 0.999 for the test set and significantly outperformed ICD-10 coding in detecting PCs. Sensitivity of the models when compared with manual curation ranged from 0.701 to 1.00, except for myocardial infarction (0.500). Positive predictive value (PPV) ranged from 0.0165 to 0.947, and negative predictive value from 0.995 to 1.00. Using a Human-in-the-Loop approach, only 16.3% of cases required manual review to reach a PPV of 1.00. CONCLUSIONS The NLP models alone were able to detect PCs at an acceptable level and outperformed ICD-10 codes. While combining NLP with manual review (Human-in-the-Loop) improved overall accuracy and reduced workload, the models still failed to identify some complications. Therefore, NLP algorithms may support (but not replace) manual surveillance and present a potential solution for more scalable PC monitoring.
Related Concept Videos
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters assessment...
Methods of Documentation VII: EMR