Scalable Identification of Clinically Relevant Chronic Obstructive Pulmonary Disease Documents in Large-Scale
Mohammed Al-Garadi1, Sharon E Davis1, Michael E Matheny1,2,3,4
1Department of Biomedical Informatics, Vanderbilt University Medical Center, 2525 West End Ave, Suite 1475, Nashville, TN, 37203, United States, 1 (615) 936-6867.
This study introduces a new method to filter clinical documents for relevance, improving data quality for machine learning. The approach uses lightweight embedding and random forest classification for better accuracy in identifying important patient information.
Area of Science:
- Clinical Informatics
- Machine Learning
- Natural Language Processing
Background:
- Electronic health records generate vast amounts of clinical notes, posing challenges for data analysis.
- Learning algorithms and large language models are sensitive to noise and irrelevant data, leading to performance degradation and inaccuracies.
- Efficiently filtering millions of clinical documents for relevance is crucial, especially in resource-constrained environments.
Purpose of the Study:
- To develop a novel framework for determining document relevance in clinical settings.
- To address the challenge of efficiently filtering large volumes of clinical notes for advanced language model processing.
- To improve the accuracy and efficiency of information retrieval from electronic health records.
Main Methods:
- Developed a framework using weak supervision and domain-expert heuristics to create "silver standard" labels.
- Utilized various text representation techniques including bag of words, TF-IDF, lightweight document embeddings, and UMLS concept extraction.
- Trained and evaluated random forest, extreme gradient boosting, and k-nearest neighbor classifiers on expert-annotated and held-out test datasets.
Main Results:
- The combination of lightweight document embedding and a random forest classifier achieved the highest performance.
- This model demonstrated a precision of 0.73, recall of 0.86, and F1-score of 0.80 for identifying relevant chronic obstructive pulmonary disease (COPD) documents.
- The proposed framework significantly outperformed baseline heuristics and other tested methods.
Conclusions:
- A novel framework for identifying COPD-relevant clinical documents was presented, utilizing lightweight embedding and machine learning.
- This approach effectively filters pertinent documents, enhancing information retrieval precision and scalability.
- The framework's minimal annotation needs show promise for diverse healthcare applications, potentially optimizing clinical outcomes through efficient document selection.
Related Concept Videos
Chronic Obstructive Pulmonary Disease-IV: Assessement and Diagnostic Studies
Medical History
Chronic Obstructive Pulmonary Disease-I: Introduction
Chronic Obstructive Pulmonary Disease I: Introduction
Chronic Obstructive Pulmonary Disease
Smoking is a primary risk factor for COPD, with over 80% of patients having a history of it. Patients typically experience progressive dyspnea or labored breathing, frequent coughing, and recurrent pulmonary infections. Many eventually succumb to respiratory failure, characterized by...
Chronic Obstructive Pulmonary Disease III: Chronic Bronchitis Features
Chronic Obstructive Pulmonary Disease-V: Management
Smoking Cessation
