Related Experiment Video
Updated: Jul 16, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Contrastive diagnostic embedding (CDE) model for automated coding - A case study using emergency department
Amara Tariq1, Kris Goddard1, Praneetha Elugunti1
1Department of Administration, Mayo Clinic, AZ, USA.
Background:
Billing codes are utilized for medical reimbursement, clinical quality metric valuation and for epidemiologic purposes to report and follow disease trends and outcomes. The current paradigm of manual coding can be expensive, time-consuming, and subject to human error. Though automation of the billing codes has been widely reported in the literature via rule-based and supervised approaches, existing strategies lack generalizability and robustness towards large and constantly changing ICD hierarchical structure.
Method:
We propose a weakly supervised training strategy by leveraging contrastive learning, contrastive diagnosis embedding (CDE) to capture the fine semantic variations between the diagnosis codes. The approach consists of a two-phase contrastive training for generating the semantic embedding space adapted to incorporate hierarchical information of ICD-10 vocabulary and a weakly supervised retrieval scheme. Core strength of the proposed method is that it puts no limit on the 70 K ICD-10 codes set and can handle all rare codes for coding the diagnosis.
Results:
Our CDE model outperformed string-based partial matching and ClinicalBERT embedding on three test cases (a retrospective testset, a prospective testset, and external testset) and produced an accurate prediction of rare and newly introduced diagnosis codes. A detailed ablation study showed the importance of each phase of the proposed multi-phase training. Each successive phase of training - ICD-10 group sensitive training (phase 1.1), ICD-10 subgroup sensitive training (phase 1.2), free-text diagnosis description-based training (phase 2) - improved performance beyond the previous phase of training. The model also outperformed existing supervised models like CAML and PLM-ICD and produced satisfactory performance on the rare codes.
Conclusion:
Compared to the existing rule-based and supervised models, the proposed weakly supervised contrastive learning overcomes the limitations in terms of generalization capability and increases the robustness of the automated billing. Such a model will allow flexibility through accurate billing code automation for practice convergence and gains efficiencies in a value-based care payment environment.
More Related Videos
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Related Concept Videos
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Methods of Documentation VII: EMR
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Methods of Documentation III: PIE
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...