Related Experiment Video
Updated: May 12, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
CATI: A medical context-enhanced framework for diagnosis code assignment in the UK Biobank study
Yue Shen1, Jie Wang1, Zhe Wang1
1MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, Hefei, Anhui, 230027, China.
Insights
This study introduces CATI, a novel framework for assigning accurate diagnosis codes in large biobanks by integrating textual data and disease hierarchies. CATI improves upon existing methods, enhancing disease-related research through better patient cohort formation.
Area of Science:
- Bioinformatics
- Medical Informatics
- Computational Biology
Background:
- Accurate diagnosis coding is essential for large-scale biobank studies and downstream disease research.
- Current methods often neglect valuable textual medical information and the hierarchical nature of disease classification systems.
- Missing diagnosis codes in biobanks limit the utility of patient data for clinical research.
Purpose of the Study:
- To develop and evaluate CATI, a medical context-enhanced framework for improved diagnosis code assignment in large biobanks.
- To integrate textual details and disease hierarchy into a unified model for more accurate coding.
- To address the challenge of missing diagnosis codes for patients in biobank datasets.
Main Methods:
- CATI framework integrates textual details from key features and disease hierarchy using text embeddings from BioBERT via prompt tuning.
- A novel convolution layer is developed to effectively propagate information between adjacent diagnosis codes, leveraging the hierarchical structure.
- The study utilizes UK Biobank data, focusing on Phecodes and ICD-10 codes as standard disease formats.
Main Results:
- CATI demonstrates superior performance compared to state-of-the-art methods for both Phecodes and ICD-10 code assignment.
- Achieved at least a 5.16% improvement in average AUROC for unseen disease codes.
- Showcased an 8.68% increase in average AUPRC for disease codes with limited training instances (1000-10000).
Conclusions:
- CATI effectively enhances diagnosis code assignment by incorporating medical context from textual data and disease hierarchies.
- The framework facilitates the creation of well-defined cohorts for downstream biomedical research.
- CATI offers a valuable approach for complex healthcare tasks by leveraging rich medical information.
Abstract:
Diagnosis codes are standard code format of diseases or medical conditions. This study is aimed at assigning diagnosis codes to patients in large-scale biobanks, particularly addressing the issue of missing codes for some patients. This is crucial for downstream disease-related tasks. While recent methods primarily rely on structured biobank data for code assignment, they often overlook the valuable medical context provided by textual information in the biobanks and hierarchical structure of the disease coding system. To address this gap, we have developed CATI, a medical context-enhanced framework for diagnosis Code Assignment by integrating Textual details derived from key features and disease hIerarchy. The study is based on the UK Biobank data and considers Phecodes and ICD-10 codes as standard disease formats. We start by representing ten informative codified features using their formal names and then integrate them into CATI as text embeddings, achieved through prompt tuning on the pre-trained language model BioBERT. Recognizing the hierarchical structure of diagnosis codes, we have developed a novel convolution layer in our method that effectively propagates logits between adjacent diagnosis codes. Evaluation results demonstrate that CATI outperforms existing state-of-the-art methods in terms of both Phecodes and ICD-10 codes, boasting at least a 5.16% improvement in average AUROC for unseen disease codes and an 8.68% rise in average AUPRC for disease codes with training instances ranging in (1000,10000]. This framework contributes to the formation of well-defined cohorts for downstream studies and offers a unique perspective for addressing complex healthcare tasks by incorporating vital medical context.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
05:49Use of Magnetic Resonance Imaging and Biopsy Data to Guide Sampling Procedures for Prostate Cancer Biobanking
Published on: October 10, 2019
Related Concept Videos
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Taxonomy