Related Experiment Video
Updated: Aug 17, 2025

Biobank for Translational Medicine: Standard Operating Procedures for Optimal Sample Management
Published on: November 30, 2022
Mapping of UK Biobank clinical codes: Challenges and possible solutions
Oleg Stroganov1, Alena Fedarovich1, Emily Wong2
1Rancho BioSciences, LLC, San Diego, California, United States of America.
UK Biobank data requires cleaning and standardization for research use. This study proposes a pipeline to harmonize diagnoses by mapping Read codes to SNOMED CT and ICD, addressing data inconsistencies.
Area of Science:
- Biomedical Informatics
- Data Science
- Clinical Data Management
Background:
- The UK Biobank dataset is a valuable resource but contains heterogeneous and inconsistent medical terminology.
- Data cleaning, curation, and standardization are essential to make UK Biobank data usable for research.
- Integrating diverse data sources like hospital records, self-reported data, and primary care data requires significant harmonization efforts.
Purpose of the Study:
- To develop and propose a data integration pipeline for harmonizing diagnostic information within the UK Biobank.
- To evaluate methods for mapping primary care (GP) Read codes to standardized terminologies like ICD and SNOMED CT.
- To identify and report challenges encountered during the mapping and harmonization process.
Main Methods:
- Evaluation of multiple approaches for mapping GP clinical Read codes to International Classification of Diseases (ICD) and Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT).
- Comparison of mapping results, flagging inconsistencies, and assigning quality categories to assess overall mapping quality.
- Development of a curation and data integration pipeline focused on harmonizing diagnostic terms.
Main Results:
- Several approaches for mapping Read codes to ICD and SNOMED CT were evaluated.
- Mapping inconsistencies were identified and categorized by quality.
- Challenges in mapping, including lack of one-to-one ontology mapping and the need for additional ontologies, were reported.
Conclusions:
- A curation and data integration pipeline is proposed for harmonizing diagnoses from UK Biobank GP data.
- Challenges in mapping Read codes to ICD and SNOMED CT were identified, some general and some specific to automated mapping.
- Leveraging existing mappings and employing automated/manual curation can overcome mapping challenges.
More Related Videos
05:49Use of Magnetic Resonance Imaging and Biopsy Data to Guide Sampling Procedures for Prostate Cancer Biobanking
Published on: October 10, 2019
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020