Related Experiment Video
Updated: Oct 8, 2026

In Vivo Modeling of the Morbid Human Genome using Danio rerio
Published on: August 24, 2013
Lost in transformation: diagnostic aggregation and semantic drift in ICD-derived kidney phenotypes
Jordan G Nestor1,2, Sarath Babu Krishna Murthy2, Casey Ta3
1Department of Medicine, Division of Nephrology, Columbia University, New York, NY 10032, United States.
Objectives:
To evaluate how cross-terminology mapping of International Classification of Diseases (ICD)-coded diagnoses is associated with preservation of diagnostic distinctions and traceability in electronic health record (EHR)-derived phenotypes for genomic research.
Materials And Methods:
We conducted a retrospective, cross-sectional study of real-world ICD code utilization and semantic fidelity among kidney-related diagnoses. Kidney-relevant ICD-9-CM and ICD-10-CM codes were identified using clinician-curated OHDSI ATLAS searches and evaluated in a large academic health system. ICD codes were mapped to Phecodes, SNOMED-CT, the Human Phenotype Ontology (HPO), and Online Mendelian Inheritance in Man (OMIM) using standardized crosswalks. Mapping coverage, richness, and multiplicity were quantified, with multiplicity operationalizing diagnostic aggregation. Three nephrologists independently rated label-level semantic fidelity (0-2 scale). Ordinal logistic regression evaluated associations between mapping multiplicity and semantic fidelity.
Results:
Among 585 nephrology-relevant ICD codes (571 435 patients), mapping coverage and richness varied across terminologies. Phecodes showed high coverage and mapping multiplicity, whereas SNOMED-CT, HPO, and OMIM mappings were predominantly one-to-one. Among ICD-Phecode mappings, 39% were rated poor, whereas no poor matches were observed among mapped SNOMED-CT, HPO, or OMIM concepts. Greater mapping multiplicity was associated with lower semantic fidelity (OR, 0.33; 95% CI, 0.29-0.37; P < .001). Most ICD-9-CM codes (96%) were represented in at least 1 external cohort.
Discussion:
Diagnostic aggregation was associated with reduced semantic fidelity and traceability, with potential implications for phenotype construction and interpretation in GWAS, PheWAS, and biobank-scale studies.
Conclusion:
This study provides an empirical framework for evaluating semantic drift in EHR-derived phenotypes and highlights a central tradeoff: abstraction enables scalability, whereas traceability enables translation.