Related Experiment Video
Updated: Jul 31, 2025

Navigating MARRVEL, a Web-Based Tool that Integrates Human Genomics and Model Organism Genetics Information
Published on: August 15, 2019
Bridging the Granularity Gap in Family History Information Extracted from Clinical Narratives
Sungrim Moon1, Liwei Wang1, Xuan Chen2
1Department of Artificial Intelligence and Informatics, Mayo Clinic, Rochester, Minnesota, USA.
Standardizing family history (FH) data from electronic health records (EHRs) is crucial for disease risk assessment. This study automatically mapped FH concepts to various resources, improving data usability for clinical analytics.
Area of Science:
- Medical Informatics
- Genomics and Computational Biology
- Clinical Epidemiology
Background:
- Family history (FH) is vital for assessing disease risk and guiding prevention strategies.
- Integrating FH data from electronic health records (EHRs) into analytics is hindered by a lack of standardization.
- Automating the alignment of FH concepts is necessary for consistent data utilization.
Purpose of the Study:
- To automatically align FH concepts from clinical text to established disease classification resources.
- To evaluate the mapping coverage and granularity of FH concepts across different terminological systems.
- To address the challenge of standardizing NLP-derived FH information for clinical applications.
Main Methods:
- Utilized the Unified Medical Language System (UMLS) to map FH concepts.
- Employed parent and broader/alike relationships within UMLS for concept alignment.
- Compared mapping performance against five resources: Clinical Classification System (CCS), Phecode, Comparative Toxicogenomics Database (CTD), Human phenotype ontology, and Human disease ontology (HDO).
Main Results:
- Achieved high mapping coverage of FH concepts using UMLS.
- Comparative Toxicogenomics Database (CTD) demonstrated the highest coverage (93%) of FH concepts.
- Human disease ontology (HDO) offered the coarsest granularity, while CCS provided the finest granularity of FH disease categories.
Conclusions:
- The study successfully demonstrated a method for standardizing NLP-derived FH data.
- Leveraging UMLS and established ontologies can mitigate granularity challenges in FH analytics.
- This approach enhances the utility of EHR-derived FH information for disease risk assessment and prevention.
More Related Videos
09:43Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
Published on: November 22, 2019
09:34Targeted Next-generation Sequencing and Bioinformatics Pipeline to Evaluate Genetic Determinants of Constitutional Disease
Published on: April 4, 2018
Related Concept Videos
Pedigree Analysis
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genomic Imprinting and Inheritance
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Animal Mitochondrial Genetics
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...