Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Enhancing Explainable AI Stability with Realistic Synthetic Data for Cardiovascular Risk Prediction.

Studies in health technology and informatics·2026
Same author

Knowledge graph embedding and alignment of incomplete electronic health records for critical care applications.

Journal of biomedical semantics·2026
Same author

Comparative study on imaging characteristics of pilocytic astrocytomas in children and adolescents.

American journal of cancer research·2026
Same author

Rule-augmented constraint learning for semantic error detection in MIMIC-III knowledge graph.

International journal of medical informatics·2026
Same author

CMEO: a metadata-centric ontology for clinical studies exploration and harmonization assessment.

BMC medical informatics and decision making·2025
Same author

Privacy preservation in blockchain-based healthcare data sharing: A systematic review.

Peer-to-peer networking and applications·2025

Related Experiment Video

Updated: Feb 22, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.6K

Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata.

Wei Hu1, Amrapali Zaveri2, Honglei Qiu1

  • 1State Key Laboratory for Novel Software Technology, Nanjing University, 163 Xianlin Avenue, Nanjing, 210023, Jiangsu, China.

BMC Bioinformatics
|September 20, 2017
PubMed
Summary

This study introduces a clustering algorithm to improve the quality of biomedical metadata, addressing issues like redundancy and inconsistency in large datasets such as the Gene Expression Omnibus (GEO). The method effectively groups similar metadata keys, enhancing data discoverability for researchers.

Keywords:
BiomedicalClusteringData qualityExperimental dataGEOMetadataReusability

More Related Videos

Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
07:11

Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis

Published on: November 10, 2023

3.4K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.3K

Related Experiment Videos

Last Updated: Feb 22, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.6K
Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
07:11

Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis

Published on: November 10, 2023

3.4K
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
09:20

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications

Published on: February 23, 2019

9.3K

Area of Science:

  • Bioinformatics
  • Data Science
  • Genomics

Background:

  • High-quality metadata is crucial for efficient searching and filtering of biomedical datasets.
  • The Gene Expression Omnibus (GEO) contains over 44 million key-value pairs suffering from quality issues like redundancy and inconsistency.
  • Existing metadata quality issues hinder researchers' ability to find relevant datasets.

Purpose of the Study:

  • To develop a clustering-based approach to address data quality issues in gene expression metadata.
  • To improve the accuracy, structure, and completeness of biomedical metadata descriptions.

Main Methods:

  • Proposed a clustering-based approach using three similarity measures for metadata keys.
  • Designed a scalable agglomerative clustering algorithm to group similar keys.

Main Results:

  • The algorithm successfully clustered similar metadata keys based on name, core concept, and value.
  • Evaluated using a gold standard, the algorithm generated 18 clusters with 355 keys.
  • Achieved a superior average F-Score of 0.63 compared to four other methods.

Conclusions:

  • The clustering algorithm effectively identifies and groups similar metadata keys, resolving scalability issues for data cleaning.
  • This approach facilitates the identification of duplicate and erroneous metadata entries.
  • The algorithm is adaptable for use with other biomedical data types.