Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Classification of Illness01:17

Classification of Illness

8.4K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
8.4K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

841
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
841
Cluster Sampling Method01:20

Cluster Sampling Method

13.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.9K
Diagnostic and Statistical Manual of Mental Disorders (DSM)01:27

Diagnostic and Statistical Manual of Mental Disorders (DSM)

775
The Diagnostic and Statistical Manual of Mental Disorders (DSM) serves as the primary classification system for mental health disorders, providing standardized diagnostic criteria for clinicians and researchers. First published by the American Psychiatric Association (APA) in 1952, the DSM has undergone several revisions to reflect evolving psychiatric understanding. The fifth edition, DSM-5, released in 2013, introduced key updates that expanded diagnostic categories and modified diagnostic...
775
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

42.2K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
42.2K
Formulating and Validating Nursing Diagnosis II01:25

Formulating and Validating Nursing Diagnosis II

3.5K
Nursing diagnoses represent a problem validated by major defining characteristics. There are four categories of nursing diagnoses: problem-focused, risk, health promotion or wellness, and syndrome. The anatomy of a nursing diagnosis includes three components: problem statement or diagnostic label, defining characteristics, and related factors.
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multimodal prediction for radiotherapy-induced hematologic toxicity in rectal cancer patients.

NPJ precision oncology·2026
Same author

Interactive Fluorescence Cell Counting via User-Guided Correction.

IEEE transactions on bio-medical engineering·2026
Same author

Missing value replacement in strings and applications.

Data mining and knowledge discovery·2025
Same author

DiagSWin: A multi-scale vision transformer with diagonal-shaped windows for object detection and segmentation.

Neural networks : the official journal of the International Neural Network Society·2024
Same author

Clustering Demographics and Sequences of Diagnosis Codes.

IEEE journal of biomedical and health informatics·2021
Same author

Anonymizing datasets with demographics and diagnosis codes in the presence of utility constraints.

Journal of biomedical informatics·2016

Related Experiment Video

Updated: Dec 31, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K

Clustering datasets with demographics and diagnosis codes.

Haodi Zhong1, Grigorios Loukides1, Robert Gwadera2

  • 1Department of Informatics, King's College London, London, UK.

Journal of Biomedical Informatics
|January 7, 2020
PubMed
Summary

This study introduces a novel method for clustering Electronic Health Record (EHR) data, effectively grouping patients by demographics and diagnosis codes. The approach enhances data analysis and classification tasks by creating meaningful patient clusters.

Keywords:
ClusteringDemographicsDiagnosis codesPattern mining

More Related Videos

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.3K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.3K

Related Experiment Videos

Last Updated: Dec 31, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

15.0K
Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.3K
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
07:50

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts

Published on: September 20, 2018

16.3K

Area of Science:

  • Health Informatics
  • Data Science
  • Computational Biology

Background:

  • Clustering Electronic Health Record (EHR) data is crucial for understanding patient clinical profiles and aiding downstream analysis like classification.
  • The heterogeneity of EHR data presents significant challenges for traditional clustering algorithms, necessitating novel approaches.
  • Existing methods struggle with the complex, multi-faceted nature of patient data, including demographic information and diagnosis codes.

Purpose of the Study:

  • To propose and evaluate a novel clustering approach specifically designed for heterogeneous EHR data.
  • To enable the discovery of relationships between patient demographics and diagnosis codes within large datasets.
  • To provide an efficient and scalable method for patient data clustering as a preprocessing step for analysis.

Main Methods:

  • Representing EHR data in a binary format, incorporating selected demographic values.
  • Identifying and utilizing combinations of frequent and correlated diagnosis codes as features.
  • Employing cosine similarity for measuring record similarity in the binary representation.
  • Applying hierarchical clustering to identify compact and well-separated patient clusters.

Main Results:

  • The proposed approach successfully constructed clusters exhibiting correlated demographics and diagnosis codes.
  • Experiments on two large, publicly available EHR datasets (over 26,000 and 52,000 records) validated the method's effectiveness.
  • The clustering approach demonstrated efficiency and scalability, suitable for large-scale EHR data analysis.

Conclusions:

  • The novel binary representation and hierarchical clustering method effectively addresses the challenges of EHR data heterogeneity.
  • This approach facilitates the discovery of meaningful patient subgroups based on integrated demographic and diagnostic information.
  • The method offers a scalable and efficient solution for preprocessing EHR data, improving subsequent analytical tasks.