Self-tuned healthy homogeneous core: Addressing heterogeneities in biomedical datasets

Abhidnya Patharkar1, Yutong Wen1, Jiajing Huang2

  • 1School of Computing and Augmented Intelligence, Arizona State University, 699 S Mill Avenue, Tempe, 85281, AZ, USA; ASU-Mayo Center for Innovative Imaging, Arizona State University, Tempe, 85281, AZ, USA.

Insights

Curating a homogeneous healthy core (H2C) improves biomedical classification by reducing noise from heterogeneous healthy data. This method enhances model performance and supports building higher-quality, smaller healthy reference cohorts for machine learning.

Area of Science:

  • Biomedical Machine Learning
  • Computational Biology
  • Data Science

Background:

  • Healthy population heterogeneity can negatively impact disease classification accuracy.
  • Decision boundaries between healthy and diseased groups are often distorted by this heterogeneity.
  • Existing methods may not adequately address the variability within healthy cohorts.

Purpose of the Study:

  • To develop a method for creating a homogeneous healthy reference subset (H2C) from heterogeneous healthy data.
  • To demonstrate that training on H2C improves classification performance compared to using the full healthy cohort.
  • To provide a classifier-agnostic strategy for healthy cohort curation in biomedical modeling.

Main Methods:

  • Decomposition of the healthy cohort into borderline and homogeneous core (H2C) subtypes.
  • Utilizing an information-theoretic framework to analyze classification error bounds.
  • Developing a healthy-only kernel density estimation with bootstrap-stability-based self-tuning.
  • Evaluating the H2C approach on three biomedical datasets: myocardial infarction, arrhythmia, and migraine.

Main Results:

  • Training on H2C consistently improved classification performance (F1, precision, recall, accuracy) across all tested datasets compared to the full healthy baseline.
  • Significant improvements in mean F1 scores were observed: 0.899 to 0.952 on PTB-XL+, 0.693 to 0.854 on Arrhythmia, and 0.643 to 0.741 on Migraine.
  • The H2C method showed the clearest advantage under a balanced training setup and subset-gated evaluation.

Conclusions:

  • Curating a stable, homogeneous healthy reference subset (H2C) enhances downstream classification performance in biomedical machine learning.
  • The H2C strategy offers a practical approach for constructing smaller, higher-quality healthy reference cohorts.
  • This self-tuned, classifier-agnostic method addresses the critical issue of healthy cohort curation in biomedical modeling.

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons01:03

¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons

Protons in identical electronic environments within a molecule are chemically equivalent and have the same chemical shift. The replacement test is a useful tool to identify chemical equivalence and predict NMR spectra. A substituent replaces each of the protons being examined and the resulting molecules are compared. If the same molecule is obtained, the protons are equivalent or homotopic. Replacement of any hydrogens in ethane by chlorine yields chloroethane because all six protons are...
Test for Homogeneity01:23

Test for Homogeneity

The goodness–of–fit test can be used to decide whether a population fits a given distribution, but it will not suffice to decide whether two populations follow the same unknown distribution. A different test, called the test for homogeneity, can be used to conclude whether two populations have the same distribution. To calculate the test statistic for a test for homogeneity, follow the same procedure as with the test of independence. The hypotheses for the test for homogeneity can be stated as...
DNA Microarrays02:34

DNA Microarrays

Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Genetic Variation01:25

Genetic Variation

Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles, which...