mbSparse: an autoencoder-based imputation method to address sparsity in microbiome data

Changlu Qi1, Yiting Cai1, Guoyou He1

  • 1College of Bioinformatics Science and Technology, Harbin Medical University, Harbin, HL, China.

Gut Microbes
|September 1, 2025
PubMed

Insights

We developed mbSparse, a deep learning algorithm, to address zero-inflation in microbiome data. This method significantly improves imputation accuracy and enhances disease detection in complex datasets.

Area of Science:

  • Microbiome research
  • Bioinformatics
  • Computational biology

Background:

  • Gut microbiota plays a vital role in host physiology.
  • High sparsity (numerous zeros) in microbiome data poses significant analytical challenges.
  • Existing methods struggle with accurate imputation of sparse microbiome data.

Purpose of the Study:

  • To develop a novel deep learning-based algorithm, mbSparse, for accurate imputation of sparse microbiome data.
  • To evaluate the performance of mbSparse compared to existing methods.
  • To assess the utility of mbSparse in a colorectal cancer analysis.

Main Methods:

  • Developed mbSparse, an imputation algorithm using a feature autoencoder and a conditional variational autoencoder (CVAE).
  • Leveraged deep learning for learning sample representations and data reconstruction.
  • Applied mbSparse to simulated and real microbiome datasets, including colorectal cancer data.

Main Results:

  • mbSparse achieved superior imputation accuracy, reducing mean squared error by up to 4.1 compared to existing methods.
  • In colorectal cancer analysis, mbSparse increased the detection of disease-associated taxa from 7 to 27 and improved predictive accuracy (AUC from 0.85 to 0.93).
  • mbSparse effectively restored over 88% of removed counts, preserving taxonomic relationships with a Pearson correlation of 0.9354.

Conclusions:

  • mbSparse offers a powerful deep learning solution for accurate microbiome data imputation, overcoming challenges posed by data sparsity.
  • The CVAE component is crucial for mbSparse's enhanced accuracy.
  • mbSparse improves biological insights and predictive power in microbiome-associated disease studies.

Related Concept Videos

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation01:24

One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation

This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
708
Estimating Population Mean with Unknown Standard Deviation01:22

Estimating Population Mean with Unknown Standard Deviation

In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
8.3K
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Truncation in Survival Analysis01:09

Truncation in Survival Analysis

Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
299
Bootstrapping01:24

Bootstrapping

The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
670
Microbial Growth Measurement: Indirect Methods01:27

Microbial Growth Measurement: Indirect Methods

Estimating microbial growth is essential for understanding population dynamics and environmental adaptations. Indirect methods provide valuable insights by measuring parameters such as turbidity, metabolic activity, and biomass, enabling efficient and reproducible assessments.During exponential growth, microbial cells scatter light proportionally to their biomass, a principle used in turbidity measurements. About one million cells per milliliter produce detectable scattering, which a...
162