A deep learning architecture for combining and imputing heterogeneous metabolomics datasets
Sadi Celik1, Baris Can1, Mehmet Ali Erdogan1
1Istanbul Technical University, 34467, Maslak, Istanbul, Turkey.
BMC Bioinformatics
|July 17, 2026
Summary
This study introduces novel methods for merging sparse metabolomics datasets, improving machine learning model training. These approaches enhance data imputation for better disease mechanism analysis.
Area of Science:
- Metabolomics
- Bioinformatics
- Computational Biology
Background:
- Public metabolomics databases contain numerous datasets valuable for understanding disease mechanisms.
- Existing datasets are often sparse, measuring only a fraction of metabolites, hindering machine learning model development.
- Combining sparse datasets directly results in high dimensionality and missing values, posing challenges for analysis.
Purpose of the Study:
- To develop and evaluate novel methods for merging sparse metabolomics datasets.
- To improve the imputation of missing metabolite data for enhanced machine learning model training.
- To effectively combine diverse metabolomics datasets while minimizing data gaps.
Main Methods:
- Proposing two novel dataset merging approaches: Iterative similarity-based merging and Model-guided agglomerative merging.
- Ensuring a minimum sparsity threshold is maintained during dataset merging.
- Employing Variational Autoencoders (VAE) for imputation model training on merged datasets.
- Evaluating the methods on the complete set of public datasets from the Metabolomics Workbench.
Main Results:
- The proposed merging and imputation methods significantly outperform the state-of-the-art approach.
- Achieved substantially improved imputation performance on sparse metabolomics data.
- Demonstrated the effectiveness of combining diverse datasets while managing metabolite overlap.
Conclusions:
- The novel merging strategies effectively address sparsity in public metabolomics datasets.
- VAE-based imputation on merged datasets enhances the ability to analyze complex molecular mechanisms in diseases.
- These methods provide a robust framework for leveraging large-scale metabolomics data for biomedical research.
