Related Experiment Video
Updated: Oct 2, 2026

CorrelationCalculator and Filigree: Tools for Data-Driven Network Analysis of Metabolomics Data
Published on: November 10, 2023
A high-dimensional benchmark of objective functions and biological resolutions for personalized metabolic phenotyping
Mehmet Ali Erdoğan1, Ali Cakmak1
1Department of Computer Engineering, Istanbul Technical University, Ayazağa Campus, Reşitpaşa Mahallesi, Sarıyer, 34469 Istanbul, Türkiye.
Abstract:
Genome-scale personalized metabolic modeling relies on objective functions to simulate intracellular flux phenotypes, yet the optimal formulation for characterizing complex human diseases remains unknown. We ask which combination of objective function, biological resolution, and feature-selection strategy yields the most reliable diagnostic signal, treating machine-learning architecture as a controlled factor. We benchmarked these choices across 57 600 experimental configurations spanning six diverse pathologies, from oncology to neurodegeneration. Within this benchmark, pathway-level minimum reaction flux was the best-performing representation, indicating that diagnostic signals are more strongly associated with rate-limiting metabolic bottlenecks than with aggregate pathway activity. Under a leakage-free evaluation, diagnostic accuracy was largely insensitive to feature density (mean $F_{1}$-score varying by <0.003 between the top-10% subset and the full feature space), so aggressive feature selection was not required. Among classifiers, only Logistic Regression lost accuracy at full density, whereas SVM was essentially unaffected and tree-based ensembles were resilient. Statistical regularization (robust_sigma) ensured consistent generalization, while topological objectives (local_k) were highly sensitive to metabolic noise, excelling only in specific high-signal cohorts. Finally, integrating objective function values as competitive features recovered the regulatory context lost during pathway aggregation. Accurate metabolic phenotyping was best captured by isolating restrictive bottleneck reactions. Predictive accuracy was largely insensitive to feature density; tree-based architectures tolerated dense omics data, while feature selection benefited mainly Logistic Regression. Overall, this benchmark offers systematic, evidence-based guidance for personalized metabolic modeling, complementing heuristic parameter tuning with quantitative comparison.

