Related Experiment Video
Updated: Apr 11, 2026

02:44
Author Spotlight: Enhanced Isolation of Interaction-Null Mutants in Yeast
Published on: December 29, 2023
1.1K
Leakage-safe diffusion augmentation with KAN-based models for imbalanced microarray gene-expression classification
Bich-Chung Phan1, Thanh Ma1, Thanh-Nghi Do1
1College of Information and Communication Technology, Can Tho University, 3/2 Street, Can Tho City, 900000, Viet Nam.
Computational Biology and Chemistry
|April 9, 2026
Summary
FoDiKAN, a novel framework, enhances gene-expression classification on imbalanced microarray data by integrating diffusion augmentation with Kolmogorov-Arnold Networks (KANs). It achieves superior performance by preventing information leakage during cross-validation.
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- Gene-expression classification is crucial for functional genomics, disease subtyping, and biomarker discovery.
- Public microarray datasets often present challenges like high dimensionality, small sample sizes, and severe class imbalance.
- Existing methods can suffer from information leakage, leading to unreliable performance estimates and biological interpretations.
Purpose of the Study:
- To introduce FoDiKAN, a leakage-aware framework for robust gene-expression classification.
- To develop algorithms (SafeCV and AugTrain) that integrate fold-local diffusion augmentation with Kolmogorov-Arnold Networks (KANs).
- To ensure reliable performance estimation and interpretation in high-dimensional, imbalanced transcriptomic data.
Main Methods:
- SafeCV implements leakage-safe outer cross-validation, fitting operators on fold-local inner-training splits.
- AugTrain employs hybrid gene selection (mRMR, Boruta), trains a diffusion model, and generates synthetic minority samples using anchored DDIM sampling.
- Synthetic samples are used exclusively for training with down-weighting, integrated within a KAN backbone.
Main Results:
- The best FoDiKAN configuration achieved a macro-F1 score of 88.6% across 25 microarray datasets.
- This performance surpassed fixed-reference Gradient Boosting (by 4.2%) and XGBoost (by 2.4%) while competing with balanced baselines.
- Analyses quantified real-synthetic mismatch and assessed design choices through ablation and sensitivity studies.
Conclusions:
- FoDiKAN provides a robust and leakage-aware approach for gene-expression classification on challenging microarray data.
- The framework effectively leverages diffusion augmentation and KANs to improve classification performance and reliability.
- The study highlights the importance of leakage-aware methodologies and provides diagnostics for interpreting augmentation effects.

