Integrating Mutation-Derived and Expression Features from Single-Cell RNA Sequencing: Pitfalls of Standard
Aidyn Kunikeyev1, Amankeldi A Salybekov2, Aigerim Yerimbetova3,4
1Institute of Automation and Information Technologies, Satbayev University, Almaty 050013, Kazakhstan.
International Journal of Molecular Sciences
|July 28, 2026
Summary
Standard cross-validation is unreliable for small-cohort single-cell RNA sequencing (scRNA-seq) machine learning. Our analysis highlights pitfalls and proposes a robust, auditable framework for hypothesis generation.
Area of Science:
- Genomics
- Computational Biology
- Bioinformatics
Background:
- Single-cell RNA sequencing (scRNA-seq) integrates gene expression and mutation data.
- Small cohorts with repeated measures pose cross-validation challenges, risking optimistic bias.
Purpose of the Study:
- Evaluate mutation-derived, expression-only, and combined features in scRNA-seq.
- Assess cross-validation pitfalls in small-cohort machine learning.
- Establish a reliable framework for hypothesis generation in scRNA-seq.
Main Methods:
- Reanalyzed PRJNA736095 dataset using GATK for variant calling and gene-burden analysis.
- Employed leakage-safe preprocessing within validation folds.
- Utilized run-level and GSM-grouped cross-validation, including permutation testing.
Main Results:
- High within-dataset accuracy for GATK gene-burden features (0.973 +/- 0.113) with run-level validation.
- Significantly reduced accuracy (0.708) with GSM-grouped validation; permutation testing was non-significant (p=0.257).
- Expression-only and combined features did not outperform variant-only approaches; expression-mutation overlap was not significant.
Conclusions:
- Standard cross-validation is overly optimistic for small-cohort scRNA-seq.
- The proposed workflow offers an auditable, hypothesis-generating framework.
- Careful validation strategies are crucial for reliable machine learning in scRNA-seq.

