Related Experiment Video
Updated: Aug 23, 2026

Imaging and Analysis for Quantifying Maize (Zea mays) Abiotic Stress Phenotypes
Published on: March 28, 2025
Comparative identification of abiotic stress-responsive differentially expressed genes in chickpea using unsupervised
Mohammad Shafiq1,2, Rimshah Sabir1,2, Yossma Waheed1,2
1Higher Engineering School of Agrobiotechnology, National Research Tomsk State University, Lenin Ave, 36, Tomsk, Tomsk Oblast, Russia, 634050.
Background:
Climate change poses a growing threat to chickpea production through abiotic stresses such as drought, heat, and salinity. Understanding the molecular stress response of chickpea is critical for improving resilience and minimizing yield loss.
Objective:
To identify robust, stress-responsive differentially expressed genes (DEGs) in chickpea under abiotic stress (drought, salt, and salinity) by comparing three complementary large-scale RNA-seq data analytical approaches.
Methods:
This study employed three complementary approaches, traditional meta-analysis, conventional statistical testing, and unsupervised machine learning, on publicly available chickpea RNA-seq data to identify robust differentially expressed genes (DEGs) under drought, salt, and salinity stress. Available RNA-seq datasets were integrated regardless of variety, tissue type, or geographical origin. The HDBSCAN clustering algorithm was optimized through distance metric evaluation and grid search hyperparameter tuning, with Euclidean distance optimal for drought and salinity and Manhattan distance for salt stress.
Results:
Standard deviation-based feature engineering on the top 3,000 most variable genes yielded the most stress-specific DEGs, with high fold enrichments for cytochrome P450, phenylpropanoid biosynthesis, and heme binding under drought. For salt and salinity stress, limited sample availability constrained the feature space, reducing HDBSCAN clustering resolution and DEG specificity demonstrating that ML performance scales markedly with sample size and feature richness, where DESeq2 showed comparatively greater robustness. For the drought dataset spanning 11 bioprojects, Limma-voom and DESeq2 with bioproject correction both returned non-specific housekeeping enrichment, in direct contrast to the stress-specific signal recovered by HDBSCAN, empirically demonstrating the superior biological specificity of unsupervised ML-based outlier detection in heterogeneous multi-bioproject data.
Conclusion:
Compared to HN-score meta-analysis, HDBSCAN demonstrated superior stress specificity by leveraging complex high-dimensional expression patterns, offering a powerful and scalable strategy for stress-responsive gene identification in chickpea and other crops.

