Related Experiment Video
Updated: Aug 28, 2026

Three Differential Expression Analysis Methods for RNA Sequencing: limma, EdgeR, DESeq2
Published on: September 18, 2021
Revisiting differential expression analysis: An updated six-dimensional comparative study
Jianxiong Wu1,2, Shaoke Lu1, Hui Yao1
1Department of Colorectal Surgery and Oncology of the Second Affiliated Hospital, Centre of Biomedical Systems and Informatics of Zhejiang University-University of Edinburgh Institute (ZJU-UoE Institute), Zhejiang University School of Medicine, Zhejiang University, Hangzhou, Zhejiang, China.
Abstract:
Differential expression (DE) analysis is probably the most prevalent task for transcriptomic studies. However, recent technological advances have seen a revival of methodological interest in DE algorithms. In this study, we performed a comprehensive updated comparative study of 12 representative DE methods using 80 simulated and real datasets. We assessed the adaptability of these methods across varying sample sizes and diverse data scenarios. This evaluation compiled a six-dimensional overview of key properties: detection accuracy, sensitivity at a low false discovery rate, false positives, stability, robustness to outliers, and robustness under noisy conditions. Strikingly, no single methods outperformed others across all evaluation criteria and sample sizes, emphasizing data-specific and scenario-specific method choice. At the widely adopted small-sample size of n = 3, ABSSeq generally outperformed other methods. As sample size increased to n = 5, the sensitivity of DESeq2 and two edgeR v4 algorithms (QLF slightly better than LRT) also raise up under a stringent false-positive control. DESeq had even fewer false positives than DESeq2, at the price of reduced sensitivity. In terms of robustness, Wilcoxon and ROTS are robust to noises for small sample sizes. Moreover, Wilcoxon is also robust to outliers, together with several other methods (ABSSeq, voom, and T.test). NBPSeq and most methods had a good stability even at small sample sizes, except three methods (ROTS, DSS, and T.test). For larger sample sizes (n > 30), all methods performed much better. Finally, we provided a "BaGua (eight trigrams)" map summarizing the multi-dimensional performances of methods, as well as a tree diagram guiding practical method selection. Together, this study outlines a systematic and updated benchmarking framework for DE analysis, emphasizing a balance between accuracy and consistency.

