Related Experiment Video
Updated: Feb 22, 2026

Analyzing Multifactorial RNA-Seq Experiments with DiCoExpress
Published on: July 29, 2022
Statistically controlled identification of differentially expressed genes in one-to-one cell line comparisons of the
Jun He1, Haidan Yan1, Hao Cai1
1Department of Bioinformatics, Key Laboratory of Ministry of Education for Gastrointestinal Cancer, Fujian Medical University, Fuzhou, 350122, China.
Background:
The Connectivity Map (CMAP) database, an important public data source for drug repositioning, archives gene expression profiles from cancer cell lines treated with and without bioactive small molecules. However, there are only one or two technical replicates for each cell line under one treatment condition. For such small-scale data, current fold-changes-based methods lack statistical control in identifying differentially expressed genes (DEGs) in treated cells. Especially, one-to-one comparison may result in too many drug-irrelevant DEGs due to random experimental factors. To tackle this problem, CMAP adopts a pattern-matching strategy to build "connection" between disease signatures and gene expression changes associated with drug treatments. However, many drug-irrelevant genes may blur the "connection" if all the genes are used instead of pre-selected DEGs induced by drug treatments.
Methods:
We applied OneComp, a customized version of RankComp, to identify DEGs in such small-scale cell line datasets. For a cell line, a list of gene pairs with stable relative expression orderings (REOs) were identified in a large collection of control cell samples measured in different experiments and they formed the background stable REOs. When applying OneComp to a small-scale cell line dataset, the background stable REOs were customized by filtering out the gene pairs with reversal REOs in the control samples of the analyzed dataset.
Results:
In simulated data, the consistency scores of overlapping genes between DEGs identified by OneComp and SAM were all higher than 99%, while the consistency score of the DEGs solely identified by OneComp was 96.85% according to the observed expression difference method. The usefulness of OneComp was exemplified in drug repositioning by identifying phenformin and metformin related genes using small-scale cell line datasets which helped to support them as a potential anti-tumor drug for non-small-cell lung carcinoma, while the pattern-matching strategy adopted by CMAP missed the two connections. The implementation of OneComp is available at https://github.com/pathint/reoa .
Conclusions:
OneComp performed well in both the simulated and real data. It is useful in drug repositioning studies by helping to find hidden "connections" between drugs and diseases.
Insights
OneComp effectively identifies differentially expressed genes (DEGs) in small-scale cell line datasets, improving drug repositioning by revealing hidden drug-disease connections missed by current methods.
Area of Science:
- Genomics
- Bioinformatics
- Pharmacology
Background:
- The Connectivity Map (CMAP) database aids drug repositioning but has limited replicates in gene expression data.
- Current methods for identifying differentially expressed genes (DEGs) in small datasets lack statistical control and can yield irrelevant genes.
- CMAP's pattern-matching strategy can be obscured by drug-irrelevant genes when identifying drug-disease connections.
Purpose of the Study:
- To develop a statistically robust method for identifying DEGs in small-scale cell line gene expression datasets.
- To enhance the accuracy of drug repositioning by improving the identification of drug-induced gene expression changes.
- To overcome limitations of existing methods in analyzing limited experimental replicates.
Main Methods:
- Applied OneComp, a customized version of RankComp, to identify DEGs in small-scale cell line datasets.
- Established background stable relative expression orderings (REOs) from large control datasets.
- Customized background REOs by filtering gene pairs with reversed REOs in control samples of the analyzed dataset.
Main Results:
- OneComp demonstrated high consistency with existing methods (SAM) on simulated data (over 99% overlap).
- OneComp identified DEGs solely with high consistency (96.85%) on simulated data.
- Successfully identified phenformin and metformin as potential anti-tumor drugs for non-small-cell lung carcinoma, connections missed by CMAP.
Conclusions:
- OneComp performs effectively on both simulated and real-world small-scale gene expression data.
- The method enhances drug repositioning by uncovering previously hidden drug-disease associations.
- OneComp offers a valuable tool for analyzing limited biological datasets in drug discovery.

