Related Experiment Video
Updated: Feb 6, 2026

Pooled shRNA Screen for Reactivation of MeCP2 on the Inactive X Chromosome
Published on: March 2, 2018
Discovering Highly Potent Molecules from an Initial Set of Inactives Using Iterative Screening
Isidro Cortés-Ciriano1, Nicholas C Firth2,3, Andreas Bender1
1Centre for Molecular Informatics, Department of Chemistry , University of Cambridge , Lensfield Road , Cambridge CB2 1EW , United Kingdom.
This study shows simple regression models outperform random forest and similarity searching for early drug discovery when only inactive compounds are known. These methods efficiently identify potent drug candidates by extrapolating from low to high activity ranges.
Area of Science:
- Computational chemistry
- Medicinal chemistry
- Drug discovery
Background:
- Quantitative structure-activity relationships (QSAR) and similarity searching excel at interpolating compound activity within known ranges.
- Their performance in early-stage drug discovery, particularly extrapolating from inactive to active compound data (extrapolation), remains underexplored.
Purpose of the Study:
- To evaluate and compare the effectiveness of various computational methods for virtual screening in early drug discovery scenarios with limited active data.
- To identify the optimal strategy for identifying highly potent compounds through iterative screening when starting with a large set of inactive molecules.
Main Methods:
- An iterative virtual screening strategy was developed and tested on 25 diverse ChEMBL bioactivity datasets.
- Methods benchmarked included random forest (RF), multiple linear regression, ridge regression, similarity searching, and random selection.
- Performance was measured by the number of iterations needed to find a highly active compound.
Main Results:
- Linear and ridge regression methods demonstrated superior performance, reducing the iterations required to find active compounds by over 50% compared to RF and similarity searching.
- Simple regression models showed better extrapolation capabilities to high-bioactivity ranges than RF.
- Scaffold diversity influenced performance, with similarity searching and RF sometimes requiring twice as many iterations as random selection.
Conclusions:
- Linear and ridge regression are highly effective for early-stage drug discovery virtual screening when only inactive data is available.
- The developed iterative screening framework can be extended to multi-target drug discovery, as demonstrated with COX-1 and COX-2 data.
- This study provides a validated approach and identifies optimal computational setups for discovering potent compounds in data-scarce early drug discovery phases.
Related Concept Videos
Initiation of Translation
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
Initiation of Translation
Transcription Initiation
The promoters and enhancers and their accessory proteins allow tight regulation of...
Negative Regulator Molecules
Molecules and Compounds
Setting Time of Cement

