Related Experiment Video
Updated: Mar 31, 2026

Using Human Differentially Expressed Gene Lists to Perform Downstream Pathway Enrichment Analysis and Target Prioritization
Published on: October 3, 2025
Target prediction utilising negative bioactivity data covering large chemical space
Lewis H Mervin1, Avid M Afzal1, Georgios Drakakis1
1Department of Chemistry, Centre for Molecular Informatics, University of Cambridge, Lensfield Road, Cambridge, CB2 1EW UK.
This study integrates inactive compound data into in silico target prediction models, significantly improving accuracy. The novel approach enhances predictions for orphan compounds, aiding drug discovery efforts.
Area of Science:
- Computational chemistry
- Cheminformatics
- Pharmacology
Background:
- In silico analyses are crucial for drug discovery but often underutilize inactive compound data.
- Chemogenomic repositories contain vast amounts of bioactivity data, including inactive compounds, which can enhance predictive models.
- Target prediction for orphan compounds is challenging and can benefit from comprehensive data integration.
Purpose of the Study:
- To integrate large-scale inactive bioactivity data into in silico target prediction models.
- To develop a novel method for predicting the probability of activity and inactivity for a range of targets.
- To improve the prediction of drug targets for orphan compounds using comprehensive datasets.
Main Methods:
- Constructed a novel human bioactivity dataset from ChEMBL and PubChem, incorporating over 195 million data points.
- Applied a sphere-exclusion algorithm to oversample presumed inactive compounds.
- Trained and evaluated a Bernoulli Naïve Bayes algorithm using fivefold cross-validation and external datasets (WOMBAT).
Main Results:
- The model achieved high precision (99.7%) and recall (99.6%) for inactive compounds and good performance for active compounds (67.7% recall, 63.8% precision).
- Model performance was influenced by training data similarity, class size, and oversampling.
- External validation using WOMBAT data yielded improved precision-recall AUC (0.56) and BEDROC (0.85) scores compared to models trained solely on active data (0.45 AUC, 0.76 BEDROC).
Conclusions:
- Incorporating inactive data into in silico models significantly enhances target prediction accuracy, improving AUC and early recognition capabilities.
- Model performance varied based on internal and external validation, highlighting the importance of robust validation strategies.
- The developed target prediction protocol (PIDGIN) is publicly available, facilitating its use in drug discovery research.
More Related Videos
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023