Related Experiment Video
Updated: Sep 26, 2025

Prediction and Validation of Gene Regulatory Elements Activated During Retinoic Acid Induced Embryonic Stem Cell Differentiation
Published on: June 21, 2016
Development and Experimental Validation of Regularized Machine Learning Models Detecting New, Structurally Distinct
Steffen Hirte1, Oliver Burk2, Ammar Tahir3
1Division of Pharmaceutical Chemistry, Department of Pharmaceutical Sciences, Faculty of Life Sciences, University of Vienna, 1090 Vienna, Austria.
Abstract:
The pregnane X receptor (PXR) regulates the metabolism of many xenobiotic and endobiotic substances. In consequence, PXR decreases the efficacy of many small-molecule drugs and induces drug-drug interactions. The prediction of PXR activators with theoretical approaches such as machine learning (ML) proves challenging due to the ligand promiscuity of PXR, which is related to its large and flexible binding pocket. In this work we demonstrate, by the example of random forest models and support vector machines, that classifiers generated following classical training procedures often fail to predict PXR activity for compounds that are dissimilar from those in the training set. We present a novel regularization technique that penalizes the gap between a model's training and validation performance. On a challenging test set, this technique led to improvements in Matthew correlation coefficients (MCCs) by up to 0.21. Using these regularized ML models, we selected 31 compounds that are structurally distinct from known PXR ligands for experimental validation. Twelve of them were confirmed as active in the cellular PXR ligand-binding domain assembly assay and more hits were identified during follow-up studies. Comprehensive analysis of key features of PXR biology conducted for three representative hits confirmed their ability to activate the PXR.
Insights
Predicting pregnane X receptor (PXR) activators is difficult. A new machine learning regularization technique improves prediction accuracy for novel compounds, identifying new PXR activators.
Area of Science:
- Pharmacology
- Computational Chemistry
- Drug Discovery
Background:
- The pregnane X receptor (PXR) plays a crucial role in metabolizing xenobiotics and endobiotic substances.
- PXR activation can reduce the efficacy of small-molecule drugs and lead to drug-drug interactions.
- Predicting PXR activators using machine learning (ML) is challenging due to PXR's ligand promiscuity and flexible binding pocket.
Purpose of the Study:
- To develop and validate a novel regularization technique for machine learning models to predict PXR activators.
- To improve the prediction of PXR activity for compounds structurally dissimilar to those in training datasets.
- To identify novel PXR activators through validated computational methods.
Main Methods:
- Implementation and evaluation of random forest and support vector machine models.
- Development of a novel regularization technique penalizing the gap between training and validation performance.
- Experimental validation of computationally selected compounds using cellular PXR ligand-binding domain assembly assays.
Main Results:
- Classical ML training procedures showed limitations in predicting PXR activity for dissimilar compounds.
- The novel regularization technique improved Matthew correlation coefficients (MCCs) by up to 0.21 on a challenging test set.
- Twelve out of 31 structurally distinct compounds selected by regularized ML models were confirmed as PXR activators.
Conclusions:
- The developed regularization technique enhances the predictive power of ML models for PXR activators.
- This approach successfully identified novel PXR-activating compounds with potential therapeutic implications.
- The findings contribute to more accurate in silico drug screening and drug-drug interaction prediction.

