Related Experiment Video
Updated: Feb 2, 2026

Method for Efficient Refolding and Purification of Chemoreceptor Ligand Binding Domain
Published on: December 12, 2017
Molecular Similarity-Based Domain Applicability Metric Efficiently Identifies Out-of-Domain Compounds
Ruifeng Liu1, Anders Wallqvist1
1Department of Defense Biotechnology High Performance Computing Software Applications Institute, Telemedicine and Advanced Technology Research Center, U.S. Army Medical Research and Materiel Command , Fort Detrick , Maryland 21702 , United States.
Abstract:
Domain applicability (DA) is a concept introduced to gauge the reliability of quantitative structure-activity relationship (QSAR) predictions. A leading DA metric is ensemble variance, which is defined as the variance of predictions by an ensemble of QSAR models. However, this metric fails to identify large prediction errors in melting point (MP) data, despite the availability of large training data sets. In this study, we examined the performance of this metric on MP data and found that, for most molecules, ensemble variance increased as their structural similarity to the training molecules decreased. However, the metric decreased for "out-of-domain" molecules, i.e., molecules with little to no structural similarity to the training compounds. This explains why ensemble variance fails to identify large prediction errors. In contrast, a new molecular similarity-based DA metric that considers the contributions of all training molecules in gauging the reliability of a prediction successfully identified predictions of MP data for which the errors were large. To validate our results, we used four additional data sets of diverse molecular properties. We divided each data set into a training set and a test set at a ratio of approximately 2:1, ensuring a small fraction of the test compounds are out of the training domain. We then trained random forest (RF) models on the training data and made RF predictions for the test set molecules. Results from these data sets confirm that the new DA metric significantly outperformed ensemble variance in identifying predictions for out-of-domain compounds. For within-domain compounds, the two metrics performed similarly, with ensemble variance marginally but consistently outperforming the new DA metric. The new DA metric, which does not rely on an ensemble of QSAR models, can be deployed with any machine-learning method, including deep neural networks.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Membrane Domains
Protein Domains
The membrane comprises a group of distinct proteins responsible for carrying out a cell's specific function. For example, the plasma membrane of the human sperm, or a single germ cell, contains a unique set of proteins in the...
Three Developmental Domains
Physical Development
Physical processes, also known as maturation, encompass the biological changes that occur across an individual's life. These changes begin with genetic inheritance and continue through various stages, including growth in height and weight,...
Three-Domain System of Life
Conservation of Protein Domains
Molecular Compounds: Formulas and Nomenclature

