Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Survival Tree01:19

Survival Tree

514
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
514
Randomized Experiments01:13

Randomized Experiments

6.3K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
6.3K
Multiple Regression01:25

Multiple Regression

3.4K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.4K
Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

734
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
734
Random Sampling Method01:09

Random Sampling Method

11.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
11.9K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

7.1K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
7.1K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Enhancing Permeability Prediction of Heterobifunctional Degraders Using Machine Learning and Metadynamics-Informed 3D Molecular Descriptors.

Journal of chemical information and modeling·2025
Same author

Development and Evaluation of Conformal Prediction Methods for Quantitative Structure-Activity Relationship.

ACS omega·2024
Same author

Stability of Prediction in Production ADMET Models as a Function of Version: Why and When Predictions Change.

Journal of chemical information and modeling·2022
Same author

Prediction Accuracy of Production ADMET Models as a Function of Version: Activity Cliffs Rule.

Journal of chemical information and modeling·2022
Same author

Nearest Neighbor Gaussian Process for Quantitative Structure-Activity Relationships.

Journal of chemical information and modeling·2020
Same author

Correction: QSAR without borders.

Chemical Society reviews·2020

Related Experiment Video

Updated: May 6, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.3K

Using random forest to model the domain applicability of another random forest model.

Robert P Sheridan1

  • 1Cheminformatics Department, Merck Research Laboratories , RY800-D133, Rahway, New Jersey 07065, United States.

Journal of Chemical Information and Modeling
|October 25, 2013
PubMed
Summary

Building a quantitative structure-activity relationship (QSAR) error model can be simplified. This study shows a useful error model can be built using just two metrics, TREE_SD and PREDICTED, for QSAR domain applicability.

More Related Videos

Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
12:26

Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM

Published on: October 11, 2016

13.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

898

Related Experiment Videos

Last Updated: May 6, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.3K
Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM
12:26

Integrating Remote Sensing with Species Distribution Models; Mapping Tamarisk Invasions Using the Software for Assisted Habitat Modeling SAHM

Published on: October 11, 2016

13.2K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

898

Area of Science:

  • Computational chemistry
  • Cheminformatics
  • Quantitative Structure-Activity Relationships (QSAR)

Background:

  • QSAR models predict molecular activity using chemical descriptors.
  • Domain applicability aims to assess the reliability of QSAR predictions.
  • Existing methods for assessing reliability often involve multiple metrics.

Purpose of the Study:

  • To develop a quantitative model for predicting QSAR prediction reliability (an "error model").
  • To explore the use of multiple metrics for building robust error models.
  • To determine the most effective metrics for constructing useful error models.

Main Methods:

  • Utilized a random forest QSAR approach to build an error model.
  • Evaluated ten different datasets and seven distinct reliability metrics.
  • Investigated the practicality of using multiple metrics beyond the 'bin paradigm'.

Main Results:

  • Demonstrated the feasibility of constructing a useful error model using QSAR methods.
  • Identified that only two specific metrics (TREE_SD and PREDICTED) were sufficient for building effective error models across examined datasets.
  • These selected metrics avoid computationally expensive similarity/distance calculations.

Conclusions:

  • A quantitative error model for QSAR domain applicability can be effectively built using random forest.
  • The metrics TREE_SD and PREDICTED are highly effective and sufficient for constructing reliable error models.
  • This approach offers a computationally efficient alternative to traditional similarity-based methods.