Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multiple Regression01:25

Multiple Regression

3.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.3K
The Equilibrium Binding Constant and Binding Strength02:18

The Equilibrium Binding Constant and Binding Strength

10.6K
The equilibrium binding constant (Kb) quantifies the strength of a protein-ligand interaction. Kb can be calculated as follows when the reaction is at equilibrium:
10.6K
Ligand Binding Sites02:40

Ligand Binding Sites

11.8K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
11.8K
Conserved Binding Sites01:49

Conserved Binding Sites

4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Regression Analysis01:11

Regression Analysis

7.2K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.2K
Survival Tree01:19

Survival Tree

498
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
498

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Prognostic Impact of Sarcopenia Following Myocardial Infarction: A Systematic Review and Meta-analysis of Mortality and Recurrent MI.

European journal of preventive cardiology·2026
Same author

High-performance tunnel-junction micro-LEDs grown by MOCVD without post-growth annealing.

Optics express·2026
Same author

Molecular mechanism of MAFB transcriptional activation of PPARD in regulating adipose browning and protecting against vascular endothelial cell injury.

Experimental cell research·2026
Same author

Integrating rumen microbiome and host metabolome to investigate feed conversion ratio across different fattening stages in Hu sheep.

Animal bioscience·2026
Same author

Mechanistic Insights into the Transient Reactions of Environmentally Persistent Free Radicals on Common Microplastics: An Important Role of Air Humidity.

Environmental science & technology·2026
Same author

Comment on "Periodizing Exercise Medicine Prescription for Patients with Cancer: A Narrative Opinion".

Sports medicine (Auckland, N.Z.)·2026

Related Experiment Video

Updated: Apr 25, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K

Substituting random forest for multiple linear regression improves binding affinity prediction of scoring functions:

Hongjian Li1, Kwong-Sak Leung, Man-Hon Wong

  • 1Department of Computer Science and Engineering, Chinese University of Hong Kong, Hong Kong, China. jackyleehongjian@gmail.com.

BMC Bioinformatics
|August 28, 2014
PubMed
Summary

Classical scoring functions for protein-ligand docking have accuracy limitations. Machine learning, specifically random forest (RF), significantly improves binding affinity prediction by capturing complex relationships, outperforming traditional methods like multivariate linear regression (MLR).

More Related Videos

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

872

Related Experiment Videos

Last Updated: Apr 25, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.0K
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

872

Area of Science:

  • Computational chemistry
  • Drug discovery
  • Bioinformatics

Background:

  • Protein-ligand docking is crucial for drug discovery, but scoring function accuracy limits its effectiveness.
  • Classical scoring functions, often using multivariate linear regression (MLR), have plateaued in predictive performance.
  • These methods rely on predetermined additive functional forms for numerical features.

Purpose of the Study:

  • To investigate the performance of machine learning techniques, specifically random forest (RF), in improving protein-ligand binding affinity prediction.
  • To compare RF performance against traditional MLR and established scoring functions like Cyscore.

Main Methods:

  • Applied random forest (RF) machine learning to predict binding affinities from structural features.
  • Investigated the impact of training sample size and the number of structural features on RF performance.
  • Utilized RF variable importance tool to analyze feature contributions.
  • Compared RF performance against multivariate linear regression (MLR) and Cyscore.

Main Results:

  • Random forest (RF) significantly outperforms multivariate linear regression (MLR) in predicting binding affinities.
  • RF effectively captures non-linear relationships between structural features and binding affinities, given sufficient training data.
  • Increased structural features and training samples enhance RF predictive performance.
  • RF analysis identified key structural features influencing binding affinity.

Conclusions:

  • Machine learning scoring functions, like RF, offer a fundamental advantage over classical methods by avoiding fixed functional forms.
  • RF demonstrates superior ability to leverage more structural features and training data for higher prediction accuracy.
  • The growing availability of structural data will further enhance RF-based scoring functions, highlighting the need to replace MLR in scoring function development.