Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Variability: Analysis01:11

Variability: Analysis

163
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
163
Classification of Systems-I01:26

Classification of Systems-I

236
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
236
Classification of Signals01:30

Classification of Signals

574
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
574
Aggregates Classification01:29

Aggregates Classification

355
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
355
Classification of Systems-II01:31

Classification of Systems-II

192
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
192
Force Classification01:22

Force Classification

1.3K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Evolution of assay interference concepts in drug discovery.

Expert opinion on drug discovery·2021
Same author

Adapting the DeepSARM approach for dual-target ligand design.

Journal of computer-aided molecular design·2021
Same author

Data set of competitive and allosteric protein kinase inhibitors confirmed by X-ray crystallography.

Data in brief·2021
Same author

Evaluation of multi-target deep neural network models for compound potency prediction under increasingly challenging test conditions.

Journal of computer-aided molecular design·2021
Same author

Predicting Isoform-Selective Carbonic Anhydrase Inhibitors via Machine Learning and Rationalizing Structural Features Important for Selectivity.

ACS omega·2021
Same author

Systematic comparison of competitive and allosteric kinase inhibitors reveals common structural characteristics.

European journal of medicinal chemistry·2021

Related Experiment Video

Updated: Aug 3, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K

Differences in learning characteristics between support vector machine and random forest models for compound

Friederike Maite Siemers1, Jürgen Bajorath2

  • 1B-IT, LIMES Program Unit Chemical Biology and Medicinal Chemistry, Department of Life Science Informatics and Data Science, Rheinische Friedrich-Wilhelms-Universität, Friedrich-Hirzebruch-Allee 5/6, 53115, Bonn, Germany.

Scientific Reports
|April 12, 2023
PubMed
Summary

Random Forest (RF) and Support Vector Machine (SVM) models show similar compound predictions but differ in learning characteristics. Explainable AI methods reveal distinct origins for accurate predictions from these molecular machine learning algorithms.

More Related Videos

Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

20.0K
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

144

Related Experiment Videos

Last Updated: Aug 3, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.6K
Cross-Modal Multivariate Pattern Analysis
13:51

Cross-Modal Multivariate Pattern Analysis

Published on: November 9, 2011

20.0K
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

144

Area of Science:

  • Molecular machine learning
  • Computational chemistry
  • Cheminformatics

Background:

  • Random Forest (RF) and Support Vector Machine (SVM) are key algorithms in molecular machine learning (ML) for predicting compound properties.
  • Understanding the prediction mechanisms of these ML models is crucial for reliable compound property prediction.

Purpose of the Study:

  • To investigate the prediction mechanisms of RF and SVM binary classification models in molecular ML.
  • To apply and extend explainable artificial intelligence (XAI) techniques, specifically Shapley values, for model interpretability.

Main Methods:

  • Utilized RF and SVM algorithms for large-scale activity-based compound classification.
  • Employed Shapley value analysis, adapted from game theory, to interpret model predictions.
  • Trained models using training sets of increasing size to observe learning characteristics.

Main Results:

  • RF and SVM models with the Tanimoto kernel produced highly similar predictions in compound classification tasks.
  • Shapley value analysis demonstrated systematic differences in the learning characteristics of RF and SVM.
  • Chemically intuitive explanations for accurate predictions derived from RF and SVM models originated from different feature attributions.

Conclusions:

  • Despite similar predictive performance, RF and SVM models exhibit distinct internal learning processes in molecular ML.
  • XAI methods like Shapley values are essential for uncovering these underlying differences and providing chemically meaningful interpretations.
  • This research enhances the understanding of how ML models make predictions in drug discovery and chemical informatics.