Related Experiment Video
Updated: Aug 28, 2025

Author Spotlight: Efficient Image Recognition Using Directional Gradient Histogram Technique and Support Vector Machines
Published on: January 5, 2024
Calculation of exact Shapley values for support vector machines with Tanimoto kernel enables model interpretation
Christian Feldmann1, Jürgen Bajorath1
1Department of Life Science Informatics and Data Science, B-IT, LIMES Program Unit Chemical Biology and Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany.
We developed SVETA for exact Shapley value calculation in Support Vector Machine (SVM) models. This method provides reliable explanations for drug discovery predictions, unlike approximate methods.
Area of Science:
- Computational Chemistry
- Cheminformatics
- Drug Discovery
Background:
- Support Vector Machine (SVM) algorithms are widely used in chemistry and drug discovery but often function as black boxes.
- Interpreting SVM predictions typically involves feature weighting or model-agnostic methods like Shapley Additive Explanations (SHAP), which approximate Shapley values (SVs).
- Accurate explanation of complex models is crucial for rationalizing predictions in sensitive fields like drug discovery.
Purpose of the Study:
- To introduce a novel algorithm, SV-expressed Tanimoto similarity (SVETA), for the exact calculation of Shapley values (SVs) in SVM models.
- To assess the reliability of exact SVs compared to approximate SHAP values in a drug discovery context.
- To demonstrate the utility of SVETA in providing consistent and interpretable explanations for SVM-based compound classification.
Main Methods:
- Development and implementation of the SV-expressed Tanimoto similarity (SVETA) algorithm for exact SV calculation.
- Application of SVETA to an SVM model utilizing the Tanimoto kernel for molecular similarity assessment.
- Comparison of exact SVs derived from SVETA with SHAP values in an SVM-based compound classification task.
- Atom-based mapping of prioritized features to identify substructures responsible for predictions.
Main Results:
- The SVETA algorithm enables the exact calculation of Shapley values for SVM models with Tanimoto kernels.
- A limited correlation was observed between exact SVs and SHAP values in a drug discovery classification task, indicating potential limitations of SHAP for rationalizing these specific SVM predictions.
- Atom-based feature mapping using exact SVs identified coherent substructures, consistent with explanations from independent Random Forest models.
Conclusions:
- SVETA provides an accurate method for explaining SVM models in cheminformatics and drug discovery.
- Exact Shapley values are essential for reliable interpretation, as approximate methods like SHAP may not fully capture the model's decision-making process.
- The consistency of explanations across different modeling techniques (SVM with SVETA and Random Forest) enhances confidence in the identified molecular substructures.
Related Concept Videos
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Calculating and Interpreting the Linear Correlation Coefficient
Chebyshev's Theorem to Interpret Standard Deviation
Kendall's Tau Test
A τ value...

