通过精确的Shapley值计算解释支向量机器预测的协议
Andrea Mastropietro1, Jürgen Bajorath2
1Deparment of Computer, Control and Management Engineering "Antonio Ruberti", Sapienza University of Rome, Via Ariosto 25, 00185 Rome, Italy.
本研究介绍了一个准确计算Shapley值的协议,以解释机器学习预测,特别是支持向量机器. 该方法提供定量特征分析和可视化,以便更好地解释模型.
科学领域:
- 计算化学是一种计算化学.
- 机器学习是机器学习.
- 化学信息学 化学信息学
背景情况:
- 沙普利值通常用于解释机器学习 (ML) 预测,但对于大型特征集来说通常是近似值.
- 精确计算Shapley值是计算密集的,特别是对于复杂的模型,如支持向量机 (SVMs).
研究的目的:
- 介绍一个准确的Shapley值计算协议,以解释SVM预测.
- 为应用这些算法提供实用的工具和方法.
- 为了实现定量特征分析和重要特征的可视化.
主要方法:
- 开发了两种技术的协议,使SVM能够精确计算Shapley值.
- 提供了准备好使用的Python脚本和定制代码以实现.
- 专注于解释ML中大型特征集的预测.
主要成果:
- 该协议允许精确计算Shapley值,克服近似限制.
- 生成定量特征分析和特征重要性映射用于可视化.
- 证明了对SVM的精确沙普利值计算的应用.
结论:
- 精确的沙普利值计算是可行的,并有利于解释SVM预测.
- 提供的协议和脚本有助于准确的特征分析和模型可解释性.
- 这种方法增强了对ML模型中特征贡献的理解.
更多相关视频
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
07:12Using Informational Connectivity to Measure the Synchronous Emergence of fMRI Multi-voxel Information Across Time
Published on: July 1, 2014
相关概念视频
Variation
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
Chebyshev's Theorem to Interpret Standard Deviation
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
