Related Experiment Video
Updated: Jun 5, 2026

08:12
A Psychophysics Paradigm for the Collection and Analysis of Similarity Judgments
Published on: March 1, 2022
A note on two novel easy-to-interpret feature effect measures for partial dependence plots in a classification
1Regional Cancer Centre Stockholm-Gotland, Region Stockholm, Stockholm, Sweden.
Journal of Applied Statistics
|June 4, 2026
Summary
New measures, relative risk of marginal effects (RRME) and odds ratio of marginal effects (ORME), interpret black box models. These methods offer robust predictions and handle missing data better than traditional models.
Area of Science:
- Statistics
- Machine Learning
- Biostatistics
Background:
- Parametric models like logistic regression are traditional for classification tasks.
- Black box supervised learning models (BBSLMs) often exceed parametric models in prediction accuracy.
- BBSLMs lack easily interpretable feature effect measures comparable to logistic regression's odds ratio (OR).
Purpose of the Study:
- To derive novel feature effect measures for binary classification using BBSLMs.
- To introduce the relative risk of marginal effects (RRME) and odds ratio of marginal effects (ORME).
- To assess the performance and interpretability of these new measures.
Main Methods:
- Derivation of RRME and ORME based on partial dependence plots for BBSLMs.
- Application to a dataset of patients admitted to hospital with myocardial infarction.
- Comparison of BBSLMs with RRME/ORME against traditional logistic regression models.
Main Results:
- BBSLMs demonstrated superior predictive ability compared to logistic regression.
- The derived RRME and ORME for anterior infarct (a key risk factor) were 1.8, comparable to the logistic regression OR of 1.9.
- RRME and ORME proved more robust, handling observations with missing feature values effectively.
Conclusions:
- The novel RRME and ORME measures provide interpretable feature effects for BBSLMs in binary classification.
- These measures offer comparable interpretability to traditional OR while leveraging the superior predictive power of BBSLMs.
- The robustness of RRME and ORME to missing data enhances their practical applicability in statistical modeling.
Related Concept Videos
Two-Way ANOVA
The two-way ANOVA is an extension of the one-way ANOVA. It is a statistical test performed on three or more samples categorized by two factors - a row factor and a column factor. Ronald Fischer mentioned it in 1925 in his book 'Statistical Methods for Researchers.'
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
The two-way ANOVA analysis initially begins by stating the null hypothesis that there is an interaction effect between the two factors of a dataset. This effect can be visualized using line segments formed by joining the means for...
Factorial Design
Factorial Analysis is an experimental design that applies Analysis of Variance (ANOVA) statistical procedures to examine a change in a dependent variable due to more than one independent variable, also known as factors. Changes in worker productivity can be reasoned, for example, to be influenced by salary and other conditions, such as skill level. One way to test this hypothesis is by categorizing salary into three levels (low, moderate, and high) and skills sets into two levels (entry level...
Friedman Two-way Analysis of Variance by Ranks
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures from...
Residual Plots
A residual plot is a statistical representation of data used to analyze correlation and regression results. It helps verify the requirements for drawing specific conclusions about correlation and regression. To obtain the residual plot, first, the residual for each data value is calculated, which is simply the vertical distance between the observed and the predicted value obtained from the regression equation.
When the residual values are plotted against the variable x, it is called a residual...
When the residual values are plotted against the variable x, it is called a residual...
Introduction to Test of Independence
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
Scatter Plot
The most common and easiest way to display the relationship between two variables, x and y, is a scatter plot. A scatter plot shows the direction of a relationship between the variables. A clear direction happens when there is either:

