11

Model Interpretability

Cantonese podcast title: 模型解釋性

Learning Objectives

  1. Define the Shapley value as the unique attribution satisfying efficiency, symmetry, null-player and additivity, and show that those four axioms jointly pin down a single formula.
  2. Compute partial dependence and individual conditional expectation for a fitted model, and explain why PDP averages out feature dependence while ICE preserves it.
  3. Implement a local surrogate (LIME), state the assumption whose violation makes its explanation inconsistent with the model's true behaviour, and contrast it with SHAP.
  4. Choose between SHAP, PDP/ICE, LIME and surrogate models by the question being asked, and name the failure mode of each.
  5. Construct a counterfactual explanation for a rejected loan and audit it for proximity, plausibility and the model's monotonicity assumptions.
Model Interpretability — visual guide
SHAP waterfall for one prediction SHAP waterfall: f(x) - E[f(X)] decomposed per feature efficiency axiom guarantees the bars sum to the deviation from baseline f(x) E[f(X)] baseline 0.18 +0.12 income +0.19 credit_score -0.09 debt_ratio -0.13 delinquencies +0.16 employment +0.13 utilisation +0.21 other f(x) = 0.77 Sum of phi_j equals f(x) - E[f(X)] by the efficiency axiom; red bars subtract, green add. TreeSHAP is exact for trees; KernelSHAP / DeepSHAP approximate under stated assumptions.

Assumes you know from ML-101

This lesson builds on ML-101 Lesson 9 (Model Evaluation), ML-101 Lesson 12 (Ensemble Learning), and ML-101 Lesson 14 (Neural Networks & Backprop). The reader is assumed to know that a fitted model ff maps inputs x∈Rpx \in \mathbb{R}^p to predictions y^\hat{y}, that ensembles average or vote over a collection of weak learners, and that a neural network is a composition of linear maps and element-wise nonlinearities trained by gradient descent. The reader is also presumed to know that models are evaluated by a held-out score, not by inspecting the coefficients.

What ML-101 did not do is ask how to attribute a prediction to its inputs in a principled way, what the partial dependence of y^\hat{y} on a single feature means once the other features are not independent of it, or why local surrogate models explain a single prediction but not the model. This lesson derives Shapley attributions from the four cooperative-game-theory axioms, derives partial dependence and ICE, and explains why LIME is local but not Shapley-faithful.

Learning Objectives

  1. Define the Shapley value as the unique attribution satisfying efficiency, symmetry, null-player and additivity, and show that those four axioms jointly pin down a single formula.
  2. Compute partial dependence and individual conditional expectation for a fitted model, and explain why PDP averages out feature dependence while ICE preserves it.
  3. Implement a local surrogate (LIME), state the assumption whose violation makes its explanation inconsistent with the model's true behaviour, and contrast it with SHAP.
  4. Choose between SHAP, PDP/ICE, LIME and surrogate models by the question being asked, and name the failure mode of each.
  5. Construct a counterfactual explanation for a rejected loan and audit it for proximity, plausibility and the model's monotonicity assumptions.

The Shapley value, derived from axioms

The setup is a cooperative game with pp players and a value function v:2{1,…,p}→Rv : 2^{\{1,\dots,p\}} \to \mathbb{R}. The Shapley value assigns to each player jj a payoff ϕj\phi_j that represents their fair contribution to the total v({1,…,p})v(\{1,\dots,p\}).

The four axioms that pin down ϕj\phi_j uniquely are:

  • Efficiency: ∑j=1pϕj=v({1,…,p})\sum_{j=1}^{p} \phi_j = v(\{1,\dots,p\}) — the attributions sum to the total.
  • Symmetry: if v(S∪{j})=v(S∪{k})v(S \cup \{j\}) = v(S \cup \{k\}) for every SS, then ϕj=ϕk\phi_j = \phi_k.
  • Null player: if v(S∪{j})=v(S)v(S \cup \{j\}) = v(S) for every SS, then ϕj=0\phi_j = 0.
  • Additivity: for any decomposition v=va+vbv = v_a + v_b, the attributions add: ϕj(v)=ϕj(va)+ϕj(vb)\phi_j(v) = \phi_j(v_a) + \phi_j(v_b).

The Shapley value is the unique function mapping vv to a vector (ϕ1,…,ϕp)(\phi_1,\dots,\phi_p) that satisfies all four. The proof is constructive: the weighted average

ϕj  =  ∑S⊆{1,…,p}∖{j} ∣S∣! (p−∣S∣−1)!p! [v(S∪{j})−v(S)]\phi_j \;=\; \sum_{S \subseteq \{1,\dots,p\}\setminus\{j\}}\,\frac{\lvert S \rvert !\,(p - \lvert S \rvert - 1)!}{p!}\,\bigl[v(S \cup \{j\}) - v(S)\bigr]

is the only one. The weight ∣S∣!(p−∣S∣−1)!/p!\lvert S \rvert ! (p - \lvert S \rvert - 1)!/p! is the fraction of orderings of the pp players that place jj immediately after some subset SS — the marginal contribution v(S∪{j})−v(S)v(S \cup \{j\}) - v(S) is averaged uniformly over all positions in the ordering.

For an ML model with prediction function ff evaluated at xx, the value function is

v(S)  =  E[ f(X)  ∣  XS=xS ],v(S) \;=\; \mathbb{E}\bigl[\,f(X) \;\big|\; X_S = x_S \,\bigr],

the expected model output when the features in SS are pinned to their values at xx and the rest are integrated over a background distribution. This is the move that takes the cooperative-game definition and makes it a tool for ML: v(∅)=E[f(X)]v(\emptyset) = \mathbb{E}[f(X)] is the baseline prediction, v({1,…,p})=f(x)v(\{1,\dots,p\}) = f(x) is the actual prediction, and the efficiency axiom guarantees

∑j=1pϕj  =  f(x)−E[f(X)].\sum_{j=1}^{p}\phi_j \;=\; f(x) - \mathbb{E}[f(X)].

The attributions sum to the deviation from the baseline — a model-faithful, locally additive decomposition of that prediction, not of the model in general.

import numpy as np
from itertools import chain, combinations

def powerset(iterable):
    s = list(iterable)
    return chain.from_iterable(combinations(s, r) for r in range(len(s) + 1))

def exact_shapley(values: dict[frozenset, float], n: int) -> np.ndarray:
    """
    Exact Shapley values by enumerating all 2^n subsets.

    `values` is a dict mapping frozenset of player indices to v(S). Exhaustive
    enumeration is only feasible for n <= ~14; for larger n use KernelSHAP
    or TreeSHAP.
    """
    phi = np.zeros(n)
    for j in range(n):
        others = [i for i in range(n) if i != j]
        for s in powerset(others):
            S = frozenset(s)
            S_with_j = S | {j}
            w = np.math.factorial(len(s)) * np.math.factorial(n - len(s) - 1) \
                / np.math.factorial(n)
            phi[j] += w * (values[S_with_j] - values[S])
    return phi

The code is a literal transcription of the Shapley formula. For n≤14n \le 14 the 2n2^n subsets fit in memory; beyond that, the exponential is prohibitive and an estimator is required. The exact form is the ground truth against which estimators are measured — knowing it exists and what it looks like is what lets a practitioner read a SHAP output as a number that obeys the four axioms, not as an opaque feature score.

SHAP estimators for tree, kernel and deep models

Three estimators are common and they make different trade-offs.

KernelSHAP approximates the Shapley values by regression. It samples coalitions S⊆{1,…,p}S \subseteq \{1,\dots,p\}, evaluates v(S)v(S) by marginalising out the absent features, and fits a weighted linear surrogate whose coefficients are interpretable as Shapley values. The sampling distribution weights coalitions by the same Shapley kernel weights that appear in the exact formula. KernelSHAP is model-agnostic: it works for any fitted ff that can be queried. Its cost is O(n p M)O(n\,p\,M) where MM is the number of Monte Carlo samples, and it inherits the background-distribution assumption of the value function — sample the background from the training joint or the explanation drifts.

TreeSHAP computes exact Shapley values for tree ensembles in polynomial time by exploiting the additive structure of decision trees. For a tree with TT internal nodes, the exact value can be computed in O(TLD2)O(TLD^2) for LL leaves and DD tree depth by traversing the splits and accumulating the marginal contributions at each. Across an ensemble of BB trees, total cost is O(B TLD2)O(B\,TLD^2), which for a forest of 300 trees of depth 6 is fast enough to be interactive. The assumption is that the tree structure is the entire model; a stacking wrapper that combines a tree ensemble with a logistic regression on top is not covered, and TreeSHAP applied to the wrapper produces attributions for the trees only, not the wrapper.

DeepSHAP propagates Shapley values through a neural network by linearising each layer around the background activations. It is exact for purely additive layers (linear maps, element-wise nonlinearities with vanishing second-order effects) and approximate otherwise. The deeper the network, the more the linearisation drifts from the true Shapley value; the practical rule of thumb is that DeepSHAP is reliable up to a handful of layers and degrades after that. For vision transformers and large language models, the same estimator is used but the background sample size must be carefully chosen — too few samples and the attributions have high variance.

EstimatorModels coveredCostFaithful?Caveat
Exactany, p≤14p \le 14O(2p⋅c)O(2^p \cdot c)exactexponential in pp
KernelSHAPanyO(M⋅p⋅c)O(M \cdot p \cdot c)approximatebackground sample must be joint
TreeSHAPtree ensemblesO(B⋅T⋅L⋅D2)O(B \cdot T \cdot L \cdot D^2)exact for treesdoes not cover stack wrappers
DeepSHAPneural networksO(N⋅forward pass)O(N \cdot \text{forward pass})approximate, drifts with depthlinearisation assumption

Partial dependence and ICE plots

Partial dependence answers a different question: averaged over the data, how does the prediction change as a single feature is varied? For feature jj with value xjx_j, the partial dependence function is

PDj(xj)  =  EX∖j[ f(xj,X∖j) ]  =  1n∑i=1nf(xj,x∖j(i)).\mathrm{PD}_j(x_j) \;=\; \mathbb{E}_{X_{\setminus j}}\bigl[\,f(x_j, X_{\setminus j})\,\bigr] \;=\; \frac{1}{n}\sum_{i=1}^{n} f(x_j, x_{\setminus j}^{(i)}).

The expectation is empirical: fix xjx_j, replace the rest of the features with their values from each training row, and average the predictions.

Partial dependence has two failure modes that practitioners routinely miss. First, when features are correlated, f(xj,x∖j(i))f(x_j, x_{\setminus j}^{(i)}) may evaluate the model at points that are not realistic — combinations of xjx_j and x∖jx_{\setminus j} that do not occur in the data. The function is then the model's behaviour on synthetic inputs, not on the true distribution. Second, PDP averages out everything except xjx_j, which means it shows the marginal effect; if the effect of jj depends on XkX_k, the dependence is invisible.

Individual Conditional Expectation (ICE) plots preserve the per-row trajectory. For each row ii:

ICEj(i)(xj)  =  f(xj,x∖j(i)).\mathrm{ICE}_j^{(i)}(x_j) \;=\; f(x_j, x_{\setminus j}^{(i)}).

Plotting one line per row shows the heterogeneity that PDP averages away. When the ICE lines are parallel, the effect of xjx_j is additive; when they fan out, the effect interacts with the rest of the row. The PDP is the average of the ICE curves; the ICE plot is the unbundled version.

import numpy as np

def partial_dependence(model, X: np.ndarray, feature_idx: int,
                       grid: np.ndarray) -> np.ndarray:
    """
    Compute the partial dependence function for `feature_idx` over `grid`.

    For each value in `grid`, replace column `feature_idx` of every row in
    X with that value, score the model, and average.
    """
    X_eval = X.copy()
    out = np.empty_like(grid, dtype=float)
    for k, v in enumerate(grid):
        X_eval[:, feature_idx] = v
        out[k] = model.predict(X_eval).mean()
    return out

The script is a literal transcription of the PDP definition. Its assumption is that the model evaluates plausibly off the training distribution. When that assumption fails — high correlation between features, or a model with sharp nonlinearities far from the data — PDP plots are an interpolation over regions where the model has not been validated. The fix is to overlay the training marginal and refuse to plot PDP outside its support.

Local surrogates and LIME

LIME constructs an interpretable model gg (typically a sparse linear model) that approximates a complex model ff in a neighbourhood around a single instance xx. The surrogate is fit on perturbed samples drawn from a neighbourhood πx\pi_x (in the original feature space for tabular data, in a superpixel-based space for images), with sample weights wi=exp⁡(−d(x,xi)2/σ2)w_i = \exp(-d(x, x_i)^2/\sigma^2) that decay with distance from xx.

The local fidelity of gg to ff is measured by

L(f,g,πx)  =  ∑i wi (f(xi)−g(xi′))2,\mathcal{L}(f, g, \pi_x) \;=\; \sum_{i}\,w_i\,\bigl(f(x_i) - g(x_i')\bigr)^2,

where xi′x_i' is the interpretable representation of xix_i. LIME minimises L\mathcal{L} subject to a complexity constraint on gg — usually ∥βg∥0≤K\lVert \beta_g \rVert_0 \le K for a sparse linear model with KK nonzero coefficients.

The result is a per-instance explanation: feature jj pushed the prediction up by βg,j\beta_{g,j} standard deviations around xx. But the explanation is local: gg approximates ff in πx\pi_x, not globally. Two instances near each other in πx\pi_x get similar explanations, but two instances in different regions of the input space can get opposite-signed explanations of the same feature. This is the source of the LIME-is-unstable critique: small changes in the kernel width σ\sigma or in the neighbourhood sampling produce visibly different explanations for the same instance.

The 201-level judgement is that LIME is a story, not a measurement. The surrogate's coefficients have no axiom-level guarantee: they are not Shapley values, they do not satisfy efficiency, and they change with σ\sigma. LIME is appropriate when a quick, intuitive, local explanation is needed and the practitioner is explicit about what the explanation is — a local linear approximation, not a measurement of feature importance. SHAP is the appropriate tool when a faithful per-instance attribution is required.

import numpy as np
from sklearn.linear_model import Ridge

def lime_explanation(model, x: np.ndarray, X_bg: np.ndarray,
                     sigma: float = 1.0, n_samples: int = 500,
                     n_features: int = 5) -> tuple[np.ndarray, float]:
    """
    Fit a local linear surrogate around `x` and return its top coefficients.

    Returns (coefs sorted by absolute value, surrogate R^2 on the perturbed
    samples). The R^2 is the honest score: low R^2 means the explanation
    is unreliable.
    """
    rng = np.random.default_rng(0)
    perturbations = rng.normal(0, 1, size=(n_samples, x.shape[0]))
    X_pert = x + sigma * perturbations
    y_pert = model.predict(X_pert)
    weights = np.exp(-np.sum(perturbations ** 2, axis=1))
    surrogate = Ridge(alpha=1e-2).fit(X_pert, y_pert, sample_weight=weights)
    r2 = surrogate.score(X_pert, y_pert, sample_weight=weights)
    order = np.argsort(np.abs(surrogate.coef_))[::-1][:n_features]
    return surrogate.coef_[order], r2

The output includes a local R2R^2 for the surrogate — the honest score of how well the explanation approximates the model in πx\pi_x. A low R2R^2 means the linear surrogate does not capture the model's behaviour even in the neighbourhood, and the explanation should be discarded. LIME implementations that do not report this number are reporting an explanation whose faithfulness is unknown.

Counterfactual explanations

A counterfactual for a rejected instance xx is a nearby instance x′x' such that f(x′)f(x') produces the desired outcome. The simplest definition asks for the minimum-distance perturbation:

x′  =  arg⁡min⁡z d(z,x)subject tof(z)=ydesired.x' \;=\; \arg\min_{z}\,d(z, x) \quad \text{subject to} \quad f(z) = y_{\text{desired}}.

The distance dd is typically a weighted L1 or L2 norm, and the search is over a constrained set (features that can change, features that are immutable like age or ethnicity). Counterfactuals are actionable in a way Shapley values are not: they tell the user what to change, not just what mattered.

The construction has three practical properties to audit. Validity: does f(x′)f(x') actually produce the desired outcome? Proximity: is x′x' close enough to be useful? Plausibility: is x′x' in a region where the model has been trained, or is it an out-of-distribution point that flips the prediction by accident?

import numpy as np

def counterfactual_l2(model, x: np.ndarray, target: float,
                      step: float = 0.1, max_iter: int = 200) -> np.ndarray:
    """
    Greedy L2 counterfactual search.

    At each step, perturb x in the direction of the model's gradient to
    move the prediction toward `target`. The result is approximate; real
    systems use DiCE, ALICE or MIP-based methods.
    """
    x_cf = x.copy().astype(float)
    for _ in range(max_iter):
        y_now = model.predict(x_cf[None, :])[0]
        if abs(y_now - target) < 1e-3:
            break
        # Numerical gradient (in production: use model's grad if available)
        eps = 1e-4
        grad = np.zeros_like(x_cf)
        for j in range(len(x_cf)):
            x_plus = x_cf.copy(); x_plus[j] += eps
            x_minus = x_cf.copy(); x_minus[j] -= eps
            grad[j] = (model.predict(x_plus[None, :])[0]
                       - model.predict(x_minus[None, :])[0]) / (2 * eps)
        x_cf -= step * np.sign(y_now - target) * grad
    return x_cf

The greedy search is illustrative; production counterfactual libraries use mixed-integer programming (for categorical features with constraints), growing-spheres methods (for black-box models without gradients), or genetic search (for high-dimensional image inputs). The assumption under which any counterfactual is fair is that the model's predictions are smooth and well-behaved in the region between xx and x′x'. When the model has sharp discontinuities — a deep network with ReLU gates — a counterfactual can land on the other side of a boundary that is geometrically one feature away, and the explanation reads as "change feature 7 by 0.001" rather than a meaningful intervention.

Choosing a method, naming its failure mode

QuestionToolFailure mode
Why did the model predict yy for this row?SHAPbackground sample is not joint; attributions drift
What does the model say if I change only feature jj?PDP / ICEoff-distribution evaluation when features are correlated
Give me a quick, intuitive local storyLIMEsurrogate R2R^2 may be low; explanation is unstable in σ\sigma
What is the smallest change that flips the outcome?counterfactualmay land on out-of-distribution points that exploit model brittleness
Is the model globally linear in feature jj?partial dependence with a fitted smooth curvelinearity assumption fails on tree models
Which features are most important globally?mean ∣ϕj∣\lvert \phi_j \rvert over a sampledepends on the background; varies with the sample

The unifying judgement is that every method has a stated assumption and a stated failure mode. The 201-level practitioner picks the method whose assumption matches the dataset, audits the assumption before trusting the output, and names the failure mode explicitly when communicating the result. "Why did the model reject my loan?" is a different question from "What would change the decision?" and deserves a different tool — a SHAP attribution is not a counterfactual, and a counterfactual is not a SHAP attribution.

Integrated gradients and attribution in deep networks

Saliency maps — the gradient of the model's output with respect to the input — are the 101-level "interpretation" of a neural network. They are also wrong as a per-instance attribution: they can saturate (a ReLU gate whose pre-activation is far from zero has zero gradient regardless of how much the input pushed it) and they ignore the path through which the input had to travel to affect the output. Integrated gradients fix both problems by averaging the gradient along a straight-line path from a baseline x′x' to the input xx:

IGj(x)  =  (xj−xj′) ∫α=01 ∂f(x′+α(x−x′))∂xj dα.\mathrm{IG}_j(x) \;=\; (x_j - x'_j)\,\int_{\alpha=0}^{1}\,\frac{\partial f\bigl(x' + \alpha(x - x')\bigr)}{\partial x_j}\,d\alpha.

The integral is over α∈[0,1]\alpha \in [0, 1] parameterising the straight-line path; the prefactor xj−xj′x_j - x'_j converts the integrated gradient back to an attribution in the input scale. Integrated gradients satisfy two axioms that raw saliency does not: sensitivity (if an input differs from the baseline in one feature but the output does not change, the attribution for that feature is zero) and implementation invariance (two networks that are functionally identical — same function from inputs to outputs — have identical attributions for every input).

The baseline x′x' matters as much as the formula. The default in vision is a black image; in text, a sequence of padding tokens; in tabular data, the per-feature mean of the training set. Each choice produces a different attribution, because the integral runs from a different start. The 201-level practice is to report the attribution for at least two baselines (e.g. the mean and a zero vector) and to flag features whose sign flips across baselines — those are features whose attribution is baseline-dependent and therefore not a stable property of the model.

import numpy as np

def integrated_gradients(model, x: np.ndarray, baseline: np.ndarray,
                         steps: int = 50) -> np.ndarray:
    """
    Approximate integrated gradients by the Riemann sum over `steps`.

    `model.predict` is called once per step; in production this is replaced
    by a batched forward pass. The Riemann sum is a numerical approximation
    of the path integral.
    """
    diff = x - baseline
    avg_grad = np.zeros_like(x)
    for alpha in np.linspace(0.0, 1.0, steps):
        x_step = baseline + alpha * diff
        eps = 1e-4
        # Numerical gradient; replace by autograd in production
        for j in range(len(x)):
            x_plus = x_step.copy(); x_plus[j] += eps
            x_minus = x_step.copy(); x_minus[j] -= eps
            avg_grad[j] += (model.predict(x_plus[None])[0]
                             - model.predict(x_minus[None])[0]) / (2 * eps)
    return diff * avg_grad / steps

The Riemann approximation with KK steps has error O(∥x−x′∥∞/K)O(\|x - x'\|_\infty / K), so doubling KK halves the error. The deep SHAP estimator in the previous section is a stochastic approximation to the same integral over an empirical baseline distribution; when the baseline is a single point, the two coincide.

Anchor explanations and rule-based surrogates

LIME and SHAP produce a real number per feature; anchors produce an IF-THEN rule that is sufficient to lock the prediction. An anchor is a set of feature predicates A={x1∈S1,x2∈S2,… }A = \{x_1 \in S_1, x_2 \in S_2, \dots\} such that the model's prediction is the same for every perturbation of xx that satisfies AA, with high probability:

ED(x∣A)[ 1{f(z)=f(x)} ]  ≥  1−ϵ.\mathbb{E}_{\mathcal{D}(x \mid A)}\bigl[\,\mathbb{1}\{f(z) = f(x)\}\,\bigr] \;\ge\; 1 - \epsilon.

The construction is by a beam search over candidate predicates, with the precision estimated by sampling perturbations that respect AA. The output is interpretable in a way continuous attributions are not: "if age ∈[30,50)\in [30, 50) and income ∈[med,∞)\in [\mathrm{med}, \infty), the model predicts approve with precision 0.97 on perturbations".

The 201-level judgement is that anchors are appropriate when the consumer of the explanation is a person who needs a decision rule — a regulator, a clinician, a customer-service agent. They are not appropriate when the consumer is a data scientist who wants to compare two models' behaviour at a fine-grained level. The two tools answer different questions, and the lesson's deeper point is that the choice of explanation is a communication decision, not a measurement decision.

Worked example: a regression model, three attribution methods

The following walks a single prediction through SHAP, LIME and integrated gradients and shows where they agree and where they disagree.

import numpy as np
from sklearn.linear_model import Ridge
import shap

rng = np.random.default_rng(0)
n, p = 500, 8
X = rng.normal(size=(n, p))
beta = np.array([3.0, -2.0, 0.0, 0.0, 1.5, 0.0, 0.0, 0.5])
y = X @ beta + 0.1 * rng.normal(size=n)

model = Ridge(alpha=0.1).fit(X, y)

# SHAP on a linear model is exact and matches the coefficient * (x - mean)
def use_explainer():
    return shap.LinearExplainer(model, shap.maskers.Independent(X))

# Pick an interesting row
x = X[0]
shap_values = use_explainer().shap_values(x)
print("SHAP       :", shap_values.round(3))
print("true beta  :", beta.round(3))
print("x - mean(x):", (x - X.mean(axis=0)).round(3))
print("expected   :", (beta * (x - X.mean(axis=0))).round(3))

On a linear model the SHAP attribution is exactly β⊙(x−E[X])\beta \odot (x - \mathbb{E}[X]): the contribution of feature jj to the deviation of the prediction from the mean. For a non-linear model the same identity does not hold and the SHAP attribution differs from the linear coefficient; that difference is the interaction effect the model has learned. The 201-level exercise is to plot the SHAP attribution against βj⋅(xj−xˉj)\beta_j \cdot (x_j - \bar{x}_j) for many rows and look at the residuals — features with large residuals are where the model is doing non-linear work on top of the linear signal, and the explanation has to account for both.

Global summary: mean absolute SHAP, signed SHAP, and the interaction decomposition

The default global summary of SHAP is

ϕˉj  =  1n ∑i=1n ∣ϕj(xi)∣,\bar{\phi}_j \;=\; \frac{1}{n}\,\sum_{i=1}^{n}\,\lvert \phi_j(x_i) \rvert,

the mean absolute attribution per feature across the dataset. This is a defensible scalar summary under the Shapley axioms but loses the sign: it cannot distinguish a feature that pushes every prediction up from one that pushes every prediction down by the same amount. The signed summary 1n∑iϕj(xi)\frac{1}{n}\sum_i \phi_j(x_i) is informative for understanding the average direction but is silent about heterogeneity. The 201-level practice is to report both, plus a histogram of ϕj\phi_j over the dataset — features with bimodal histograms are pushing some predictions up and others down, which the scalar summaries cannot surface.

The interaction index is the natural generalisation when the SHAP library's interaction_values API is available. For features jj and kk, the interaction ϕj,k(x)\phi_{j, k}(x) satisfies

ϕj(x)  =  ϕjmain(x)  +  ∑k≠j ϕj,k(x),\phi_j(x) \;=\; \phi_j^{\mathrm{main}}(x) \;+\; \sum_{k \neq j}\,\phi_{j, k}(x),

a decomposition of the per-feature attribution into a main effect plus pairwise interactions. Reporting only ϕj\phi_j and not ϕj,k\phi_{j,k} is the LIME-style mistake: the explanation looks complete but the interaction is rolled into the main effect and the consumer cannot tell which.

Key Takeaways

  • Shapley values are the unique attribution that satisfies efficiency, symmetry, null-player and additivity; the axioms pin down the formula, not the other way around.
  • SHAP estimators trade off exactness against model coverage: TreeSHAP is exact for trees, KernelSHAP is model-agnostic but approximate, DeepSHAP drifts with network depth.
  • Partial dependence averages out feature dependence and is misleading on correlated inputs; ICE preserves the per-row trajectory and surfaces interactions.
  • LIME is a local story, not a measurement; its surrogate R2R^2 is the honest score, and low R2R^2 means the explanation should be discarded.
  • Counterfactuals are actionable but require auditing for plausibility; a counterfactual that lands off the training distribution is exploiting model brittleness, not exposing truth.
  • Integrated gradients fix the saturation and path problems of raw saliency; anchors deliver rule-based surrogates for human consumers; mean absolute SHAP and the interaction index together describe global behaviour without flattening it into one number.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

The Shapley value is the *unique* function from value functions to per-player payoffs that satisfies which four?

0 of 8 answered

Pick a lesson to start the audio.