Explainability Principles

Quantitative Responsible AI: Principles, Governance, and Methods

Author

Fei Huang, UNSW Sydney

Learning objectives

  • Define explainability, and recognise interpretability and transparency as closely related terms often used interchangeably.
  • Explain why explainability matters in high-stakes domains such as lending, hiring, healthcare, and insurance.
  • Map stakeholder needs to different types of explanation.
  • Classify explainability methods along the key taxonomic dimensions.

What is explainability?

Between 2013 and 2019, the Dutch tax authority used a self-learning algorithm to flag childcare-benefit applications as high-risk for fraud. The system weighted dual nationality as a risk factor, a proxy for ethnicity, and wrongly flagged around 26,000 families, disproportionately those with a migration background, who were then forced to repay tens of thousands of euros they did not owe. Affected families could not get a meaningful account of why they had been flagged, the oversight bodies meant to review individual cases could not properly interrogate the algorithm’s reasoning either, and the resulting financial ruin left over a thousand children removed from their homes. The scandal, known in the Netherlands as the toeslagenaffaire, brought down the Dutch government in January 2021 (Amnesty International 2021).

No single official decided to do this. Somewhere between “an algorithm flagged this family” and “children were taken into care,” there was no point at which a family, a caseworker, or a court could get a clear, checkable account of why. That gap, between a model producing an output and a human being able to understand and contest it, is what this chapter is about.

Core definitions

Explainability converts a model’s internal logic into a form a specific audience can understand, for a specific purpose. A model is explainable to the degree that a person can understand the cause of its decision (Biran and Cotton 2017), or, put operationally, to the degree that a user can correctly and efficiently predict what it will do (Kim et al. 2016).

Note

Two closely related terms appear throughout the literature, and this course will not police the boundary between them sharply. Interpretability is often used for the same underlying idea. Transparency sometimes refers more specifically to understanding a model’s internal structure directly, as opposed to a post-hoc explanation constructed after training without exposing what the model actually computes internally (Lipton 2018). Rudin (2019) argues that the terminological churn between these words has itself strayed from the needs of real problems. What matters in each case is: for whom, about what, and for what purpose?

Why does explainability matter?

A model can be highly accurate and still be a liability if nobody can explain what it is doing. Several distinct reasons drive the case for explainability, and they don’t all point in the same direction (Doshi-Velez and Kim 2017; Molnar 2025):

Reason Description
Human learning People need explanations to update mental models when outcomes are unexpected
Safety and testing High-risk applications require confidence that models behave correctly
Bias detection Explainability is a debugging tool for finding discriminatory patterns
Social acceptance Transparent systems gain greater user trust and legitimacy
Auditing Models can only be properly evaluated when their decisions are visible
Right to explanation Regulatory requirements (GDPR, ASIC, EU AI Act) mandate explanations for automated decisions

When is explainability not required?

Explainability is not free. It can cost accuracy, engineering time, and sometimes security. The case for it weakens when (Doshi-Velez and Kim 2017; Molnar 2025):

  • The application has low stakes and no significant consequences from errors
  • The method is well-established and thoroughly tested (e.g. standard actuarial tables, validated clinical scoring rules)
  • Transparency enables gaming or manipulation (e.g. fraud scoring systems)

But most consequential domains this course covers (lending, hiring, healthcare, criminal justice, insurance) involve high-stakes decisions and regulatory scrutiny by default. Explainability is generally expected.

Demands for explanation across domains

Regulators, decision subjects, and internal reviewers each add specific demands for explainability:

  • Regulators expect models to be documented, justified, and auditable
  • Decision subjects have a right to understand why they were declined, screened out, or charged a particular price
  • Practitioners signing off on a model’s use need to verify that outputs behave as expected
  • Boards and senior management need to understand model risk

A black-box XGBoost model (a highly accurate model built by combining many decision trees, a technique called gradient boosting) may outperform a simpler model on predictive accuracy, yet be harder to document, harder to explain to an affected individual, and harder to audit for fairness.

NoteCase Study: insurance — regulatory demands

Actuarial and insurance regulatory settings illustrate this concretely. Prudential and market-conduct regulators (e.g. APRA, ASIC, US state insurance departments) expect pricing and underwriting models to be filed and justified. Policyholders (the customers who hold an insurance contract) have a right to understand why they are charged a particular premium (the price paid for insurance cover) or denied a claim (a request for the insurer to pay out after a loss). Actuaries signing off on a model must verify its outputs behave as expected.

Goals of explainability

Three goals of explainability

Explainability serves several distinct purposes, not just one (Rudin et al. 2022):

1. Improving the model

Explainability helps identify when models learn unintended patterns. A model can have strong test-set accuracy while relying on non-causal features, for example using background snow in a photo to classify wolves vs dogs. In insurance, this might mean using a proxy variable (a feature that indirectly stands in for another) for a protected attribute (a legally protected personal characteristic, such as race or gender, that a model should not rely on) without detection.

2. Justifying predictions

Different stakeholders require different explanations. Decision-makers need to verify forecasts against domain knowledge. Customers need to understand and contest decisions. Regulators need to verify compliance.

3. Discovering insights

Beyond predictions, explainability reveals relationships between inputs and outputs, useful for product design, pricing strategy, and scientific understanding.

Characteristics of good explanations

Research on how people actually process explanations, not just what is technically correct, identifies what makes them effective (Miller 2019):

Principle Implication
Contrastive “Why this prediction rather than another?” is more useful than a full causal chain
Selective 1–2 key reasons are accepted even when many factors contribute
Social Explanations vary by audience — a customer needs a different explanation than an actuary
Abnormality focus Rare or unexpected features make stronger explanations than common ones
Consistent with beliefs Explanations aligning with prior knowledge are more persuasive (but watch for confirmation bias)
Faithful Explanations should truthfully reflect model behaviour, not just sound plausible

Stakeholders

Who needs explanations?

A useful way to organise stakeholder needs is by the role someone plays in relation to a model, since the same person can occupy different roles for different systems (Tomsett et al. 2018). Six roles recur across domains:

Stakeholder Need
Model creators (data scientists) Debugging, feature engineering, model selection
Operators (e.g. insurers, banks, hospitals, employers) Compliance, governance, sign-off
Executors (e.g. underwriters, loan officers, clinicians, hiring managers) Verify model outputs before acting on them
Decision subjects (e.g. policyholders, applicants, patients) Understand why an outcome or decision differs
Auditors (regulators, internal audit) Validate fairness, accuracy, compliance
Data subjects Challenge automated decisions (GDPR right to explanation)

Different stakeholders need different types of explanation, such as local versus global, or technical versus plain language.

Mapping stakeholders to explanations

Concretely, the mapping of stakeholder to explanation type looks different in every domain, but the structure is the same. Identify the question being asked, then match scope (global vs local) and audience (technical vs plain language).

NoteCase Study: insurance — stakeholder mapping
Stakeholder Typical question Explanation type needed
Actuary signing off Does the model behave sensibly across the rating factors (the input variables used to set a price, e.g. age or location)? Global, feature effects
Underwriter Why is this applicant’s premium higher than expected? Local, individual prediction
Customer Why did my premium increase this year? Local, contrastive, plain language
Regulator / auditor Does this model treat protected groups fairly? Global, group-level, with statistical tests
Board / management What are the top drivers of claims costs this year? Global, feature importance

The same mapping recurs elsewhere. A loan officer asking why an applicant was declined needs a local explanation, and a regulator auditing a hiring-screen tool for disparate impact (when a seemingly neutral policy ends up harming one group more than another) needs a global, group-level one.

Taxonomy of methods

Two primary categories

Explainability by design: train models that are inherently interpretable.

  • Linear and logistic regression
  • Generalised linear models (GLMs, a model that predicts an outcome as a weighted sum of the input features) and generalised additive models (GAMs, a more flexible version that allows curved rather than straight-line effects)
  • Decision trees (models that split data by a series of branching yes/no questions) and decision rules
  • RuleFit (a model combining decision rules with a linear model) (Friedman and Popescu 2008)

Post-hoc explainability: apply methods after model training to explain an existing (possibly black-box) model.

The choice between them is a fundamental design decision made before modelling, not an afterthought.

Dimension 1: Intrinsic vs post-hoc

Intrinsic (by design) Post-hoc
How Interpretable model structure Explanation applied after training
Examples GLM, decision tree, GAM SHAP, PDP, LIME
Advantage Explanation is exact Works on any model
Limitation May sacrifice predictive power Approximation; may not be faithful

In regulated scoring settings (insurance pricing, credit scoring), GLMs are the standard because they are intrinsically interpretable and easy to document. Machine learning models require post-hoc explanation.

Dimension 2: Model-agnostic vs model-specific

Model-agnostic: works with any model. Treats the model as a black box and queries it for predictions.

  • Uses the SIPA principle: Sample data → Perform an intervention → get Predictions → Aggregate (Molnar 2025)
  • Examples: SHAP (KernelSHAP), LIME, PDP, permutation feature importance

Model-specific: designed for a particular model class.

  • Examples: GLM coefficients, decision tree paths, TreeSHAP (for trees), saliency maps (for neural networks)
  • Faster and often more accurate, but tied to the model architecture

Dimension 3: Global vs local

Global: describes average model behaviour across the dataset.

  • “What features does this model rely on overall?”
  • Examples: permutation feature importance, PDP, ALE, surrogate models (simple models trained to approximate a complex one)

Local: explains a single prediction for one individual.

  • “Why did this particular applicant receive this premium?”
  • Examples: SHAP values for one observation, LIME, counterfactual explanations (statements like “if this input had been different, the outcome would have changed”)

Both are needed in practice. Global methods support auditing and governance. Local methods support customer explanations and contestability.

Summary taxonomy

The three dimensions are independent axes, not a nested hierarchy. A method’s position on one does not determine its position on the others.

Dimension Values Examples
Origin Intrinsic vs post-hoc GLM coefficients, decision-tree paths = intrinsic; KernelSHAP, TreeSHAP, PDP, LIME = post-hoc
Applicability Model-agnostic vs model-specific KernelSHAP, PDP, PFI, LIME = model-agnostic; GLM coefficients, decision-tree paths, TreeSHAP = model-specific
Scope Global vs local Mean absolute SHAP values (aggregated), PFI, PDP = global; SHAP value for one observation, LIME = local

Any combination is possible, e.g.:

  • GLM coefficients and decision-tree paths: intrinsic + model-specific
  • KernelSHAP: post-hoc + model-agnostic
  • TreeSHAP: post-hoc + model-specific
  • A single SHAP value: local
  • The mean absolute SHAP value for a feature across all observations: global

Explainable models

Generalised linear models

Generalised linear models are a standard tool in regulated scoring settings (credit scoring, clinical risk, insurance pricing) precisely because they are intrinsically interpretable.

A Poisson GLM, e.g. for a count outcome such as claim frequency (how often a policyholder files a claim) or hospital admissions:

\log \mathbb{E}[Y_i] = \beta_0 + \beta_1 x_{i1} + \cdots + \beta_p x_{ip} + \log(\text{exposure}_i)

Each coefficient \beta_j has a direct interpretation. A one-unit increase in x_j multiplies expected frequency by e^{\beta_j}, holding all else equal.

Limitations: assumes a specific functional form and cannot capture complex non-linear interactions automatically.

Decision trees

A decision tree partitions the feature space into rectangular regions and assigns a constant prediction to each.

Strengths: highly interpretable, can be visualised as a flow chart.

Limitations: unstable (small data changes cause large tree changes), tend to overfit, and poor at smooth non-linear relationships.

Decision trees are rarely used directly for pricing or scoring in regulated settings, but are useful for communicating segmentation logic to non-technical stakeholders, for example in insurance, lending, and clinical triage.

GAMs: extending GLMs

A Generalised Additive Model (GAM) (Hastie and Tibshirani 1990) replaces linear terms with smooth functions:

g(\mathbb{E}[Y]) = \beta_0 + f_1(x_1) + f_2(x_2) + \cdots + f_p(x_p)

Each f_j is a smooth, visualisable function, often a spline (a flexible curve built from smooth pieces). GAMs are more flexible than GLMs while remaining interpretable through plots of each f_j.

Neural additive models (NAMs) (Agarwal et al. 2021) extend the idea to deep learning with interpretable additive structure.

Limits of explainability

The accuracy–explainability trade-off

More flexible models (gradient boosting, neural networks) tend to be more accurate but less explainable. Explainable models (GLMs, decision trees) are transparent but may underfit complex patterns.

This trade-off is widely assumed, but Rudin (2019) argues it is often a myth, particularly on structured data with meaningful, well-engineered features: once such features exist, complex black-box models frequently show little or no accuracy advantage over a well-constructed interpretable one. The trade-off is more likely to be real where useful features are not hand-built, such as raw images or audio, but it should be checked empirically for a given problem rather than assumed. Where a genuine trade-off does exist, methods like GAMs and constrained gradient boosting can narrow it. Post-hoc explanation of a black-box model is always an approximation, and the explanation may not be fully faithful to the model’s actual reasoning.

Human factors in explanation

Technical accuracy is not sufficient on its own. An explanation only does its job if a person can actually use it (Miller 2019):

  • Cognitive load: simpler explanations are preferred even if incomplete
  • Selectivity: 1–2 reasons are easier to act on than 20 feature contributions
  • Trust calibration: overly confident explanations can cause over-reliance on a flawed model

In professional settings, the risk of a plausible but misleading explanation is at least as serious as no explanation at all.

Summary

  • Explainability is not an end in itself. It is a means to accountability, trust, and governance.
  • The choice of method depends on who needs the explanation, what they need to understand, and what decision follows.
  • The fundamental taxonomy: intrinsic vs post-hoc, model-agnostic vs model-specific, global vs local.
  • In regulated scoring domains (insurance, credit, clinical risk), GLMs provide explainability by design. Machine learning models require careful post-hoc explanation.

Post-hoc explainability methods

Overview

This section covers the four most widely used post-hoc, model-agnostic methods in practice:

Method Scope What it answers
Permutation Feature Importance (PFI) Global Which features matter most to this model?
Partial Dependence Plots (PDP) Global How does a feature affect predictions on average?
SHAP Local + Global How much did each feature contribute to this prediction?
LIME Local What simple rule approximates the model near this prediction?

All methods treat the model as a black box and work for any model type (GLM, XGBoost, neural network).

Permutation Feature Importance

Break a feature, watch the damage

Breiman (2001) introduced this idea as a variable-importance measure for random forests, but the underlying logic transfers to any fitted model, not just random forests: if a model genuinely relies on a feature, scrambling that feature’s values should hurt its predictions. If shuffling a column barely changes performance, the model was not using it.

The appeal is that this works without touching the model itself: no retraining, no access to internals, just repeated scoring on deliberately corrupted data.

Computing it

  1. Fit the model and compute baseline error e_0 on test data.
  2. For each feature j = 1, \ldots, p:
    • Permute (shuffle) column j in the test set.
    • Re-score the model on the shuffled data, with no retraining.
    • Compute error e_j on the permuted data.
    • Importance: I_j = e_j / e_0 (ratio) or e_j - e_0 (difference).
  3. Rank features by I_j.

What it tells you, and where it breaks down

PFI is model-agnostic, captures interaction effects (shuffling one feature disrupts every interaction it participates in), and compresses an entire model into a single ranked list, all without retraining. But it is a coarse instrument. Permuting a feature that is correlated with others creates combinations that never occur in real data, so an importance score can reflect an artefact of the shuffle rather than genuine model behaviour. It also only shows that a feature matters, not how, in which direction, or over what range, and a single run is noisy enough that results should be averaged over several permutations. If the model itself does not generalise well, PFI faithfully reports the model’s reliance on training noise rather than real signal.

Reading a PFI ranking

A PFI ranking surfaces two kinds of signal at once. It shows which legitimate factors the model actually relies on, and which non-legitimate or proxy features contribute more than domain knowledge would suggest.

NoteCase Study: insurance — PFI ranking

In a motor insurance pricing model, PFI might reveal:

  • Age and Bonus (a driver’s no-claims discount level) are the most important legitimate rating factors
  • InsuranceScore (a non-legitimate proxy) contributes significantly, signalling that fairness scrutiny is needed
  • Density and Value contribute less than expected from domain knowledge, prompting a check of data quality

This is often the first diagnostic run after fitting a new model, in any domain. A hiring model whose top feature turns out to be a proxy for age or gender warrants the same scrutiny.

Partial Dependence Plots

Averaging out everything except one feature

Friedman (2001) introduced PDP alongside gradient boosting, as a way to ask a question that a coefficient cannot answer for a black-box model: if every observation in the data shared the same value of one feature, what would the average prediction look like? The construction is direct. Fix the feature of interest to a value, force every observation to take it, score the model, and average the result. Repeat across the feature’s range and the resulting curve is the partial dependence function.

Formula

For feature x_j, the partial dependence function is:

\hat{f}_j(x_j) = \mathbb{E}_{X_{-j}}\left[\hat{f}(x_j, X_{-j})\right] \approx \frac{1}{n}\sum_{i=1}^n \hat{f}(x_j, x_{-j}^{(i)})

We substitute x_j for all observations and average, marginalising over the joint distribution of the remaining features.

What the curve tells you

A PDP is intuitive to read: direction (does the prediction rise or fall as the feature increases?), shape (linear, monotone, or non-linear, such as U-shaped with age), and magnitude (how much the prediction moves across the observed range) are all visible at a glance, for any model type, without assuming a functional form. A causal reading of that shape is not automatic. It requires additional causal assumptions on top of the plot itself (Zhao and Hastie 2021).

For two features, a 2D PDP surface can reveal possible interaction, but the joint surface also contains both features’ main effects, so confirming genuine interaction requires comparing it against the sum of the two one-feature PDPs, or using a dedicated interaction measure (H-statistic, SHAP interaction values, Ch5). Visualisation also becomes difficult beyond two variables.

Where the average misleads you

Two assumptions do most of the damage. PDP assumes feature independence: when features are correlated, it can extrapolate to combinations that are rarely or never observed in the data, and that extrapolation region is exactly where the plot is least trustworthy. And because PDP is an average, it can mask heterogeneous effects: if the effect differs across subgroups, or points in opposite directions for different individuals, the average can look flat or misleading even though no individual actually experiences that average effect. On top of this, PDP is only practical for one or two features at a time, and it does not show the underlying feature distribution, which invites over-interpreting effects in sparse regions of the plot.

Note

Individual Conditional Expectation (ICE) curves (Goldstein et al. 2015) address the heterogeneous-effects limitation. A PDP is the average of many ICE curves, one per observation, each showing how a single instance’s prediction would change as the feature of interest varies, holding everything else about that instance fixed. Overlaying ICE curves on top of a PDP reveals whether the average effect is representative of most observations, or whether it masks subgroup effects that point in different directions and cancel out in the average. ICE curves still assume feature independence, so they remain vulnerable to the same extrapolation problem as PDP when features are correlated.

Note

Accumulated Local Effects (ALE) plots (Apley and Zhu 2020) address the independence assumption by conditioning on local neighbourhoods rather than marginalising globally. Prefer ALE when features are correlated.

NoteCase Study: manipulating PD plots to hide discrimination

The two weaknesses above are not just textbook caveats. Xin et al. (2025) show they can be deliberately exploited. Their adversarial framework leaves a black-box model’s predictions untouched on the real, densely-observed region of the data, but substitutes a separately chosen, non-discriminatory output whenever an input falls in the sparse “extrapolation” region that PDP’s averaging step actually depends on. Because that region is rarely or never observed in practice, almost none of the model’s real-world predictions change, yet the resulting PD plot for the manipulated feature can be engineered to look essentially flat.

Applied to real motor insurance claims data and to the COMPAS recidivism dataset, the attack produced PD plots for a protected or proxy variable (age, race) that showed no visible disparity, while the underlying model’s predictions, and its discriminatory behaviour, were preserved almost entirely, with over 95% accuracy retained relative to the original model’s outputs. This directly undermines the “the protected-attribute PD plot looks flat, so the model is fine” check used elsewhere in this course. ICE curves, correlation checks between features, and ALE plots (above) make the manipulation harder to construct convincingly, but none of them make it impossible. Treat a flat PD plot as one input to a fairness review, not as proof of one.

Reading a PDP in practice

A PDP for a numeric feature such as age often shows a non-linear, non-monotonic effect that a GLM coefficient cannot represent directly, but the PDP makes it visible for any model type.

NoteCase Study: insurance — reading a PDP

A PDP for Age in a motor insurance frequency model might show high predicted frequency for young drivers (18–25), declining steeply to ~30, then gradually rising after 65.

The same non-monotonic shape shows up elsewhere. Years of prior work experience rarely pushes a hiring model’s output in one straight line, and tumour size rarely moves a clinical risk model’s prediction linearly either. A PDP makes that curvature visible without having to guess its functional form in advance.

SHAP

Splitting credit for a prediction

SHAP (SHapley Additive exPlanations) (Lundberg and Lee 2017) borrows its answer from cooperative game theory: if several players jointly produce a payoff, how should it be split among them, based on each one’s actual contribution? Treat a model’s features as the players and its prediction as the shared payoff, and the fair split is the Shapley value, the average marginal contribution of feature j across every possible combination of the other features it could appear alongside.

Explanation model:

g(z') = \phi_0 + \sum_{j=1}^p \phi_j z_j'

where \phi_j is the Shapley value for feature j and \phi_0 is the baseline (average prediction).

Why Shapley values, specifically

Game theory guarantees this particular split satisfies three properties no other feature-attribution method achieves simultaneously (Lundberg and Lee 2017): local accuracy (the contributions sum exactly to the gap between this prediction and the baseline, \sum_j \phi_j = \hat{f}(x) - \mathbb{E}[\hat{f}]), missingness (a feature absent from the model gets zero credit), and consistency (if a feature’s true marginal contribution goes up after a model change, its Shapley value cannot go down).

Two ways to compute it

KernelSHAP is model-agnostic: it samples feature coalitions, scores them, and fits a weighted linear regression with the SHAP kernel to recover the \phi_j. It works for any model but is slow.

TreeSHAP (Lundberg et al. 2020) exploits tree structure to compute exact values in polynomial rather than exponential time, for XGBoost, LightGBM, and random forests. It is the default choice for gradient boosting models across domains, including insurance pricing, credit scoring, and clinical risk.

Reading a SHAP output

A waterfall plot decomposes one prediction into the baseline plus each feature’s push or pull: start at the baseline E[\hat{f}] (the average prediction across the population), let each feature push the prediction up (positive SHAP value) or down (negative value), and the sum of every push and pull lands exactly on this individual’s prediction.

NoteCase Study: insurance — SHAP waterfall

For a single policyholder’s motor premium prediction, age (young) pushes the prediction up and bonus (low, i.e. a strong no-claims record) pushes it down. This is the natural format for a customer-facing explanation: “Your premium is higher than average mainly because of your age and driving history.”

The same waterfall format works for a declined loan application or a clinical risk score, expressed as a sum of contributions from income, credit history, or lab values relative to a baseline. Beyond the waterfall, a summary plot shows importance, direction, and feature value across every observation at once; a dependence plot shows SHAP value against feature value for one feature, the local counterpart to a PDP; the mean absolute SHAP value aggregates local values into a global importance ranking; and interaction values show how pairs of features jointly move a prediction.

What SHAP buys you, and what it does not

The game-theoretic guarantees are real strengths: local accuracy, missingness, and consistency hold exactly, local values aggregate cleanly into a global summary, and TreeSHAP makes all of this fast for tree-based models. But the guarantees have limits. KernelSHAP is slow without a tree-specific shortcut, and it assumes feature independence, an assumption TreeSHAP’s conditional variant relaxes at the cost of sometimes assigning non-zero credit to a correlated feature the model never actually used. Shapley values are attributions, not causes, and treating a large \phi_j as evidence that a feature causes the outcome is a category error. And, like any post-hoc explanation, SHAP explanations can be adversarially manipulated to hide biases in the underlying model (Slack et al. 2020).

LIME

Zoom in until the model looks simple

LIME (Local Interpretable Model-Agnostic Explanations) (Ribeiro et al. 2016) starts from a different bet than SHAP’s game-theoretic split: a model that is hopelessly complex globally may behave almost linearly in a small enough neighbourhood around one point. Fit a simple, interpretable model in that neighbourhood, and its coefficients become the explanation for the complex model’s prediction there.

Building the local model

  1. Select the instance x^* to explain.
  2. Perturb the data: generate samples near x^* by sampling from a normal distribution (tabular), removing words (text), or toggling superpixels (images).
  3. Score the black-box model on perturbed samples.
  4. Weight samples by proximity to x^*.
  5. Fit an interpretable model (e.g. sparse linear regression) on the weighted perturbed data.
  6. Use the local model’s coefficients as the explanation.

The fitting step minimises

\mathcal{L}(f, g, \pi_{x^*}) + \Omega(g)

where \mathcal{L} is the approximation error between explanation g and black-box f, \pi_{x^*} is the proximity kernel, and \Omega(g) penalises complexity.

A fast explanation with a fragile foundation

LIME works across tabular, text, and image data, stays model-agnostic so the black box underneath can be swapped freely, and produces short, selective explanations that are easy to hand to a non-technical audience, with a fidelity score attached to say how well the local model actually approximates the black box at that point. The fragility shows up in the construction itself. The neighbourhood is defined by a kernel width with no principled default, and that choice can change the explanation substantially. Perturbed samples are drawn without regard to feature correlations, so they can land on combinations that never occur in real data. Two nearby points can receive noticeably different explanations, and, like SHAP, a model can be deliberately built to produce LIME explanations that look reasonable while masking biased behaviour. Explanation quality also depends on the choice of interpretable representation used to build it, such as superpixels for images or word presence for text.

NoteCase Study: healthcare — LIME on a readmission-risk score

A hospital’s 30-day readmission-risk model flags one patient as high-risk. LIME perturbs that patient’s record (age, prior admissions, lab values, length of stay) around the instance, fits a local linear model, and returns a short list: a recent prior admission and one elevated lab value push risk upward, time since the last visit pushes it down. A discharge nurse can act on that short list directly, unlike a summary of 40 model inputs, but the explanation is a local approximation, not a full account of the model’s reasoning, so a second, more stable method (SHAP) is worth checking before it drives a real staffing or discharge decision.

Choosing between SHAP and LIME

SHAP LIME
Theoretical basis Game theory (Shapley values) Local surrogate model
Consistency Guaranteed (axiomatic) Not guaranteed
Stability High (exact for TreeSHAP) Low (kernel-width sensitive)
Speed Fast (TreeSHAP) Slower
Output Feature contributions (additive) Local model coefficients
Best for Tree-based models in production Quick exploration, non-tree models

In domains that use gradient boosting (XGBoost, LightGBM), such as insurance, credit, and healthcare, SHAP (TreeSHAP) is preferred over LIME for this reason.

Validation and communication

Are explanations faithful?

A key risk is that an explanation can be plausible without accurately reflecting the model’s reasoning. Doshi-Velez and Kim (2017) distinguish three ways to test this: against a real end-task (application-grounded), against a simplified task with human judges (human-grounded), and against a formal proxy criterion with no humans involved (functionally-grounded). The tests below draw on all three.

Tests to check:

  • Sanity check: do the most important features match domain knowledge?
  • Consistency: do similar individuals get similar explanations?
  • Simulation test: perturb one feature and check whether the SHAP value for that feature changes in the expected direction
  • Adversarial check: could these explanations be generated by a biased model that was designed to look fair? (Slack et al. 2020)
NoteCase Study: an unfaithful explanation of COMPAS

Rudin (2019) describes a concrete instance of this risk. ProPublica’s investigation of the COMPAS recidivism tool fit a linear explanation model that depended on race, and used it to argue the underlying, proprietary black-box model was racially biased. COMPAS’s developers disputed this, arguing the actual model may not depend on race beyond its correlation with age and criminal history, since COMPAS is plausibly nonlinear and never used race as an explicit predictor. If so, it was ProPublica’s explanation model, not the original black box, that depended on race, an unfaithful proxy mistaken for the thing it was meant to approximate. Whichever side of that specific dispute is correct, the general lesson holds: an explanation model can look damning, or reassuring, while not faithfully representing the black box it claims to explain.

Communicating explanations by audience

Explanations should be calibrated to the audience. The same SHAP output needs different framing for a decision subject vs an auditor. Roles vary by domain (validator, front-line decision-maker, decision subject, regulator, board), but the mapping of method to format is the same.

These methods are illustrative, not audience-exclusive. The same method (e.g. SHAP) can serve several audiences depending on the scope and framing applied to it. What actually differs by audience is the main question, the required scope, the level of technical detail, and the communication format.

NoteCase Study: insurance — communicating by audience
Audience Main question Scope Possible methods Communication
Actuary / validator Does the model behave sensibly? Global + local PFI, PDP/ALE, SHAP summary and local cases Technical report
Underwriter Why is this individual prediction unusual? Local SHAP waterfall, similar-case comparison Internal dashboard
Customer Why did I receive this outcome rather than another? Local, plain language Distilled local attribution or counterfactual explanation Letter or portal message
Regulator / auditor Is the model justified, fair and auditable? Global + subgroup + selected local PFI, PDP/ALE, SHAP summaries, fairness analysis and case reviews Regulatory documentation
Board What are the main risks and business implications? Aggregated global Key findings derived from several diagnostics Executive dashboard

The same structure recurs outside insurance. A hiring platform owes a rejected applicant a short, plain-language reason; a bank’s model-risk committee needs the aggregated technical picture; a health regulator auditing a triage algorithm needs the subgroup-level fairness view.

Common pitfalls in practice

  1. Explanation \neq causation: a high SHAP value for Age does not mean Age causes higher claims. It means the model uses Age heavily.

  2. Global \neq local: a feature ranked unimportant globally may be the dominant driver for an individual prediction.

  3. Correlation blindness: PDP and KernelSHAP marginalise over the data distribution and can extrapolate into unrealistic regions when features are correlated. ALE is designed specifically to fix this. It averages differences in predictions over conditional, not marginal, distributions (Apley and Zhu 2020). TreeSHAP is not a fix for the same problem. It is a tree-specific algorithm valued for computational efficiency and exactness. Whether it extrapolates into unrealistic combinations still depends on the SHAP convention chosen (interventional vs. tree-path-dependent).

  4. Instability of LIME: do not present a single LIME explanation as definitive. Run multiple times and check consistency.

  5. Confirmation bias: explanations that match prior beliefs are persuasive, but may be wrong. Always validate against held-out data.

  6. A flat PD plot is not proof of fairness: PD plots are commonly treated as a global sanity check for a protected or proxy variable, but their own extrapolation and averaging weaknesses can be deliberately exploited to produce a flat-looking plot while the underlying model stays discriminatory (Xin et al. 2025).

Regulatory expectations

Explainability is increasingly a regulatory requirement, not just good practice:

  • EU AI Act (2024–): high-risk AI systems, including credit scoring, employment screening, and life and health insurance risk assessment and pricing, must provide explanations to affected individuals
  • GDPR Article 22: automated decisions must be explainable and contestable
  • China’s PIPL (Article 24): grants individuals a right to an explanation of an automated decision, and a right to refuse a decision made solely by automated means where it significantly affects them (National People’s Congress (China) 2021); China’s 2026 NFRA guidance for banking and insurance separately names transparency and explainability as a required governance element for AI systems (National Financial Regulatory Administration (China) 2026)
  • Singapore’s MAS FEAT Transparency principle: requires institutions to document their AI systems internally and give customers meaningful explanations of AI-driven decisions that affect them (Monetary Authority of Singapore 2018)

Sector-specific rules add further detail: adverse-action notices for credit decisions (e.g. US ECOA/FCRA), bias-audit and notice requirements for automated employment-screening tools (e.g. NYC Local Law 144), and model documentation requirements for insurance rate filings (e.g. US state regulators, APRA CPG 234 in Australia).

NoteRegulatory evidence: the NAIC Model Review Manual

The NAIC Regulatory Review of Predictive Models white paper, which US state insurance regulators use when reviewing rating models, originally focused on GLMs. Its tree-based-model appendix extends the same review framework to random forests and gradient-boosted models, and explicitly asks insurers to submit variable-importance plots and interpretability plots, including PDPs, ALE plots, and Shapley plots, with at least one plot per model variable, “accompanied by commentary on why the visualized relationship is reasonable for variables of concern.” This is a concrete example of the methods introduced in this chapter being used as practical regulatory evidence, not just classroom exercises, which also means the PD-plot vulnerability described above is not a hypothetical concern for this industry. It targets precisely the review evidence regulators ask insurers to submit (Xin et al. 2025).

NoteRegulatory evidence: ASIC’s transparency review

Requiring an explanation on paper does not guarantee one in practice. A 2026 review by Australia’s securities and investments regulator, ASIC, examined the quote and renewal documents of five motor vehicle insurers and found that none explained the key factors that affected the calculation of the premium, or why the premium had changed from the previous year (Australian Securities and Investments Commission 2026). This is the compliance gap the methods in this chapter exist to close: a right to an explanation of automated pricing is only meaningful if insurers actually surface which factors moved the price and why, not just that a model produced a number (Huang 2026).

Explainability is not just good practice. It is becoming a compliance obligation across domains.

Explainability for large language models

The faithfulness question above becomes sharper, not different, for large language models (LLMs). Everything covered so far (PDP, ALE, SHAP, LIME) assumes a fixed, interpretable set of input features. An LLM has neither. Its “features” are subword tokens in a variable-length prompt, and its parameters number in the billions, so none of these methods apply directly.

Chain-of-thought is not a faithful explanation. A natural workaround is to ask the model to “think step by step” and treat the resulting text as its explanation. Turpin et al. (2023) show this is unreliable. When a biasing feature is quietly added to a prompt (e.g. always placing the correct multiple-choice answer in position A), models systematically shift their answers toward it without ever mentioning the bias in their stated reasoning, instead producing a plausible-sounding justification for the (now-biased) answer. This is the LLM version of the sanity/adversarial checks above. A stated explanation can look reasonable while not describing what actually drove the output.

Mechanistic interpretability is the emerging alternative. Rather than trusting the model’s self-report, this line of work opens up the model’s internals directly. Templeton et al. (2024) use sparse autoencoders (a neural-network technique that decomposes a model’s internal activity into distinct, nameable components) to extract millions of human-interpretable features from a production-scale LLM (Claude 3 Sonnet), showing that individual internal directions correspond to identifiable concepts. This is closer in spirit to a global, structural explanation than to a SHAP value, but it is still a research frontier, not yet a routine audit tool the way PDP or SHAP are for tabular models.

Practical implication. Don’t treat an LLM’s stated reasoning as equivalent to a SHAP value or a PDP curve. It has not been validated the same way, and Turpin et al. (2023)’s evidence suggests it often isn’t faithful. If an LLM’s output feeds a consequential decision, the audience-communication principle from earlier still applies. Be explicit about what kind of explanation you are actually giving (a self-report, not a validated attribution), since overclaiming faithfulness is itself a communication failure.

Summary

  • Permutation feature importance gives a global ranking of which features matter most.
  • PDP shows the average marginal effect of a feature (global, visual, intuitive). Use ALE when features are correlated.
  • SHAP decomposes each prediction into feature contributions. It is the most principled and versatile method, and TreeSHAP is fast for tree models.
  • LIME fits a local surrogate, useful for exploration but less stable than SHAP.
  • All explanations require validation. Plausibility is not the same as faithfulness.
  • Communication should be calibrated to the audience: technical, managerial, and customer-facing formats all differ.

References

Agarwal, Rishabh, Levi Melnick, Nicholas Frosst, et al. 2021. “Neural Additive Models: Interpretable Machine Learning with Neural Nets.” Advances in Neural Information Processing Systems 34.
Amnesty International. 2021. Xenophobic Machines: Discrimination Through Unregulated Use of Algorithms in the Dutch Childcare Benefits Scandal. Amnesty International. https://www.amnesty.org/en/documents/eur35/4686/2021/en/.
Apley, Daniel W, and Jingyu Zhu. 2020. “Visualizing the Effects of Predictor Variables in Black Box Supervised Learning Models.” Journal of the Royal Statistical Society Series B: Statistical Methodology 82 (4): 1059–86.
Australian Securities and Investments Commission. 2026. Report 838: Road Testing Transparency in Car Insurance Premiums. Australian Securities; Investments Commission. https://www.asic.gov.au/regulatory-resources/find-a-document/reports/rep-838-road-testing-transparency-in-car-insurance-premiums.
Biran, Or, and Courtenay Cotton. 2017. “Explanation and Justification in Machine Learning: A Survey.” IJCAI-17 Workshop on Explainable AI (XAI).
Breiman, Leo. 2001. “Random Forests.” Machine Learning 45 (1): 5–32.
Doshi-Velez, Finale, and Been Kim. 2017. “Towards a Rigorous Science of Interpretable Machine Learning.” arXiv Preprint arXiv:1702.08608.
Friedman, Jerome H. 2001. “Greedy Function Approximation: A Gradient Boosting Machine.” The Annals of Statistics 29 (5): 1189–232.
Friedman, Jerome H., and Bogdan E. Popescu. 2008. “Predictive Learning via Rule Ensembles.” The Annals of Applied Statistics 2 (3): 916–54.
Goldstein, Alex, Adam Kapelner, Justin Bleich, and Emil Pitkin. 2015. “Peeking Inside the Black Box: Visualizing Statistical Learning with Plots of Individual Conditional Expectation.” Journal of Computational and Graphical Statistics 24 (1): 44–65.
Hastie, Trevor J, and Robert J Tibshirani. 1990. Generalized Additive Models. Chapman; Hall/CRC.
Huang, Fei. 2026. “Insurers Have a Lot of Data about Us. Where Do They Get It, and How Do They Use It?” The Conversation, August. https://theconversation.com/insurers-have-a-lot-of-data-about-us-where-do-they-get-it-and-how-do-they-use-it-290053.
Kim, Been, Rajiv Khanna, and Oluwasanmi Koyejo. 2016. “Examples Are Not Enough, Learn to Criticize! Criticism for Interpretability.” Advances in Neural Information Processing Systems.
Lipton, Zachary C. 2018. “The Mythos of Model Interpretability.” Communications of the ACM 61 (10): 36–43.
Lundberg, Scott M, Gabriel Erion, Hugh Chen, et al. 2020. “From Local Explanations to Global Understanding with Explainable AI for Trees.” Nature Machine Intelligence 2 (1): 56–67.
Lundberg, Scott M, and Su-In Lee. 2017. “A Unified Approach to Interpreting Model Predictions.” Advances in Neural Information Processing Systems 30.
Miller, Tim. 2019. “Explanation in Artificial Intelligence: Insights from the Social Sciences.” Artificial Intelligence 267: 1–38.
Molnar, Christoph. 2025. Interpretable Machine Learning: A Guide for Making Black Box Models Explainable. 3rd ed. https://christophm.github.io/interpretable-ml-book/.
Monetary Authority of Singapore. 2018. “Principles to Promote Fairness, Ethics, Accountability and Transparency (FEAT) in the Use of Artificial Intelligence and Data Analytics in Singapore’s Financial Sector.” https://www.mas.gov.sg/publications/monographs-or-information-paper/2018/feat.
National Financial Regulatory Administration (China). 2026. Guidance on the Safe Development and Application of Artificial Intelligence in the Banking and Insurance Sectors (关于银行业保险业人工智能安全开发应用的指导意见, Jin Fa [2026] No. 8). https://www.nfra.gov.cn/cn/view/pages/governmentDetail.html?docId=1261784&generaltype=1.
National People’s Congress (China). 2021. Personal Information Protection Law of the People’s Republic of China (中华人民共和国个人信息保护法). http://www.npc.gov.cn/npc/c2/c30834/202108/t20210820_313088.html.
Ribeiro, Marco Tulio, Sameer Singh, and Carlos Guestrin. 2016. “Why Should i Trust You?: Explaining the Predictions of Any Classifier.” Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–44.
Rudin, Cynthia. 2019. “Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead.” Nature Machine Intelligence 1 (5): 206–15.
Rudin, Cynthia, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022. “Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges.” Statistics Surveys 16: 1–85.
Slack, Dylan, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020. “Fooling LIME and SHAP: Adversarial Attacks on Post Hoc Explanation Methods.” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 180–86.
Templeton, Adly, Tom Conerly, Jonathan Marcus, et al. 2024. “Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet.” Transformer Circuits Thread. https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html.
Tomsett, Richard, Dave Braines, Dan Harborne, Alun Preece, and Supriyo Chakraborty. 2018. “Interpretable to Whom? A Role-Based Model for Analyzing Interpretable Machine Learning Systems.” arXiv Preprint arXiv:1806.07552.
Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023. “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.” Advances in Neural Information Processing Systems 36.
Xin, Xi, Giles Hooker, and Fei Huang. 2025. “Pitfalls in Machine Learning Interpretability: Manipulating Partial Dependence Plots to Hide Discrimination.” Insurance: Mathematics and Economics 125: 103135. https://doi.org/10.1016/j.insmatheco.2025.103135.
Zhao, Qingyuan, and Trevor Hastie. 2021. “Causal Interpretations of Black-Box Models.” Journal of Business & Economic Statistics 39 (1): 272–81.