Explainability Principles
Quantitative Responsible AI: Principles, Governance, and Methods
Learning objectives
- Define explainability, and recognise interpretability and transparency as closely related terms often used interchangeably.
- Explain why explainability matters in high-stakes domains such as lending, hiring, healthcare, and insurance.
- Map stakeholder needs to different types of explanation.
- Classify explainability methods along the key taxonomic dimensions.
What is explainability?
Between 2013 and 2019, the Dutch tax authority used a self-learning algorithm to flag childcare-benefit applications as high-risk for fraud. The system weighted dual nationality as a risk factor, a proxy for ethnicity, and wrongly flagged around 26,000 families, disproportionately those with a migration background, who were then forced to repay tens of thousands of euros they did not owe. Affected families could not get a meaningful account of why they had been flagged, the oversight bodies meant to review individual cases could not properly interrogate the algorithm’s reasoning either, and the resulting financial ruin left over a thousand children removed from their homes. The scandal, known in the Netherlands as the toeslagenaffaire, brought down the Dutch government in January 2021 (Amnesty International 2021).
No single official decided to do this. Somewhere between “an algorithm flagged this family” and “children were taken into care,” there was no point at which a family, a caseworker, or a court could get a clear, checkable account of why. That gap, between a model producing an output and a human being able to understand and contest it, is what this chapter is about.
Core definitions
Explainability converts a model’s internal logic into a form a specific audience can understand, for a specific purpose. A model is explainable to the degree that a person can understand the cause of its decision (Biran and Cotton 2017), or, put operationally, to the degree that a user can correctly and efficiently predict what it will do (Kim et al. 2016).
Two closely related terms appear throughout the literature, and this course will not police the boundary between them sharply. Interpretability is often used for the same underlying idea. Transparency sometimes refers more specifically to understanding a model’s internal structure directly, as opposed to a post-hoc explanation constructed after training without exposing what the model actually computes internally (Lipton 2018). Rudin (2019) argues that the terminological churn between these words has itself strayed from the needs of real problems. What matters in each case is: for whom, about what, and for what purpose?
Why does explainability matter?
A model can be highly accurate and still be a liability if nobody can explain what it is doing. Several distinct reasons drive the case for explainability, and they don’t all point in the same direction (Doshi-Velez and Kim 2017; Molnar 2025):
| Reason | Description |
|---|---|
| Human learning | People need explanations to update mental models when outcomes are unexpected |
| Safety and testing | High-risk applications require confidence that models behave correctly |
| Bias detection | Explainability is a debugging tool for finding discriminatory patterns |
| Social acceptance | Transparent systems gain greater user trust and legitimacy |
| Auditing | Models can only be properly evaluated when their decisions are visible |
| Right to explanation | Regulatory requirements (GDPR, ASIC, EU AI Act) mandate explanations for automated decisions |
When is explainability not required?
Explainability is not free. It can cost accuracy, engineering time, and sometimes security. The case for it weakens when (Doshi-Velez and Kim 2017; Molnar 2025):
- The application has low stakes and no significant consequences from errors
- The method is well-established and thoroughly tested (e.g. standard actuarial tables, validated clinical scoring rules)
- Transparency enables gaming or manipulation (e.g. fraud scoring systems)
But most consequential domains this course covers (lending, hiring, healthcare, criminal justice, insurance) involve high-stakes decisions and regulatory scrutiny by default. Explainability is generally expected.
Demands for explanation across domains
Regulators, decision subjects, and internal reviewers each add specific demands for explainability:
- Regulators expect models to be documented, justified, and auditable
- Decision subjects have a right to understand why they were declined, screened out, or charged a particular price
- Practitioners signing off on a model’s use need to verify that outputs behave as expected
- Boards and senior management need to understand model risk
A black-box XGBoost model (a highly accurate model built by combining many decision trees, a technique called gradient boosting) may outperform a simpler model on predictive accuracy, yet be harder to document, harder to explain to an affected individual, and harder to audit for fairness.
Actuarial and insurance regulatory settings illustrate this concretely. Prudential and market-conduct regulators (e.g. APRA, ASIC, US state insurance departments) expect pricing and underwriting models to be filed and justified. Policyholders (the customers who hold an insurance contract) have a right to understand why they are charged a particular premium (the price paid for insurance cover) or denied a claim (a request for the insurer to pay out after a loss). Actuaries signing off on a model must verify its outputs behave as expected.
Goals of explainability
Three goals of explainability
Explainability serves several distinct purposes, not just one (Rudin et al. 2022):
1. Improving the model
Explainability helps identify when models learn unintended patterns. A model can have strong test-set accuracy while relying on non-causal features, for example using background snow in a photo to classify wolves vs dogs. In insurance, this might mean using a proxy variable (a feature that indirectly stands in for another) for a protected attribute (a legally protected personal characteristic, such as race or gender, that a model should not rely on) without detection.
2. Justifying predictions
Different stakeholders require different explanations. Decision-makers need to verify forecasts against domain knowledge. Customers need to understand and contest decisions. Regulators need to verify compliance.
3. Discovering insights
Beyond predictions, explainability reveals relationships between inputs and outputs, useful for product design, pricing strategy, and scientific understanding.
Characteristics of good explanations
Research on how people actually process explanations, not just what is technically correct, identifies what makes them effective (Miller 2019):
| Principle | Implication |
|---|---|
| Contrastive | “Why this prediction rather than another?” is more useful than a full causal chain |
| Selective | 1–2 key reasons are accepted even when many factors contribute |
| Social | Explanations vary by audience — a customer needs a different explanation than an actuary |
| Abnormality focus | Rare or unexpected features make stronger explanations than common ones |
| Consistent with beliefs | Explanations aligning with prior knowledge are more persuasive (but watch for confirmation bias) |
| Faithful | Explanations should truthfully reflect model behaviour, not just sound plausible |
Stakeholders
Who needs explanations?
A useful way to organise stakeholder needs is by the role someone plays in relation to a model, since the same person can occupy different roles for different systems (Tomsett et al. 2018). Six roles recur across domains:
| Stakeholder | Need |
|---|---|
| Model creators (data scientists) | Debugging, feature engineering, model selection |
| Operators (e.g. insurers, banks, hospitals, employers) | Compliance, governance, sign-off |
| Executors (e.g. underwriters, loan officers, clinicians, hiring managers) | Verify model outputs before acting on them |
| Decision subjects (e.g. policyholders, applicants, patients) | Understand why an outcome or decision differs |
| Auditors (regulators, internal audit) | Validate fairness, accuracy, compliance |
| Data subjects | Challenge automated decisions (GDPR right to explanation) |
Different stakeholders need different types of explanation, such as local versus global, or technical versus plain language.
Mapping stakeholders to explanations
Concretely, the mapping of stakeholder to explanation type looks different in every domain, but the structure is the same. Identify the question being asked, then match scope (global vs local) and audience (technical vs plain language).
| Stakeholder | Typical question | Explanation type needed |
|---|---|---|
| Actuary signing off | Does the model behave sensibly across the rating factors (the input variables used to set a price, e.g. age or location)? | Global, feature effects |
| Underwriter | Why is this applicant’s premium higher than expected? | Local, individual prediction |
| Customer | Why did my premium increase this year? | Local, contrastive, plain language |
| Regulator / auditor | Does this model treat protected groups fairly? | Global, group-level, with statistical tests |
| Board / management | What are the top drivers of claims costs this year? | Global, feature importance |
The same mapping recurs elsewhere. A loan officer asking why an applicant was declined needs a local explanation, and a regulator auditing a hiring-screen tool for disparate impact (when a seemingly neutral policy ends up harming one group more than another) needs a global, group-level one.
Taxonomy of methods
Two primary categories
Explainability by design: train models that are inherently interpretable.
- Linear and logistic regression
- Generalised linear models (GLMs, a model that predicts an outcome as a weighted sum of the input features) and generalised additive models (GAMs, a more flexible version that allows curved rather than straight-line effects)
- Decision trees (models that split data by a series of branching yes/no questions) and decision rules
- RuleFit (a model combining decision rules with a linear model) (Friedman and Popescu 2008)
Post-hoc explainability: apply methods after model training to explain an existing (possibly black-box) model.
The choice between them is a fundamental design decision made before modelling, not an afterthought.
Dimension 1: Intrinsic vs post-hoc
| Intrinsic (by design) | Post-hoc | |
|---|---|---|
| How | Interpretable model structure | Explanation applied after training |
| Examples | GLM, decision tree, GAM | SHAP, PDP, LIME |
| Advantage | Explanation is exact | Works on any model |
| Limitation | May sacrifice predictive power | Approximation; may not be faithful |
In regulated scoring settings (insurance pricing, credit scoring), GLMs are the standard because they are intrinsically interpretable and easy to document. Machine learning models require post-hoc explanation.
Dimension 2: Model-agnostic vs model-specific
Model-agnostic: works with any model. Treats the model as a black box and queries it for predictions.
- Uses the SIPA principle: Sample data → Perform an intervention → get Predictions → Aggregate (Molnar 2025)
- Examples: SHAP (KernelSHAP), LIME, PDP, permutation feature importance
Model-specific: designed for a particular model class.
- Examples: GLM coefficients, decision tree paths, TreeSHAP (for trees), saliency maps (for neural networks)
- Faster and often more accurate, but tied to the model architecture
Dimension 3: Global vs local
Global: describes average model behaviour across the dataset.
- “What features does this model rely on overall?”
- Examples: permutation feature importance, PDP, ALE, surrogate models (simple models trained to approximate a complex one)
Local: explains a single prediction for one individual.
- “Why did this particular applicant receive this premium?”
- Examples: SHAP values for one observation, LIME, counterfactual explanations (statements like “if this input had been different, the outcome would have changed”)
Both are needed in practice. Global methods support auditing and governance. Local methods support customer explanations and contestability.
Summary taxonomy
The three dimensions are independent axes, not a nested hierarchy. A method’s position on one does not determine its position on the others.
| Dimension | Values | Examples |
|---|---|---|
| Origin | Intrinsic vs post-hoc | GLM coefficients, decision-tree paths = intrinsic; KernelSHAP, TreeSHAP, PDP, LIME = post-hoc |
| Applicability | Model-agnostic vs model-specific | KernelSHAP, PDP, PFI, LIME = model-agnostic; GLM coefficients, decision-tree paths, TreeSHAP = model-specific |
| Scope | Global vs local | Mean absolute SHAP values (aggregated), PFI, PDP = global; SHAP value for one observation, LIME = local |
Any combination is possible, e.g.:
- GLM coefficients and decision-tree paths: intrinsic + model-specific
- KernelSHAP: post-hoc + model-agnostic
- TreeSHAP: post-hoc + model-specific
- A single SHAP value: local
- The mean absolute SHAP value for a feature across all observations: global
Explainable models
Generalised linear models
Generalised linear models are a standard tool in regulated scoring settings (credit scoring, clinical risk, insurance pricing) precisely because they are intrinsically interpretable.
A Poisson GLM, e.g. for a count outcome such as claim frequency (how often a policyholder files a claim) or hospital admissions:
\log \mathbb{E}[Y_i] = \beta_0 + \beta_1 x_{i1} + \cdots + \beta_p x_{ip} + \log(\text{exposure}_i)
Each coefficient \beta_j has a direct interpretation. A one-unit increase in x_j multiplies expected frequency by e^{\beta_j}, holding all else equal.
Limitations: assumes a specific functional form and cannot capture complex non-linear interactions automatically.
Decision trees
A decision tree partitions the feature space into rectangular regions and assigns a constant prediction to each.
Strengths: highly interpretable, can be visualised as a flow chart.
Limitations: unstable (small data changes cause large tree changes), tend to overfit, and poor at smooth non-linear relationships.
Decision trees are rarely used directly for pricing or scoring in regulated settings, but are useful for communicating segmentation logic to non-technical stakeholders, for example in insurance, lending, and clinical triage.
GAMs: extending GLMs
A Generalised Additive Model (GAM) (Hastie and Tibshirani 1990) replaces linear terms with smooth functions:
g(\mathbb{E}[Y]) = \beta_0 + f_1(x_1) + f_2(x_2) + \cdots + f_p(x_p)
Each f_j is a smooth, visualisable function, often a spline (a flexible curve built from smooth pieces). GAMs are more flexible than GLMs while remaining interpretable through plots of each f_j.
Neural additive models (NAMs) (Agarwal et al. 2021) extend the idea to deep learning with interpretable additive structure.
Limits of explainability
The accuracy–explainability trade-off
More flexible models (gradient boosting, neural networks) tend to be more accurate but less explainable. Explainable models (GLMs, decision trees) are transparent but may underfit complex patterns.
This trade-off is widely assumed, but Rudin (2019) argues it is often a myth, particularly on structured data with meaningful, well-engineered features: once such features exist, complex black-box models frequently show little or no accuracy advantage over a well-constructed interpretable one. The trade-off is more likely to be real where useful features are not hand-built, such as raw images or audio, but it should be checked empirically for a given problem rather than assumed. Where a genuine trade-off does exist, methods like GAMs and constrained gradient boosting can narrow it. Post-hoc explanation of a black-box model is always an approximation, and the explanation may not be fully faithful to the model’s actual reasoning.
Human factors in explanation
Technical accuracy is not sufficient on its own. An explanation only does its job if a person can actually use it (Miller 2019):
- Cognitive load: simpler explanations are preferred even if incomplete
- Selectivity: 1–2 reasons are easier to act on than 20 feature contributions
- Trust calibration: overly confident explanations can cause over-reliance on a flawed model
In professional settings, the risk of a plausible but misleading explanation is at least as serious as no explanation at all.
Summary
- Explainability is not an end in itself. It is a means to accountability, trust, and governance.
- The choice of method depends on who needs the explanation, what they need to understand, and what decision follows.
- The fundamental taxonomy: intrinsic vs post-hoc, model-agnostic vs model-specific, global vs local.
- In regulated scoring domains (insurance, credit, clinical risk), GLMs provide explainability by design. Machine learning models require careful post-hoc explanation.
Post-hoc explainability methods
Overview
This section covers the four most widely used post-hoc, model-agnostic methods in practice:
| Method | Scope | What it answers |
|---|---|---|
| Permutation Feature Importance (PFI) | Global | Which features matter most to this model? |
| Partial Dependence Plots (PDP) | Global | How does a feature affect predictions on average? |
| SHAP | Local + Global | How much did each feature contribute to this prediction? |
| LIME | Local | What simple rule approximates the model near this prediction? |
All methods treat the model as a black box and work for any model type (GLM, XGBoost, neural network).
Permutation Feature Importance
Break a feature, watch the damage
Breiman (2001) introduced this idea as a variable-importance measure for random forests, but the underlying logic transfers to any fitted model, not just random forests: if a model genuinely relies on a feature, scrambling that feature’s values should hurt its predictions. If shuffling a column barely changes performance, the model was not using it.
The appeal is that this works without touching the model itself: no retraining, no access to internals, just repeated scoring on deliberately corrupted data.
Computing it
- Fit the model and compute baseline error e_0 on test data.
- For each feature j = 1, \ldots, p:
- Permute (shuffle) column j in the test set.
- Re-score the model on the shuffled data, with no retraining.
- Compute error e_j on the permuted data.
- Importance: I_j = e_j / e_0 (ratio) or e_j - e_0 (difference).
- Rank features by I_j.
What it tells you, and where it breaks down
PFI is model-agnostic, captures interaction effects (shuffling one feature disrupts every interaction it participates in), and compresses an entire model into a single ranked list, all without retraining. But it is a coarse instrument. Permuting a feature that is correlated with others creates combinations that never occur in real data, so an importance score can reflect an artefact of the shuffle rather than genuine model behaviour. It also only shows that a feature matters, not how, in which direction, or over what range, and a single run is noisy enough that results should be averaged over several permutations. If the model itself does not generalise well, PFI faithfully reports the model’s reliance on training noise rather than real signal.
Reading a PFI ranking
A PFI ranking surfaces two kinds of signal at once. It shows which legitimate factors the model actually relies on, and which non-legitimate or proxy features contribute more than domain knowledge would suggest.
In a motor insurance pricing model, PFI might reveal:
- Age and Bonus (a driver’s no-claims discount level) are the most important legitimate rating factors
- InsuranceScore (a non-legitimate proxy) contributes significantly, signalling that fairness scrutiny is needed
- Density and Value contribute less than expected from domain knowledge, prompting a check of data quality
This is often the first diagnostic run after fitting a new model, in any domain. A hiring model whose top feature turns out to be a proxy for age or gender warrants the same scrutiny.
Partial Dependence Plots
Averaging out everything except one feature
Friedman (2001) introduced PDP alongside gradient boosting, as a way to ask a question that a coefficient cannot answer for a black-box model: if every observation in the data shared the same value of one feature, what would the average prediction look like? The construction is direct. Fix the feature of interest to a value, force every observation to take it, score the model, and average the result. Repeat across the feature’s range and the resulting curve is the partial dependence function.
Formula
For feature x_j, the partial dependence function is:
\hat{f}_j(x_j) = \mathbb{E}_{X_{-j}}\left[\hat{f}(x_j, X_{-j})\right] \approx \frac{1}{n}\sum_{i=1}^n \hat{f}(x_j, x_{-j}^{(i)})
We substitute x_j for all observations and average, marginalising over the joint distribution of the remaining features.
What the curve tells you
A PDP is intuitive to read: direction (does the prediction rise or fall as the feature increases?), shape (linear, monotone, or non-linear, such as U-shaped with age), and magnitude (how much the prediction moves across the observed range) are all visible at a glance, for any model type, without assuming a functional form. A causal reading of that shape is not automatic. It requires additional causal assumptions on top of the plot itself (Zhao and Hastie 2021).
For two features, a 2D PDP surface can reveal possible interaction, but the joint surface also contains both features’ main effects, so confirming genuine interaction requires comparing it against the sum of the two one-feature PDPs, or using a dedicated interaction measure (H-statistic, SHAP interaction values, Ch5). Visualisation also becomes difficult beyond two variables.
Where the average misleads you
Two assumptions do most of the damage. PDP assumes feature independence: when features are correlated, it can extrapolate to combinations that are rarely or never observed in the data, and that extrapolation region is exactly where the plot is least trustworthy. And because PDP is an average, it can mask heterogeneous effects: if the effect differs across subgroups, or points in opposite directions for different individuals, the average can look flat or misleading even though no individual actually experiences that average effect. On top of this, PDP is only practical for one or two features at a time, and it does not show the underlying feature distribution, which invites over-interpreting effects in sparse regions of the plot.
Individual Conditional Expectation (ICE) curves (Goldstein et al. 2015) address the heterogeneous-effects limitation. A PDP is the average of many ICE curves, one per observation, each showing how a single instance’s prediction would change as the feature of interest varies, holding everything else about that instance fixed. Overlaying ICE curves on top of a PDP reveals whether the average effect is representative of most observations, or whether it masks subgroup effects that point in different directions and cancel out in the average. ICE curves still assume feature independence, so they remain vulnerable to the same extrapolation problem as PDP when features are correlated.
Accumulated Local Effects (ALE) plots (Apley and Zhu 2020) address the independence assumption by conditioning on local neighbourhoods rather than marginalising globally. Prefer ALE when features are correlated.
The two weaknesses above are not just textbook caveats. Xin et al. (2025) show they can be deliberately exploited. Their adversarial framework leaves a black-box model’s predictions untouched on the real, densely-observed region of the data, but substitutes a separately chosen, non-discriminatory output whenever an input falls in the sparse “extrapolation” region that PDP’s averaging step actually depends on. Because that region is rarely or never observed in practice, almost none of the model’s real-world predictions change, yet the resulting PD plot for the manipulated feature can be engineered to look essentially flat.
Applied to real motor insurance claims data and to the COMPAS recidivism dataset, the attack produced PD plots for a protected or proxy variable (age, race) that showed no visible disparity, while the underlying model’s predictions, and its discriminatory behaviour, were preserved almost entirely, with over 95% accuracy retained relative to the original model’s outputs. This directly undermines the “the protected-attribute PD plot looks flat, so the model is fine” check used elsewhere in this course. ICE curves, correlation checks between features, and ALE plots (above) make the manipulation harder to construct convincingly, but none of them make it impossible. Treat a flat PD plot as one input to a fairness review, not as proof of one.
Reading a PDP in practice
A PDP for a numeric feature such as age often shows a non-linear, non-monotonic effect that a GLM coefficient cannot represent directly, but the PDP makes it visible for any model type.
A PDP for Age in a motor insurance frequency model might show high predicted frequency for young drivers (18–25), declining steeply to ~30, then gradually rising after 65.
The same non-monotonic shape shows up elsewhere. Years of prior work experience rarely pushes a hiring model’s output in one straight line, and tumour size rarely moves a clinical risk model’s prediction linearly either. A PDP makes that curvature visible without having to guess its functional form in advance.
SHAP
Splitting credit for a prediction
SHAP (SHapley Additive exPlanations) (Lundberg and Lee 2017) borrows its answer from cooperative game theory: if several players jointly produce a payoff, how should it be split among them, based on each one’s actual contribution? Treat a model’s features as the players and its prediction as the shared payoff, and the fair split is the Shapley value, the average marginal contribution of feature j across every possible combination of the other features it could appear alongside.
Explanation model:
g(z') = \phi_0 + \sum_{j=1}^p \phi_j z_j'
where \phi_j is the Shapley value for feature j and \phi_0 is the baseline (average prediction).
Why Shapley values, specifically
Game theory guarantees this particular split satisfies three properties no other feature-attribution method achieves simultaneously (Lundberg and Lee 2017): local accuracy (the contributions sum exactly to the gap between this prediction and the baseline, \sum_j \phi_j = \hat{f}(x) - \mathbb{E}[\hat{f}]), missingness (a feature absent from the model gets zero credit), and consistency (if a feature’s true marginal contribution goes up after a model change, its Shapley value cannot go down).
Two ways to compute it
KernelSHAP is model-agnostic: it samples feature coalitions, scores them, and fits a weighted linear regression with the SHAP kernel to recover the \phi_j. It works for any model but is slow.
TreeSHAP (Lundberg et al. 2020) exploits tree structure to compute exact values in polynomial rather than exponential time, for XGBoost, LightGBM, and random forests. It is the default choice for gradient boosting models across domains, including insurance pricing, credit scoring, and clinical risk.
Reading a SHAP output
A waterfall plot decomposes one prediction into the baseline plus each feature’s push or pull: start at the baseline E[\hat{f}] (the average prediction across the population), let each feature push the prediction up (positive SHAP value) or down (negative value), and the sum of every push and pull lands exactly on this individual’s prediction.
For a single policyholder’s motor premium prediction, age (young) pushes the prediction up and bonus (low, i.e. a strong no-claims record) pushes it down. This is the natural format for a customer-facing explanation: “Your premium is higher than average mainly because of your age and driving history.”
The same waterfall format works for a declined loan application or a clinical risk score, expressed as a sum of contributions from income, credit history, or lab values relative to a baseline. Beyond the waterfall, a summary plot shows importance, direction, and feature value across every observation at once; a dependence plot shows SHAP value against feature value for one feature, the local counterpart to a PDP; the mean absolute SHAP value aggregates local values into a global importance ranking; and interaction values show how pairs of features jointly move a prediction.
What SHAP buys you, and what it does not
The game-theoretic guarantees are real strengths: local accuracy, missingness, and consistency hold exactly, local values aggregate cleanly into a global summary, and TreeSHAP makes all of this fast for tree-based models. But the guarantees have limits. KernelSHAP is slow without a tree-specific shortcut, and it assumes feature independence, an assumption TreeSHAP’s conditional variant relaxes at the cost of sometimes assigning non-zero credit to a correlated feature the model never actually used. Shapley values are attributions, not causes, and treating a large \phi_j as evidence that a feature causes the outcome is a category error. And, like any post-hoc explanation, SHAP explanations can be adversarially manipulated to hide biases in the underlying model (Slack et al. 2020).
LIME
Zoom in until the model looks simple
LIME (Local Interpretable Model-Agnostic Explanations) (Ribeiro et al. 2016) starts from a different bet than SHAP’s game-theoretic split: a model that is hopelessly complex globally may behave almost linearly in a small enough neighbourhood around one point. Fit a simple, interpretable model in that neighbourhood, and its coefficients become the explanation for the complex model’s prediction there.
Building the local model
- Select the instance x^* to explain.
- Perturb the data: generate samples near x^* by sampling from a normal distribution (tabular), removing words (text), or toggling superpixels (images).
- Score the black-box model on perturbed samples.
- Weight samples by proximity to x^*.
- Fit an interpretable model (e.g. sparse linear regression) on the weighted perturbed data.
- Use the local model’s coefficients as the explanation.
The fitting step minimises
\mathcal{L}(f, g, \pi_{x^*}) + \Omega(g)
where \mathcal{L} is the approximation error between explanation g and black-box f, \pi_{x^*} is the proximity kernel, and \Omega(g) penalises complexity.
A fast explanation with a fragile foundation
LIME works across tabular, text, and image data, stays model-agnostic so the black box underneath can be swapped freely, and produces short, selective explanations that are easy to hand to a non-technical audience, with a fidelity score attached to say how well the local model actually approximates the black box at that point. The fragility shows up in the construction itself. The neighbourhood is defined by a kernel width with no principled default, and that choice can change the explanation substantially. Perturbed samples are drawn without regard to feature correlations, so they can land on combinations that never occur in real data. Two nearby points can receive noticeably different explanations, and, like SHAP, a model can be deliberately built to produce LIME explanations that look reasonable while masking biased behaviour. Explanation quality also depends on the choice of interpretable representation used to build it, such as superpixels for images or word presence for text.
A hospital’s 30-day readmission-risk model flags one patient as high-risk. LIME perturbs that patient’s record (age, prior admissions, lab values, length of stay) around the instance, fits a local linear model, and returns a short list: a recent prior admission and one elevated lab value push risk upward, time since the last visit pushes it down. A discharge nurse can act on that short list directly, unlike a summary of 40 model inputs, but the explanation is a local approximation, not a full account of the model’s reasoning, so a second, more stable method (SHAP) is worth checking before it drives a real staffing or discharge decision.
Choosing between SHAP and LIME
| SHAP | LIME | |
|---|---|---|
| Theoretical basis | Game theory (Shapley values) | Local surrogate model |
| Consistency | Guaranteed (axiomatic) | Not guaranteed |
| Stability | High (exact for TreeSHAP) | Low (kernel-width sensitive) |
| Speed | Fast (TreeSHAP) | Slower |
| Output | Feature contributions (additive) | Local model coefficients |
| Best for | Tree-based models in production | Quick exploration, non-tree models |
In domains that use gradient boosting (XGBoost, LightGBM), such as insurance, credit, and healthcare, SHAP (TreeSHAP) is preferred over LIME for this reason.
Validation and communication
Are explanations faithful?
A key risk is that an explanation can be plausible without accurately reflecting the model’s reasoning. Doshi-Velez and Kim (2017) distinguish three ways to test this: against a real end-task (application-grounded), against a simplified task with human judges (human-grounded), and against a formal proxy criterion with no humans involved (functionally-grounded). The tests below draw on all three.
Tests to check:
- Sanity check: do the most important features match domain knowledge?
- Consistency: do similar individuals get similar explanations?
- Simulation test: perturb one feature and check whether the SHAP value for that feature changes in the expected direction
- Adversarial check: could these explanations be generated by a biased model that was designed to look fair? (Slack et al. 2020)
Rudin (2019) describes a concrete instance of this risk. ProPublica’s investigation of the COMPAS recidivism tool fit a linear explanation model that depended on race, and used it to argue the underlying, proprietary black-box model was racially biased. COMPAS’s developers disputed this, arguing the actual model may not depend on race beyond its correlation with age and criminal history, since COMPAS is plausibly nonlinear and never used race as an explicit predictor. If so, it was ProPublica’s explanation model, not the original black box, that depended on race, an unfaithful proxy mistaken for the thing it was meant to approximate. Whichever side of that specific dispute is correct, the general lesson holds: an explanation model can look damning, or reassuring, while not faithfully representing the black box it claims to explain.
Communicating explanations by audience
Explanations should be calibrated to the audience. The same SHAP output needs different framing for a decision subject vs an auditor. Roles vary by domain (validator, front-line decision-maker, decision subject, regulator, board), but the mapping of method to format is the same.
These methods are illustrative, not audience-exclusive. The same method (e.g. SHAP) can serve several audiences depending on the scope and framing applied to it. What actually differs by audience is the main question, the required scope, the level of technical detail, and the communication format.
| Audience | Main question | Scope | Possible methods | Communication |
|---|---|---|---|---|
| Actuary / validator | Does the model behave sensibly? | Global + local | PFI, PDP/ALE, SHAP summary and local cases | Technical report |
| Underwriter | Why is this individual prediction unusual? | Local | SHAP waterfall, similar-case comparison | Internal dashboard |
| Customer | Why did I receive this outcome rather than another? | Local, plain language | Distilled local attribution or counterfactual explanation | Letter or portal message |
| Regulator / auditor | Is the model justified, fair and auditable? | Global + subgroup + selected local | PFI, PDP/ALE, SHAP summaries, fairness analysis and case reviews | Regulatory documentation |
| Board | What are the main risks and business implications? | Aggregated global | Key findings derived from several diagnostics | Executive dashboard |
The same structure recurs outside insurance. A hiring platform owes a rejected applicant a short, plain-language reason; a bank’s model-risk committee needs the aggregated technical picture; a health regulator auditing a triage algorithm needs the subgroup-level fairness view.
Common pitfalls in practice
Explanation \neq causation: a high SHAP value for Age does not mean Age causes higher claims. It means the model uses Age heavily.
Global \neq local: a feature ranked unimportant globally may be the dominant driver for an individual prediction.
Correlation blindness: PDP and KernelSHAP marginalise over the data distribution and can extrapolate into unrealistic regions when features are correlated. ALE is designed specifically to fix this. It averages differences in predictions over conditional, not marginal, distributions (Apley and Zhu 2020). TreeSHAP is not a fix for the same problem. It is a tree-specific algorithm valued for computational efficiency and exactness. Whether it extrapolates into unrealistic combinations still depends on the SHAP convention chosen (interventional vs. tree-path-dependent).
Instability of LIME: do not present a single LIME explanation as definitive. Run multiple times and check consistency.
Confirmation bias: explanations that match prior beliefs are persuasive, but may be wrong. Always validate against held-out data.
A flat PD plot is not proof of fairness: PD plots are commonly treated as a global sanity check for a protected or proxy variable, but their own extrapolation and averaging weaknesses can be deliberately exploited to produce a flat-looking plot while the underlying model stays discriminatory (Xin et al. 2025).
Regulatory expectations
Explainability is increasingly a regulatory requirement, not just good practice:
- EU AI Act (2024–): high-risk AI systems, including credit scoring, employment screening, and life and health insurance risk assessment and pricing, must provide explanations to affected individuals
- GDPR Article 22: automated decisions must be explainable and contestable
- China’s PIPL (Article 24): grants individuals a right to an explanation of an automated decision, and a right to refuse a decision made solely by automated means where it significantly affects them (National People’s Congress (China) 2021); China’s 2026 NFRA guidance for banking and insurance separately names transparency and explainability as a required governance element for AI systems (National Financial Regulatory Administration (China) 2026)
- Singapore’s MAS FEAT Transparency principle: requires institutions to document their AI systems internally and give customers meaningful explanations of AI-driven decisions that affect them (Monetary Authority of Singapore 2018)
Sector-specific rules add further detail: adverse-action notices for credit decisions (e.g. US ECOA/FCRA), bias-audit and notice requirements for automated employment-screening tools (e.g. NYC Local Law 144), and model documentation requirements for insurance rate filings (e.g. US state regulators, APRA CPG 234 in Australia).
The NAIC Regulatory Review of Predictive Models white paper, which US state insurance regulators use when reviewing rating models, originally focused on GLMs. Its tree-based-model appendix extends the same review framework to random forests and gradient-boosted models, and explicitly asks insurers to submit variable-importance plots and interpretability plots, including PDPs, ALE plots, and Shapley plots, with at least one plot per model variable, “accompanied by commentary on why the visualized relationship is reasonable for variables of concern.” This is a concrete example of the methods introduced in this chapter being used as practical regulatory evidence, not just classroom exercises, which also means the PD-plot vulnerability described above is not a hypothetical concern for this industry. It targets precisely the review evidence regulators ask insurers to submit (Xin et al. 2025).
Requiring an explanation on paper does not guarantee one in practice. A 2026 review by Australia’s securities and investments regulator, ASIC, examined the quote and renewal documents of five motor vehicle insurers and found that none explained the key factors that affected the calculation of the premium, or why the premium had changed from the previous year (Australian Securities and Investments Commission 2026). This is the compliance gap the methods in this chapter exist to close: a right to an explanation of automated pricing is only meaningful if insurers actually surface which factors moved the price and why, not just that a model produced a number (Huang 2026).
Explainability is not just good practice. It is becoming a compliance obligation across domains.
Explainability for large language models
The faithfulness question above becomes sharper, not different, for large language models (LLMs). Everything covered so far (PDP, ALE, SHAP, LIME) assumes a fixed, interpretable set of input features. An LLM has neither. Its “features” are subword tokens in a variable-length prompt, and its parameters number in the billions, so none of these methods apply directly.
Chain-of-thought is not a faithful explanation. A natural workaround is to ask the model to “think step by step” and treat the resulting text as its explanation. Turpin et al. (2023) show this is unreliable. When a biasing feature is quietly added to a prompt (e.g. always placing the correct multiple-choice answer in position A), models systematically shift their answers toward it without ever mentioning the bias in their stated reasoning, instead producing a plausible-sounding justification for the (now-biased) answer. This is the LLM version of the sanity/adversarial checks above. A stated explanation can look reasonable while not describing what actually drove the output.
Mechanistic interpretability is the emerging alternative. Rather than trusting the model’s self-report, this line of work opens up the model’s internals directly. Templeton et al. (2024) use sparse autoencoders (a neural-network technique that decomposes a model’s internal activity into distinct, nameable components) to extract millions of human-interpretable features from a production-scale LLM (Claude 3 Sonnet), showing that individual internal directions correspond to identifiable concepts. This is closer in spirit to a global, structural explanation than to a SHAP value, but it is still a research frontier, not yet a routine audit tool the way PDP or SHAP are for tabular models.
Practical implication. Don’t treat an LLM’s stated reasoning as equivalent to a SHAP value or a PDP curve. It has not been validated the same way, and Turpin et al. (2023)’s evidence suggests it often isn’t faithful. If an LLM’s output feeds a consequential decision, the audience-communication principle from earlier still applies. Be explicit about what kind of explanation you are actually giving (a self-report, not a validated attribution), since overclaiming faithfulness is itself a communication failure.
Summary
- Permutation feature importance gives a global ranking of which features matter most.
- PDP shows the average marginal effect of a feature (global, visual, intuitive). Use ALE when features are correlated.
- SHAP decomposes each prediction into feature contributions. It is the most principled and versatile method, and TreeSHAP is fast for tree models.
- LIME fits a local surrogate, useful for exploration but less stable than SHAP.
- All explanations require validation. Plausibility is not the same as faithfulness.
- Communication should be calibrated to the audience: technical, managerial, and customer-facing formats all differ.
Recommended reading
- Molnar (2025) — christophm.github.io/interpretable-ml-book, free web textbook, Chapters on PDP, PFI, SHAP, LIME
- Rudin et al. (2022) — comprehensive peer-reviewed survey, fundamental principles and open challenges
- Miller (2019) — the social-science basis for what makes an explanation effective
- Tomsett et al. (2018) — the role-based stakeholder taxonomy used above
- Lipton (2018) — decomposing “interpretability” into transparency and post-hoc explanation
- Lundberg and Lee (2017) — original SHAP paper
- Ribeiro et al. (2016) — original LIME paper
- Slack et al. (2020) — adversarial manipulation of SHAP and LIME explanations
- Xin et al. (2025) — adversarial manipulation of PD plots to conceal discrimination, validated on real insurance and COMPAS data
- Rudin (2019) — argument for using interpretable models in high-stakes decisions, including the COMPAS explanation-fidelity case discussed above
- Doshi-Velez and Kim (2017) — “Towards a rigorous science of interpretable machine learning”
- Turpin et al. (2023); Templeton et al. (2024) — explainability specifically for large language models: chain-of-thought faithfulness and mechanistic interpretability