Slides — Chapter 5
~2-hour session, hands-on, with five parts and five discussion breaks
| Time | Part |
|---|---|
| 0:00 – 0:10 | Motivation: why interpretability, regulatory context |
| 0:10 – 0:25 | Setup, workflow and case study |
| 0:25 – 0:50 | Feature importance |
| 0:50 – 1:15 | Main effects, PDP and ALE |
| 1:15 – 1:25 | Break |
| 1:25 – 1:55 | SHAP, global and local |
| 1:55 – 2:10 | Interactions and summary |
AI adoption across industries, including insurance, has led to widespread use of complex black-box models.
What are the benefits, and the potential risks, of that shift?
Before the list: what would you say?
Interpretability, explainability, and transparency are recurring requirements across AI principle documents:
(Partially) interpretable / transparent models
Global model-agnostic methods
Local model-agnostic methods
Same sequence for any tabular model, any domain.
pg15trainingThe workflow is: fit → global importance → main effects → local explanations → interactions → validate/communicate.
Which step gets skipped most often under deadline pressure, and what’s the risk?
SubGroup2 dominates, partly an artifact. Many categories → many splitting points → inflated total gain. Bonus, Age, Density look less important after grouping, despite large individual encoded-feature gains.
Model-agnostic (works the same way for any model type). Shuffle a column and measure the performance drop, no tree structure needed (Breiman 2001).
Built-in gain says SubGroup2 dominates, partly a category-count artifact.
If you only had time for ONE importance method before a deadline, which would you trust more, and why?
Strong nonlinear effect. Very young drivers → much higher predicted frequency, declining rapidly to ~30, then stable. PDP and ALE closely agree. The effect is robust, not an artifact of correlation with other predictors.
Approximately monotonic. Lower bonus (discount, clean history) → lower predicted frequency. Again PDP and ALE agree closely. The model has learned the expected relationship.
PDP/ALE describe average effects. A regulator or policyholder asks. Why does this policy pay this premium?
SHAP bridges both, a portfolio-level ranking (consistent with PDP/ALE) and a decomposition of one prediction into feature contributions.
Important
Interpretation ≠ trustworthiness. SHAP and LIME can be adversarially manipulated to mask discriminatory behaviour (Slack et al. 2020). Pair explanations with validation and fairness assessment, not as proof on their own.
Ranks variables by mean |SHAP value|, similar spirit to PFI (permutation feature importance), based on actual contribution size.
Each point = one policy × one feature. Position = SHAP value (pushes prediction up/down), and colour = feature value.
The beeswarm shows both magnitude and direction, across every policy.
What can it tell you that a simple bar-chart importance ranking can’t?
The natural format for a customer-facing explanation. For example, “your premium is higher mainly because of X and Y.”
What would happen if one characteristic of this policy changed?
Vary one predictor at a time from the selected policy, holding the rest fixed at the median for continuous predictors and the most-common category otherwise.
Related to counterfactual explanations (finding the feature changes that would flip or substantially alter a prediction), but simpler. Not an optimal or realistic combined change, just a one-at-a-time sensitivity check.
Everything so far treats each variable’s effect independently. But tree models routinely capture interactions. For example, the young-driver uplift might be amplified by a high bonus level.
Ignoring interactions can make single-variable summaries misleading.
Measures how far the joint effect of two variables departs from additive (Friedman and Popescu 2008). 0 = no interaction, closer to 1 = stronger.
Same idea as the H-statistic, but allocated to individual predictions rather than a single portfolio-level number.
If Age and Bonus interact strongly —
what’s wrong with explaining a policy’s premium using two separate, one-at-a-time SHAP contributions?
If you could only run two of today’s five methods on every model before deployment —
which two, and what would you be willing to miss?
Chapter 6 — Privacy Principles
Bring one example of personal data you’ve disclosed to a company recently, and think about what happened to it next.
