Slides — Chapter 8
2-hour-20-minute session, the course wrap-up. Five parts, six discussion breaks
| Time | Part |
|---|---|
| 0:00 – 0:20 | Review: what we’ve covered |
| 0:20 – 0:45 | The key trade-offs |
| 0:45 – 0:55 | Break |
| 0:55 – 1:45 | An integrated framework: the lifecycle, plus a capstone exercise |
| 1:45 – 2:10 | Practical governance |
| 2:10 – 2:20 | Course summary |
| Ch | Pillar | Key methods |
|---|---|---|
| 2–3 | Fairness | Criteria (FTU, CPV, DP, CDP…); five model designs; fairness–accuracy trade-off; COMPAS (US recidivism tool) |
| 4–5 | Explainability | PFI, PDP, ALE, SHAP, LIME; global + local; interactions |
| 6–7 | Privacy | Re-identification risk, k-anonymity, differential privacy, synthetic data |
No chapter is self-contained. A real system must address all six ethics principles simultaneously. Managing the interactions is this chapter.
The limits of actuarial fairness
A criterion can be statistically valid and satisfy a named fairness definition, and still be rejected as illegitimate. Which criterion should govern a practice is a contested, political determination, not a purely quantitative one.
Chapter 1: the insurer that removed gender
Engine size and annual mileage remained. Both correlated with gender. Has the insurer achieved fairness?
Now you can answer more rigorously:
Using everything from Chapters 2–7 —
a. How would you now answer “has the insurer achieved fairness”? What’s still missing, even with all these tools?
b. Taking insurance as an example: how would you decide whether a variable like engine size or annual mileage should be used for pricing at all, versus excluded outright, even though each has a genuine, non-discriminatory reason to correlate with risk? What changes when the variable is external or non-traditional, telematics, social media activity, or credit history, rather than a variable the insurer collected for its own actuarial purpose?
1. Fairness vs. accuracy. A fairness constraint reduces usable predictive information
2. Explainability vs. performance. Simpler models are easier to explain but generally less accurate, and post-hoc tools add computation and approximation error
3. Privacy vs. utility. k-anonymity, differential privacy, and synthetic data all reduce information content
Each pillar creates obligations that can conflict with the others.
| Model | Accuracy | Disparity |
|---|---|---|
| M0 | Highest | Highest |
| MU | Moderate reduction | Moderate reduction |
| MCDP / MC | Small reduction | Large reduction |
| MDP | Largest reduction | Near zero |
Accuracy measured by RMSE (smaller error = higher accuracy); “reduction” describes the drop from M0’s accuracy level as fairness constraints tighten.
Not always a trade-off. When errors are unevenly distributed across groups to begin with, a fairness-aware model can sometimes reduce disparity and improve accuracy at once — a genuine win-win (Huang et al. 2024).
A policy decision, not a modelling decision
No statistically correct level of constraint. Depends on regulatory regime, product type, stakeholder values, market structure. Practitioners quantify the trade-off. Decision-makers choose it.
| Technique | Guarantee | Utility cost |
|---|---|---|
| k-Anonymity | None formal | Reduced precision, some suppression |
| Differential privacy (ε=1) | Strong | High noise |
| DP (ε=10) | Weak | Low noise |
| Synthetic (no DP) | None | High fidelity, memorisation risk |
Techniques are layered in practice: k-anonymise → synthesise → check NNDR (nearest-neighbour distance ratio) → DP on released summary stats. Defence-in-depth, not one silver bullet.
Revealing: a SHAP beeswarm showing Postcode dominating premium variation (premium being the price a customer pays for their policy) can flag proxy discrimination the fairness model missed.
Masking: SHAP and LIME can be adversarially manipulated (Slack et al. 2020). A model can look innocuous to an auditor while still discriminating in production.
Explanation tools generate hypotheses and guide criterion choice. They don’t replace formal fairness testing.
If explanation tools can be gamed to look fair —
what would actually make you trust a vendor’s fairness claims?
The ethical AI lifecycle. Source: Huang (2025).
Now we populate each stage with concrete responsible AI obligations.
1. Problem definition. Is AI appropriate? Which criterion applies? Who’s accountable?
2. Data collection. Legal basis? Historical bias? Quasi-identifiers? Minimisation?
3. Model development. Right model design? Proxies addressed? Explainable enough? Trade-off quantified?
A wrong choice at Stage 1 can’t be fixed by better modelling later.
4. Validation. Pre-committed audit? TOST, not plain significance? Stable out-of-time?
5. Deployment. Human review for edge cases? Contestability? DPIA (data protection impact assessment) complete?
6. Monitoring. Fairness metrics recalculated? Data drift tracked? Retraining trigger?
Of the six lifecycle stages —
which do you think most AI failures actually originate in, even though they surface much later?
A motor insurer builds an AI claims-triage system: every incoming claim gets a fraud-risk and complexity score. Low scores → fast-tracked payout within 48h. High scores → routed to manual investigation before any payment.
For each stage, answer using the chapter’s own tools, not general reasoning.
| Stage | Applied question | Ch |
|---|---|---|
| 1. Problem definition | Cost of a false positive vs. a false negative, and who bears each? Reversible? | 1 |
| 2. Data collection | Historical bias in claim-history data? Narrative text minimised? | 6 |
| 3. Model development | Which fairness criterion? Which design keeps the fraud signal? | 2, 3 |
| 4. Validation | What audit protocol proves, not just fails to disprove, fairness? | 2, 3 |
| 5. Deployment | A specific explanation and a real appeal, or a generic notice? | 4, 5 |
| 6. Monitoring | What triggers re-validation after a fraud-pattern shift? | 3, 6 |
Step back from the stage-by-stage view:
There is no answer key. You should leave with a defensible position on each question, and an honest account of which trade-offs you resolved versus only acknowledged.
Mapping between regulations, fairness criteria, and model designs. Source: Xin and Huang (2024).
Regime → criterion → model design → explainability verifies → privacy protects the pipeline.
A model card is a short, standard document summarising a model for anyone auditing, inheriting, or approving it.
Purpose · Data · Fairness criterion · Performance · Fairness audit · Explainability summaries · Privacy assessment · Limitations · Owner and reviewer
Documentation is what makes a governance claim checkable, not just asserted.
Source: National AI Centre AI6 guidance (National AI Centre 2025)
Important
Vendor accountability doesn’t transfer. The regulator looks to the deploying organisation (Ch 1’s ASIC REP 798).
Each practice, grounded in a real regulatory anchor, applied to this chapter’s own claims-triage system (see the capstone case study):
| Practice | Regulatory anchor | Applied here |
|---|---|---|
| 1. Accountability | FAR; ASIC REP 798; APRA CPS 230 | Who owns the decision to reject a claim: the vendor, the assessor, or the insurer? |
| 2. Impact assessment | Anti-discrimination law; Privacy Act ADM (Dec 2026) | Could the triage score systematically disadvantage claimants least able to absorb a delay? |
| 3. Risk management | APRA CPS 230 | Same model, different risk: a delayed motor claim vs. a delayed hardship claim |
| 4. Transparency | Privacy Act ADM (Dec 2026); ASIC’s 11 questions | Can a claimant be told, today, why their payment was delayed? |
| 5. Test & monitor | APRA CPS 230 model risk | Does the model stay calibrated as fraud patterns shift? |
| 6. Human control | FAR “reasonable steps”; APRA CPS 230 continuity | Does the reviewing assessor have real authority to override the flag? |
| Function | Responsibility |
|---|---|
| Risk & compliance | Own the framework, set risk appetite (how much risk the org accepts) |
| Technology & data | Build it into the pipeline, monitor continuously |
| Frontline users | Real oversight, knowing when and how to override |
Tools only produce a trustworthy system inside an organisation that assigns ownership and gives people real authority to intervene.
| Task | Tool |
|---|---|
| Which variables drive predictions | PFI, grouped SHAP |
| Shape of a variable’s effect | PDP, ALE |
| Explain one prediction | SHAP waterfall |
| Detect interactions | H-statistic, SHAP interaction values |
| Enforce demographic parity | MDP, MCDP |
| Audit for compliance | TOST + HC3 |
| Assess re-identification risk | sdcMicro |
| Share data externally | Synthetic data + NNDR |
Responsible AI = making innovation defensible.
Of those five limitations —
which worries you most, for a system you might build or evaluate after this course?
Not a checklist completed once before deployment. A continuous process across the full lifecycle, from problem formulation to retirement.
The practitioner’s responsibility
“The actuarial profession should lead, not follow, in building public trust in AI systems.” — Huang (2025)
The same position (quantitative fluency, accountability for a certified model, a bridge between technical and non-technical stakeholders) applies to credit risk analysts, HR analytics leads, and clinical decision-support leads alike.
Questions, feedback, and continued conversation welcome. This is the end of the syllabus, not the end of the questions.
