Slides — Chapter 3
2-hour session, hands-on, five parts, five discussion breaks
| Time | Part |
|---|---|
| 0:00 – 0:15 | From principle to practice: the workflow |
| 0:15 – 0:40 | Fitting the models (GLM + XGBoost) |
| 0:40 – 1:05 | Results: fairness–accuracy trade-off |
| 1:05 – 1:15 | Break |
| 1:15 – 1:35 | Premium redistribution |
| 1:35 – 2:00 | Case study: COMPAS classification |
Same sequence, regardless of domain.
| Model | Criterion | Approach |
|---|---|---|
| M0 | Baseline | Full model, includes Gender |
| MU | FTU (Fairness Through Unawareness) | Gender removed |
| MDP | Demographic parity | All predictors debiased |
| MCDP | Conditional DP | Only InsuranceScore debiased |
| MC | CPV (Controlling for Protected Variable) | Full model, average over Gender at scoring |
Each fitted with GLM (Poisson frequency + Gamma severity) and XGBoost.
If your organisation had to deploy one of these five designs tomorrow —
which would you pick, and why?
For MDP and MCDP, adjust continuous predictor distributions so they no longer depend on Gender, preserving overall shape.
\text{Pure Premium} = \text{Claim Frequency} \times \text{Claim Severity}
Age entered as a flexible function (\text{Age}, \log(\text{Age}), \text{Age}^2, \text{Age}^3, \text{Age}^4).
Disparate Impact Ratio (DIR): \dfrac{\mathbb{E}(\hat{Y} \mid X_P = b)}{\mathbb{E}(\hat{Y} \mid X_P = a)}
RMSE (accuracy, lower = better)
The cost of fairness here is real, but small.
Removing Gender doesn’t just shrink the gender gap in MU. It reverses which gender pays more.
What does that tell you about unawareness alone?
Compare each fair model to MU (the industry-standard benchmark):
\text{Relative Premium Difference}_{i} = \frac{\bar{Y}_{\text{model},i} - \bar{Y}_{\text{MU},i}}{\bar{Y}_{\text{MU},i}}
Positive = fair model charges more than MU · Negative = charges less
The solidarity principle: stricter criteria imply more cross-subsidy between groups.
MDP produces the largest transfer between groups.
Who would push back hardest against deploying MDP, and what would they say?
A US county is deciding whether to renew its contract for a pretrial risk-assessment tool that helps judges set bail and release conditions.
Everything above is regression (continuous cost). This decision is binary: release/detain, flag/clear.
The M0/MU framework applies unchanged. Only the model type and metrics change.
m0 <- glm(Two_yr_Recidivism ~ Number_of_Priors + Age_Above_FourtyFive +
Age_Below_TwentyFive + Female + Misdemeanor + ethnicity,
data = compas_data, family = binomial)
mu <- glm(Two_yr_Recidivism ~ Number_of_Priors + Age_Above_FourtyFive +
Age_Below_TwentyFive + Female + Misdemeanor,
data = compas_data, family = binomial)M0 includes ethnicity directly. MU removes it, mirroring the real COMPAS tool.
| Criterion | M0 disparity ratio | MU disparity ratio |
|---|---|---|
| Demographic parity | 2.11 | 1.94 |
| Equal opportunity / TPR gap | 1.75 | 1.67 |
| Predictive rate parity / precision gap | 1.11 | 1.15 |
Unawareness barely moves anything. MU costs almost no accuracy (66.5% vs 66.8%). Every disparity ratio survives nearly unchanged.
Important
Demographic parity ignores the true outcome. It fully reflects the base-rate gap (39.1% vs 52.3%) plus any model unfairness. Precision conditions on the prediction, partly absorbing that same gap.
This is Chapter 2’s impossibility result, on real data: separation and sufficiency can’t both hold when base rates differ. Choosing a criterion is choosing how much of the base-rate difference counts as “unfairness” vs. “signal.”
The three checks disagree sharply on how bad COMPAS’s disparity is.
If you were a judge deciding whether to use this tool, which check would you trust most, and why?
roc_pivot()This reject-option correction relabels predictions within theta of the cutoff, giving the disadvantaged group’s borderline cases the benefit of the doubt.
| theta | Accuracy | Demographic parity | TPR gap | Precision gap |
|---|---|---|---|---|
| 0 (MU) | 66.5% | 1.94 | 1.67 | 1.15 |
| 0.05 | 65.7% | 1.13 | 1.11 | 1.31 |
| 0.10 | 62.8% | 0.63 | 0.69 | 1.45 |
| 0.20 | 56.8% | 0.23 | 0.30 | 1.73 |
Important
Pushing demographic parity and the TPR gap toward 1 does not leave precision alone — it actively worsens it (1.15 → 1.73). Past theta ≈ 0.10, the correction overshoots and reverses the disparity, at a steep accuracy cost.
No theta satisfies all three checks at once, because none exists while base rates differ. Post-processing lets you choose where on the curve to sit. It doesn’t let you escape the curve.
Post-processing dials fairness up on one criterion, always at the cost of another.
If you had to pick a theta for a real deployed system, how would you decide where to stop?
Chapter 4 — Explainability Principles
Bring one prediction from today’s models you’d want explained to a policyholder.
