Slides — Chapter 2
2-hour session with five parts and five discussion breaks
| Time | Part |
|---|---|
| 0:00 – 0:45 | Step 1: What is fair? |
| 0:45 – 1:15 | Step 2: Designing fair models |
| 1:15 – 1:25 | Break |
| 1:25 – 1:40 | Steps 3 & 4 (brief) |
| 1:40 – 1:55 | Fairness for large language models |
| 1:55 – 2:00 | Summary |
Note
Insurance recurs as a worked example. The framework applies to any consequential automated decision.
Case Study: COMPAS recidivism scoring
ProPublica’s Machine Bias investigation found Black defendants who did not reoffend were flagged “high risk” nearly twice as often as white defendants (Angwin et al. 2016).
The vendor countered that COMPAS was well-calibrated: same risk score, similar reoffense rate across race.
Has fairness been achieved? There’s no single answer. COMPAS is calibrated the same way for both groups, but its errors are not evenly distributed. Which one counts as “fair” depends on which of those two you prioritise.
This is not a coincidence of the COMPAS data. Whenever two groups reoffend, default, or claim at genuinely different rates, satisfying both sides at once becomes mathematically near-impossible, as this chapter makes precise later.
| Lens | What it emphasises |
|---|---|
| Anti-discrimination | Legal obligation to avoid unjustified disparate outcomes |
| Competing ethical notions | Distributive justice, solidarity, pooling vs individualising |
| Consumer / user trust | Promises made in contracts, disclosures, conduct standards |
These lenses don’t always agree. The point is making trade-offs visible before a model is built.
Using an input feature isn’t a purely statistical question:
Match each example below to one of the six feature principles (Frees and Huang 2023):
1. Statistical discrimination — engine size has predictive value, which is necessary but not sufficient on its own
2. Causality — a known causal driver of the insured event
3. Control — sports car ownership is a choice, sensitive attributes are not
4. Socially valuable behaviour — discourages participating in genetic testing research
5. Mutability — changes over time versus staying fixed
6. Past discrimination — skin colour maps onto a protected, historically discriminated-against characteristic, eye colour does not
Direct (disparate treatment) means treated less favourably because a protected characteristic differs.
Indirect (disparate impact) means disproportionately affected because protected status is inferred through a neutral-seeming practice.
Removing the attribute isn’t enough
A hospital algorithm used healthcare cost as a proxy for need, to allocate care-management resources (Obermeyer et al. 2019).
Black patients historically had less access to care → lower costs for the same underlying need → the model systematically under-referred them, without ever using race as an input.
Think of a variable, in a domain you know, that could be a proxy for a protected attribute.
Identifiable, or unidentifiable? Why does that distinction change how you’d fix it?
A few examples across domains
Insurance: postcode as a proxy for race or socioeconomic status (identifiable, a single named column); claims-handling tone in free-text notes as a proxy for perceived credibility (harder to name, embedded in language).
Hiring: employment gaps as a proxy for caregiving status (identifiable); résumé-writing style learned from a demographically skewed training set (unidentifiable, no single feature to point to).
Lending: postcode or shopping patterns as a proxy for race (identifiable, the classic redlining pattern); an opaque credit score itself, if its inputs are undisclosed, can behave as an unidentifiable proxy even though it’s a single number.
Identifiable proxies can be found and removed (MCDP-style orthogonalisation, or dropping the variable). Unidentifiable proxies can’t be, since there’s no single column to act on, which is why testing outcomes directly (MDP, disparate-impact audits) matters more when the proxy is diffuse.
X_P = protected attribute, X_{NP} = other features, Y = actual outcome, \hat{Y} = prediction
| Criterion | Formal condition | Intuition |
|---|---|---|
| FTU (fairness through unawareness) | \hat{Y} = f(X_{NP}) | Remove the protected attribute |
| Fairness through awareness | D(\hat{Y}(x), \hat{Y}(y)) \le d(x, y) | Similar individuals get similar outcomes, via an explicit similarity metric d |
| CPV (controlling for the protected variable) | \hat{Y}_{CPV}(x) = E_{X_P}[f(x_{NP}, X_P)] | Average the model over X_P |
| Criterion | Formal condition | Intuition |
|---|---|---|
| Demographic parity | \hat{Y} \perp X_P | Same outcome distribution, all groups |
| Conditional DP | \hat{Y} \perp X_P \mid X_{L} | Equal predictions within legitimate segments (segments defined by factors accepted as fair grounds for differentiation, covered later) |
| Separation | \hat{Y} \perp X_P \mid Y | Equal error rates across groups |
| Sufficiency | Y \perp X_P \mid \hat{Y} | Equal calibration across groups |
Separation asks whether, conditional on truth Y, errors are the same across groups.
Sufficiency asks whether, conditional on the prediction \hat{Y}, the truth is the same across groups.
Both are desirable, but generally can’t both hold when base rates (the actual outcome rate in each group) differ. (Except special cases, like perfect prediction.)
| Event | Condition | Notion (\Pr\{\text{event} \mid \text{condition}\}) |
|---|---|---|
| \hat{Y}=1 | Y=1 | TPR, recall |
| \hat{Y}=0 | Y=1 | FNR |
| \hat{Y}=1 | Y=0 | FPR |
| \hat{Y}=0 | Y=0 | TNR |
| Y=1 | \hat{Y}=1 | PPV, precision |
| Y=0 | \hat{Y}=0 | NPV |
Same mechanism as the precision-recall trade-off, applied across groups instead of across thresholds. PPV is a function of TPR, FPR, and base rate, so matching TPR/FPR (recall) across groups with different base rates still leaves PPV (precision) unequal.
In practice, you usually can’t satisfy both separation and sufficiency at once, and have to prioritise one.
What would you look at to decide which one to prioritise for a given decision?
What the literature suggests
The relative cost of the two error types. Equalised-odds-style separation was originally motivated by asymmetric error costs (Hardt et al. 2016). If a false positive (e.g. wrongly denying parole) is a much worse harm than a false negative, prioritise equalising that error rate across groups rather than full calibration.
Whether a human interprets the score downstream, or a fixed threshold gates the decision automatically. A score a human consults (e.g. a risk score shown to a judge) needs to mean the same thing for every group, which is what calibration guarantees, so sufficiency matters more. A score that triggers an automatic cutoff is better served by equal error rates at that cutoff, which is what separation guarantees.
What the original COMPAS debate actually turned on. ProPublica’s separation-style critique focused specifically on FPR, the rate at which defendants who did not reoffend were nonetheless flagged high-risk, a punitive error. Northpointe’s sufficiency-style defence focused on calibration, arguing the score meant the same thing for every defendant. Neither side was wrong about the criterion they picked; they disagreed about which error mattered more (Kleinberg et al. 2017; Chouldechova 2017).
Individual and group fairness together are often impossible to achieve simultaneously.
This is exactly the COMPAS tension from the start of this chapter. ProPublica’s critique was separation-style, the vendor’s defence was sufficiency-style. Both were valid, and neither could hold at once given the different reoffense base rates across race.
\text{PPV}_g = \frac{p_g(1 - \text{FNR}_g)}{p_g(1-\text{FNR}_g) + (1-p_g)\,\text{FPR}_g}
Equally reliable “high risk” flags in both groups, but different reoffense rates, mean different false positive rates. That’s algebra, not evidence of biased predictions.
100 people per group, same FNR (10%) and FPR (20%) in both, so separation holds by construction.
Group A
| Predicted \hat{Y}=1 | Predicted \hat{Y}=0 | |
|---|---|---|
| Actual Y=1 | TP = 27 | FN = 3 |
| Actual Y=0 | FP = 14 | TN = 56 |
Group B
| Predicted \hat{Y}=1 | Predicted \hat{Y}=0 | |
|---|---|---|
| Actual Y=1 | TP = 54 | FN = 6 |
| Actual Y=0 | FP = 8 | TN = 32 |
\text{Flag rate}_g = \frac{TP+FP}{TP+FP+FN+TN} \quad \text{TPR}_g = \frac{TP}{TP+FN} \quad \text{FPR}_g = \frac{FP}{FP+TN} \quad \text{PPV}_g = \frac{TP}{TP+FP}
| Group A (p_0 = 30%) | Group B (p_1 = 60%) | |
|---|---|---|
| Reoffend | 30 | 60 |
| Do not reoffend | 70 | 40 |
| Correctly flagged | 27 | 54 |
| Wrongly flagged | 14 | 8 |
| Total flagged high-risk | 41 | 62 |
| Flag rate (demographic parity) | 41% | 62% |
| \text{TPR}_g,\ \text{FPR}_g (separation) | 90%, 20% | 90%, 20% |
| PPV_g (sufficiency) | 27/41 ≈ 66% | 54/62 ≈ 87% |
Same error rates, but flag rate differs (41% vs 62%) and PPV jumps from 66% to 87%, purely from the base-rate gap. Matching flag rate or PPV instead would each require different FPR/FNR in each group, violating separation.
Matching criterion to regime:
| Position | Emphasis | Logic |
|---|---|---|
| Social good | Solidarity, access, universal service | Subsidising solidarity |
| Economic commodity | Efficiency, adverse selection | Chance solidarity |
Actuarial fairness and the EU gender ban
Actuarial fairness means charging each insured person only for their own risk. In the EU’s 2012 ban on gender-based motor pricing, the ECJ held that equal treatment took precedence over gender being a statistically valid predictor (European Court of Justice 2011). The ban didn’t remove the underlying risk difference — it redistributed who pays for it.
Separation and sufficiency can’t both hold when base rates differ.
If you had to pick one, for a domain you care about — which, and what are you trading away?
Two examples, using the framework from Step 1
Criminal-justice risk scoring (like COMPAS): a false positive (wrongly flagged high-risk) is a punitive harm affecting someone’s liberty. This pushes toward separation, specifically equalising FPR, even at the cost of calibration.
Medical risk scoring (e.g. a score a clinician uses to prioritise screening): the score is interpreted by a human downstream, and a “70% risk” that means different things by group directly undermines clinical trust. This pushes toward sufficiency: calibration matters more than matching error rates exactly.
Whichever you pick, name what you’re giving up. That’s the actual skill this chapter is teaching: not “solving” the trade-off, but making the choice, and its cost, explicit and defensible.
| Stage | What happens | Models |
|---|---|---|
| Pre-processing | Adjust inputs before training | MDP, MCDP |
| In-processing | Build fairness into the objective | Constrained optimisation |
| Post-processing | Adjust outputs after training | MC |
The right stage depends on the criterion chosen in Step 1, not what’s easiest to implement. MDP/MCDP/MC are this course’s named examples per stage. The same criterion could in principle be reached by a different route (Barocas et al. 2023, ch. 3).
| Model | Criterion | Approach |
|---|---|---|
| M0 | Baseline | Uses X_P and X_{NP} |
| MU | FTU | Protected attribute removed |
| MDP | Demographic parity | Debias all X_{NP} |
| MCDP | Conditional DP | Debias only non-legitimate predictors |
| MC | CPV | Fit M0, then average over X_P at scoring |
Step 1’s criteria came from a classification example (COMPAS). Insurance pricing is usually regression — Xin and Huang (2024) formalise these five designs for continuous \hat{Y} (claim frequency, severity, pure premium). DP extends directly (equal average \hat{Y} across groups); separation/sufficiency’s TPR/FPR/PPV form needs a continuous-outcome analogue instead.
Notation from Xin and Huang (2024): X_{NP}^* is debiased X_{NP}, X_{NP_{legit}} / X_{NP_{not}} split X_{NP} for MCDP.
\hat{Y}_{M0} = f_{M0}(X_{NP}, X_P) \hat{Y}_{MU} = f_{MU}(X_{NP}) \hat{Y}_{MDP} = f_{MDP}(X_{NP}^*) \hat{Y}_{MCDP} = f_{MCDP}(X_{NP_{not}}^*,\, X_{NP_{legit}}) \hat{Y}_{MC} = \frac{1}{N}\sum_{j=1}^N \hat{f}_{M0}(X_{NP},\, X_P = x_{P_j})
Legitimate means grounds a regulator/institution accepts for differentiated treatment.
MCDP is built for this. Differences allowed only through legitimate variables.
No restriction → M0 · Prohibit X_P → MU · Prohibit proxies too → MU* · Equal average outcomes → MDP · Full pooling → community rating (same price for everyone in a group regardless of individual risk)
Note
Not every regulation maps to one criterion cleanly, common under principles-based regimes: China’s NFRA guidance (National Financial Regulatory Administration 2026a, 2026b; Sina Finance 2026), Australia’s AHRC (Australian Human Rights Commission and Actuaries Institute 2022), and Singapore’s MAS FEAT/Veritas (Monetary Authority of Singapore 2018, 2019) all avoid prescribing a single fairness criterion. The institution has to choose, using Step 1’s judgement.
French motor insurance evidence (Xin and Huang 2024):
The cost of fairness is smaller than often assumed, but not zero. Strict demographic parity trades off against something like adverse selection. Quantifiable, but still a business decision.
Root mean square error (x, lower is more accurate) vs. disparate impact ratio (y, 1 is perfect parity), five model designs, GLM and XGBoost, French motor insurance data (Xin and Huang 2024). Dotted lines mark the four-fifths rule’s 0.8–1.25 band.
The insurance case study found only a modest accuracy cost for fairness.
Do you think that generalises, or is it domain-specific? What would make the trade-off worse?
Evidence points both ways
Supporting generalisation: Rodolfa et al. (2021) find the same modest-cost pattern across education, mental health, criminal justice, and housing safety; Hardt et al. (2016) find it in credit scoring too. Several independent domains, same qualitative result.
Against blanket generalisation: the cost depends on how much of a model’s predictive power is actually carried by the protected attribute or its correlates. A large base-rate gap between groups (like COMPAS’s) means removing or orthogonalising that signal costs more. A weak correlation means fairness is closer to free.
What would make it worse: smaller training data (less signal to substitute for the removed correlation), a narrower feature set, or a genuinely strong relationship between the protected attribute and the outcome that the model can’t recover elsewhere.
A model meeting its criterion can still produce an unfair outcome once its predictions hit a real process.
A downstream step, re-ranking, a discretionary override, a separate margin or eligibility adjustment, can undo fairness achieved at the model stage if it correlates with a protected group.
Audit asks whether we can demonstrate compliance. Test asks whether we can find a problem.
A credible audit:
No fixed input schema. No column called “Gender,” prediction is open-ended text.
Group-fairness criteria compare a well-defined \hat{Y} across groups.
Open-ended generation has no single \hat{Y}. Evaluation needs templated probes (swap a name/pronoun, compare sentiment/toxicity).
Dozens of proposed metrics, little consensus (Gallegos et al. 2024). This is considerably less mature than the tabular setting this course focuses on.
A conventional bias audit compares outcomes for Male vs. Female claimants only, since that’s how gender has traditionally been recorded.
Real intake forms increasingly offer Non-binary and “prefer not to say” options too. If an LLM passed a Male/Female audit with no disparity, would you trust it to be unbiased for the categories the audit never tested? What could go wrong that a binary audit would miss entirely?
Huang et al. (2026)’s counterfactual audit:
Under the conventional Male/Female, full-information audit: none of the six models shows a significant disparity.
Once Non-binary and Not Specified are included, three models (GPT-5, GPT-4o, Gemini 2.5 Pro) show significant disparities. GPT-4o is the most biased: $128–$288 difference in predicted claim amount, favouring Not Specified claimants when the name is present, but the disparity flips to favour Non-binary over Male claimants once the name is removed.
Three models (Gemini 3 Flash, Claude 4 Sonnet, Claude 4.5 Sonnet) show no significant disparity under any of the 16 conditions tested.
A binary-only audit would have certified all six models as fair.
GPT-4o shows the bias; three other models (Gemini 3 Flash, Claude 4 Sonnet, Claude 4.5 Sonnet) don’t.
Before the next slide: as the insurer’s AI governance lead, what would you actually change? List at least two concrete options.
Options grounded in the paper and this chapter
Change the prompt, not just the model. GPT-4o’s bias pattern flips depending on whether the claimant’s name is present. Stripping the name (and other demographic-adjacent free text) from the prompt before the claim reaches the model is the LLM analogue of FTU, removing an input the model doesn’t need to do its job.
Switch models, if procurement allows it. Three of the six models tested show no significant disparity under any condition. Model choice is itself a fairness lever, not just a technical one, provided switching doesn’t trade away needed capability.
Audit is not a one-off gate. GPT-5 fixed most of GPT-4o’s disparities, but that isn’t guaranteed for every update. Every model version change should trigger a new, multiplicity-aware audit (multiple gender categories, multiple prompt configurations), not a one-time sign-off at procurement.
Notice what’s missing: there’s no MCDP-style pre-processing step. The LLM’s input isn’t a fixed feature vector X_{NP} you can orthogonalise, so the five-model toolkit from Step 2 doesn’t transfer directly to open-ended generation.
Real-estate platforms are starting to use GenAI to suggest listing prices.
Historically, human-set housing prices have been higher in white-dominant neighbourhoods than in comparable minority-dominant ones. If an LLM is trained on this same historical data, would you expect its prices to reproduce, widen, or narrow that gap? Why?
Tanlamai et al. (2026)’s comparison:
Human-set prices: houses in minority-dominant neighbourhoods are priced 22.9% lower than comparable houses in white-dominant neighbourhoods.
GPT-4o prices: the same gap shrinks to 14.2%, an 8.7 percentage-point reduction, not an amplification.
GenAI prices are lower than human prices overall (13.5% lower on average), but the reduction is larger for minority-dominant neighbourhoods, which is what narrows the gap.
The opposite direction from the insurance case study. Mechanism checks suggest it isn’t just “regression to the mean”: it’s driven by what the training data (much of it public web content) encodes about pricing generally, not a simple averaging effect.
GPT-4o narrows the racial pricing gap (22.9% → 14.2%), but doesn’t eliminate it.
Before the next slide: would you deploy GPT-4o’s prices as-is, or add a correction step? What would that correction look like, using this chapter’s own toolkit?
Options grounded in this chapter
Treat the residual 14.2% gap as an unfinished Step 2 problem, not a solved one. GPT-4o’s price is a fitted output over X_{NP} (location, size, school rating, and so on), so it can be treated like any other predicted price for this chapter’s toolkit.
MCDP-style correction: identify which of GPT-4o’s inputs are legitimate risk factors (e.g. flood risk, school rating) versus non-legitimate proxies (e.g. a postcode standing in for neighbourhood racial composition), and orthogonalise only the latter against race before the price is finalised.
MC-style correction: average GPT-4o’s price over the neighbourhood’s racial composition at scoring time, the same post-processing idea used in the insurance-pricing worked example earlier in this chapter.
Neither is free. Both require the platform to first classify legitimate vs. non-legitimate inputs (a judgement call) or to have race/neighbourhood-composition data available at scoring time, exactly the same practical costs already discussed for MCDP and MC.
If you were auditing an LLM drafting claim-denial letters —
what would “fair” even mean here, with no single \hat{Y} to compare?
No single Ŷ, but not undefined either
Templated probes, not direct criteria. Hold the claim details fixed, swap only the claimant’s name or stated demographic, and compare the generated text across the swap: tone (via sentiment analysis), specificity of the stated denial reason, or presence of an appeals-process explanation.
Decompose “fair” into measurable proxies. Even without a single outcome number, ask: is the rate of denial letters (a binary extractable from the text) equal across groups, a demographic-parity-style question? Is the letter’s reading level or actionability, a proxy for how easy it is to contest the decision, equal across groups?
This is exactly Gallegos et al. (2024)’s point: dozens of proposed metrics exist because no single one dominates. The audit design has to be built for the specific harm, not borrowed wholesale from the tabular setting.
| Step | Question | Depth here |
|---|---|---|
| 1. Define fairness | Which criterion applies? | Full |
| 2. Design a fair model | Which design enforces it? | Full |
| 3. Assess impact | Who gains/loses in practice? | Brief |
| 4. Audit the system | Can we demonstrate compliance? | Brief |
Of the four steps — define, design, assess impact, audit —
which do you think organisations skip most often in practice, and why?
Most commonly nominated: Step 4 (audit)
Steps 1 and 2 are where the interesting technical work happens, and where a data science team’s incentives naturally point, so they tend to get done. Step 3 (impact assessment) is at least partially forced by growing algorithmic-impact-assessment regulation. Step 4 (ongoing audit) is the one most often skipped, because it has no natural endpoint (a model “passes” once at deployment, but nothing forces a recheck later), and because, as the LLM case study earlier in this chapter showed, a model that passes today’s audit can fail tomorrow’s after a routine version update.
Full list, including fair.feihuang.org’s interactive tools, in the lecture notes.
Chapter 3 — Fairness Practice
Bring a laptop with R installed. We’ll fit these five model designs on real data.
