Slides — Chapter 6
~2-hour 15-minute session, five parts, five discussion breaks
| Time | Part |
|---|---|
| 0:00 – 0:20 | What is privacy? |
| 0:20 – 0:40 | Regulatory frameworks |
| 0:40 – 1:00 | Theoretical frameworks: Solove & contextual integrity |
| 1:00 – 1:10 | Break |
| 1:10 – 1:35 | Case study: connected vehicles & GM OnStar |
| 1:35 – 2:05 | Privacy-enhancing techniques |
| 2:05 – 2:15 | DPIA, LLMs, summary |
How would you define each of these, and how do they differ?
| Concept | Meaning |
|---|---|
| Privacy | The individual’s right to control information about themselves |
| Confidentiality | The data holder’s obligation not to disclose without authorisation |
| Security | Technical/organisational measures protecting data |
Privacy is the right. Confidentiality is the obligation it creates. Security is the mechanism that enforces it.
Privacy enables autonomy, making choices without surveillance, profiling, or unwanted disclosure shaping them.
Westin (1967): “the claim of individuals… to determine for themselves when, how, and to what extent information about them is communicated to others.”
Key implication: disclosure for one purpose ≠ consent for another.
Case Study: one disclosure, repurposed
An applicant discloses a chronic condition for underwriting (the process an insurer uses to assess risk and set a price). The same insurer’s marketing division later uses it to exclude them from a wellness campaign. No new disclosure occurred. The same fact was simply reused across purposes never agreed to.
For each category, can you name an example before it’s revealed?
| Category | Examples |
|---|---|
| Personal | Name, address, DOB, vehicle registration |
| Sensitive | Health, genetic, biometric, race, religion, sexual orientation |
| Derived / inferred | Risk scores, behavioural profiles, propensity scores |
Derived data is increasingly contentious. Never directly disclosed, yet can reveal sensitive attributes indirectly.
Think of a piece of derived/inferred data about you that exists somewhere, such as a risk score or a recommendation profile.
Did you ever explicitly disclose the thing it reveals?
Note
Art 22: right not to be subject to a decision based solely on automated processing that significantly affects you, unless consented, contractually necessary, or authorised by law.
13 APPs, applying to orgs >$3M turnover, and specified categories regardless of size (health providers, credit reporting bodies, businesses trading in personal information). Most insurers are covered via the turnover threshold, not an insurance-specific exemption:
Embedded from the start, not bolted on afterward. GDPR Art 25 gives this legal force; the seven principles below are Cavoukian’s earlier framework, related but distinct:
Proactive · Default setting · Embedded in design · Full functionality · End-to-end security · Visibility · User-centric
In practice: build privacy into the pipeline from day one, across underwriting models, hiring algorithms, health scoring alike.
Four groups, by how the violation arises. For each, what kind of harm belongs here?
| Group | Examples |
|---|---|
| Collection | Surveillance, interrogation |
| Processing | Aggregation, identification, secondary use, exclusion |
| Dissemination | Breach of confidentiality, disclosure, exposure |
| Invasion | Intrusion, decisional interference |
Aggregation is often the most consequential for AI. Innocuous data points combine into a profile no single item would reveal.
Privacy isn’t about secrecy. It’s about appropriate information flows (Nissenbaum 2004).
Five parameters define a context’s norms:
Try it: a policyholder shares driving data with their insurer to get a quote. The insurer’s own marketing team later uses that same data to target them with ads for other products. What’s each parameter here?
| Parameter | In this scenario |
|---|---|
| Data subject | The policyholder |
| Sender | The policyholder, via the insurer’s app |
| Recipient | The insurer, for pricing, not marketing |
| Information type | Driving-behaviour data |
| Transmission principle | For a quote, not for internal cross-selling |
Nothing here was disclosed to anyone new. The violation is entirely in the fifth parameter changing.
Would the person who shared this data expect it to flow this way? That’s the test.
Apply the contextual integrity test to a data-sharing practice you’ve encountered.
Would you have expected it to flow that way?
Traditional UBI (dongle, a small plug-in device, or app): customer opts in; data flows customer → insurer; visible, if intrusive.
Connected vehicles invert this entirely:
What happened, 2022–2026
Before the formal analysis on the next two slides:
Using what you know about privacy, confidentiality, and consent, what specifically went wrong here — and at which step could it have been stopped?
| Parameter | Customers understood | What happened |
|---|---|---|
| Sender | Me (via app) | GM |
| Recipient | GM, for coaching | LexisNexis → insurers |
| Info type | Feedback scores | Raw trip-level records |
| Transmission | Not for resale | Sold commercially |
Solove: surveillance (every trip) → aggregation (score) → secondary use + breach of confidentiality (sold) → exclusion (couldn’t see it) → decisional interference (premium hikes, no explanation).
Every parameter in the OnStar case was violated — yet nothing about it was technically a data breach.
What does that tell you about the limits of security-based privacy protection?
EU Data Act (2024): users get real-time access to connected-product data, and can direct it to a third party of their choice, breaking the OEM monopoly.
Australia: no equivalent yet. The ACCC (Australia’s competition regulator) has flagged OEM data access as a competition concern.
US: the FTC (Federal Trade Commission, the US consumer-protection regulator) opened an inquiry into GM/OnStar in 2024, and finalised a 20-year consent order in January 2026 (Federal Trade Commission 2026). Still no general federal law.
Data that looks anonymised often isn’t.
Age + postcode + claims date is often enough to re-identify an insurance policyholder (a person who holds an insurance policy) in “aggregated” data.
Why: Sweeney (1997) re-identified Governor Weld’s hospital record from “anonymised” Massachusetts health data, using only ZIP + birth date + sex, cross-referenced with $20 of voter-roll data (Sweeney 2002).
|\{r' \in D : r'[Q] = r[Q]\}| \geq k
Every record indistinguishable from at least k-1 others on quasi-identifiers (age, postcode, gender, …).
Example: three policyholders sharing the same postcode, age band, and gender form one equivalence class of size 3, satisfying 3-anonymity.
| Technique | Addresses | Requirement |
|---|---|---|
| l-Diversity (Machanavajjhala et al. 2007) | Homogeneity attack: all k records share the same sensitive value | Class must contain at least l distinct sensitive values |
| t-Closeness (Li et al. 2007) | Skewness attack: sensitive values cluster at one end | Class’s sensitive-value distribution within t of the overall distribution |
Example: five policyholders, all fraud-flagged, satisfy k-anonymity but fail l-diversity (a homogeneity attack). If two outcomes are present but 90% fraud vs. 5% portfolio-wide, that class fails t-closeness (a skewness attack) even though it is technically l-diverse.
2018: US Census Bureau researchers reconstructed and re-identified roughly one in six Americans’ 2010 Census records, using only published aggregate tables plus a commercial database (Abowd 2018). Traditional disclosure avoidance, including k-anonymity-style methods, was not safe against this.
\Pr[\mathcal{M}(D) \in S] \leq e^{\varepsilon} \cdot \Pr[\mathcal{M}(D') \in S]
In plain terms: the output barely changes whether or not any one person’s record is included, no matter what an attacker already knows.
\mathcal{M}(D) = f(D) + \text{Lap}\!\left(\frac{\Delta f}{\varepsilon}\right)
\Delta f = the worst-case swing one person’s record could cause in the true answer. \varepsilon = the privacy budget: smaller means more noise and stronger privacy.
Example: true average claim = $8,000, \Delta f = \$2{,}000, \varepsilon = 1. The released value might come out as $7,650 or $9,400, depending on the random draw.
| \varepsilon | Interpretation |
|---|---|
| \leq 1 | Strong privacy, noisy output |
| 1–10 | Moderate, common in practice |
| > 10 | Weak, close to the true value |
Why: the IRS’s Statistics of Income division has published fully synthetic public-use tax files since the early 2000s (Internal Revenue Service, Statistics of Income Division 2018), generated via sequential regression, the same family of methods synthpop uses in Chapter 7.
Fit a generative model to the real data, then sample new, artificial records from it.
Example: a synthetic record for “34-year-old policyholder, mid-density postcode, Bonus level 20, no claims”, generated from learned relationships, not copied from any real policyholder.
Balance: fidelity (useful statistics) vs. privacy (no re-identification). Key risk without DP training: memorisation of rare or extreme records.
| Technique | Guarantee | Preserves records | Modelling |
|---|---|---|---|
| k-Anonymity | No | Yes (generalised) | Reduced utility |
| Differential privacy | Yes (\varepsilon-DP) | No (adds noise) | Statistics & ML |
| Synthetic data | No (unless DP-trained) | No (new records) | Yes, if high fidelity |
Suppose you had to release a dataset for external research.
which of the three PET families would you reach for first, and what would you be trading away?
Example: three insurers jointly training a fraud-detection model each keep their claims data in-house, train locally overnight, and send back only updated weights, never a record, to be averaged into next day’s shared model.
A real production deployment, in banking, not insurance
WeBank (Tencent’s Chinese digital bank) built FATE, an open-source federated learning platform donated to the Linux Foundation in 2019, and used it to build a federated credit-rating model. WeBank has since led China’s national standardisation effort for federated machine learning (Liu et al. 2021).
An analogous insurance use case, e.g. several insurers jointly training a fraud-detection model, remains mostly at the research/pilot stage — no comparably documented production deployment yet.
Why would this matter for insurance specifically? What could several insurers gain by training a shared model this way that none of them could get alone — and what would stop them from just sharing the raw data instead?
Data Protection Impact Assessment (DPIA, called a Privacy Impact Assessment or PIA in Australia), a structured risk review before deploying a new data system.
Required under GDPR Art 35 for high-risk, large-scale, or systematic processing, and most consequential AI applications qualify.
Applied to a new UBI telematics app: identify (GPS pings, braking events, vendor → insurer → possible third parties), assess (location reveals home/work address; the same OnStar pattern could recur), recommend (trip summaries not raw GPS, purpose-limitation clauses, k-anonymity before external sharing), respond (document residual re-identification risk, get sign-off).
LLMs introduce a failure mode tabular PETs weren’t built for. The model itself can memorise and regurgitate training data (Carlini et al. 2021; Nasr et al. 2023).
A real breach, not just a research finding: a March 2023 bug in ChatGPT’s Redis library exposed some users’ chat titles to other users, and the payment details of 1.2% of ChatGPT Plus subscribers active in a nine-hour window (name, email, billing address, partial card number) (OpenAI 2023).
Italy’s Garante temporarily halted ChatGPT’s processing of Italian users’ data (2023), citing this breach alongside a lack of legal basis and transparency (Garante per la protezione dei dati personali 2023).
Two questions, every LLM workflow: could the model leak its training data? And, more immediately, what happens to this prompt’s data at the provider?
Samsung learned the second one the hard way: in three separate 2023 incidents, engineers pasted proprietary source code, meeting notes, and confidential test data into ChatGPT, after which Samsung banned the tool for staff entirely (Ray 2023).
Suppose your organisation started using a third-party LLM API for customer service.
what’s the first privacy question you’d want answered before sending real customer data into a prompt?
Lawfulness · Purpose limitation · Data minimisation · Transparency · Accuracy · Security · Accountability
All map to concrete practice: identify a legal basis, don’t repurpose silently, collect only what’s needed, explain plainly, keep data correct, protect it, document your decisions.
Chapter 7 — Privacy Practice
Bring a laptop with R installed. We’ll apply k-anonymity, differential privacy, and synthetic data generation to real insurance micro-data.
