Privacy Principles

Slides — Chapter 6

Fei Huang, UNSW Sydney

Today’s roadmap

~2-hour 15-minute session, five parts, five discussion breaks

Time Part
0:00 – 0:20 What is privacy?
0:20 – 0:40 Regulatory frameworks
0:40 – 1:00 Theoretical frameworks: Solove & contextual integrity
1:00 – 1:10 Break
1:10 – 1:35 Case study: connected vehicles & GM OnStar
1:35 – 2:05 Privacy-enhancing techniques
2:05 – 2:15 DPIA, LLMs, summary

Learning objectives

  • Define privacy, and distinguish it from confidentiality and security
  • Explain core principles from major regulatory frameworks
  • Apply Solove’s taxonomy and Nissenbaum’s contextual integrity
  • Identify privacy tensions from connected devices, using telematics (data collected from devices or apps that track how, when, and where a vehicle is driven) as a case study

Part 1 — What is privacy?

Why it matters beyond compliance

Privacy enables autonomy, making choices without surveillance, profiling, or unwanted disclosure shaping them.

  • Discrimination: health/genetic data pricing cover, behavioural data screening candidates
  • Chilling effects: patients avoiding care, applicants concealing history
  • Power asymmetry: organisations hold vast data, and individuals see little of how it’s used
  • Reputational harm: disclosure of sensitive records

Defining privacy

Westin (1967): “the claim of individuals… to determine for themselves when, how, and to what extent information about them is communicated to others.”

Key implication: disclosure for one purpose ≠ consent for another.

Case Study: one disclosure, repurposed

An applicant discloses a chronic condition for underwriting (the process an insurer uses to assess risk and set a price). The same insurer’s marketing division later uses it to exclude them from a wellness campaign. No new disclosure occurred. The same fact was simply reused across purposes never agreed to.

Sensitive vs personal data

For each category, can you name an example before it’s revealed?

Category Examples
Personal Name, address, DOB, vehicle registration
Sensitive Health, genetic, biometric, race, religion, sexual orientation
Derived / inferred Risk scores, behavioural profiles, propensity scores

Derived data is increasingly contentious. Never directly disclosed, yet can reveal sensitive attributes indirectly.

💬 Discuss (3 min)

Think of a piece of derived/inferred data about you that exists somewhere, such as a risk score or a recommendation profile.

Did you ever explicitly disclose the thing it reveals?

Part 2 — Regulatory frameworks

GDPR: six core principles (Art 5)

  • Lawfulness, fairness, transparency
  • Purpose limitation
  • Data minimisation
  • Accuracy
  • Storage limitation
  • Integrity and confidentiality

Note

Art 22: right not to be subject to a decision based solely on automated processing that significantly affects you, unless consented, contractually necessary, or authorised by law.

GDPR: key rights

  • Access (Art 15), request all data held
  • Rectification (Art 16), correct inaccuracies
  • Erasure (Art 17), “right to be forgotten”
  • Portability (Art 20), machine-readable export
  • Explanation of automated decision-making (Arts 13-15 and 22), understand the logic of automated decisions

Australian Privacy Principles

13 APPs, applying to orgs >$3M turnover, and specified categories regardless of size (health providers, credit reporting bodies, businesses trading in personal information). Most insurers are covered via the turnover threshold, not an insurance-specific exemption:

  • APP 1: open, transparent policy · APP 3: collect only what’s necessary
  • APP 5: notify at collection · APP 6: use only for primary/related purpose
  • APP 8: cross-border protection equivalence · APP 11: security
  • APP 12–13: access and correction rights

China, Singapore, and the US

  • China’s PIPL (2021): China’s GDPR-equivalent. Art 24 requires transparent, fair automated decisions, and grants a right to explanation and to refuse solely-automated decisions (National People’s Congress (China) 2021)
  • Singapore’s PDPA (2012, amended 2020): 10 obligations including consent, access/correction, and data breach notification (3-day notice to the PDPC) (Personal Data Protection Commission Singapore 2012); no standalone explanation right, addressed instead through Chapter 1’s voluntary AI-governance frameworks
  • US: no federal privacy law. California’s CCPA/CPRA gives know/delete/correct/opt-out rights (California Privacy Protection Agency 2023); insurers aren’t automatically exempt where existing sector law (IIPPA, GLBA) doesn’t already cover a practice

Privacy by design

Embedded from the start, not bolted on afterward. GDPR Art 25 gives this legal force; the seven principles below are Cavoukian’s earlier framework, related but distinct:

Proactive · Default setting · Embedded in design · Full functionality · End-to-end security · Visibility · User-centric

In practice: build privacy into the pipeline from day one, across underwriting models, hiring algorithms, health scoring alike.

Part 3 — Theoretical frameworks

Solove’s taxonomy

Four groups, by how the violation arises. For each, what kind of harm belongs here?

Group Examples
Collection Surveillance, interrogation
Processing Aggregation, identification, secondary use, exclusion
Dissemination Breach of confidentiality, disclosure, exposure
Invasion Intrusion, decisional interference

Aggregation is often the most consequential for AI. Innocuous data points combine into a profile no single item would reveal.

Contextual integrity

Privacy isn’t about secrecy. It’s about appropriate information flows (Nissenbaum 2004).

Five parameters define a context’s norms:

  1. Data subject, who it’s about
  2. Sender, who provides it
  3. Recipient, who receives it
  4. Information type
  5. Transmission principle, under what conditions

Try it: a policyholder shares driving data with their insurer to get a quote. The insurer’s own marketing team later uses that same data to target them with ads for other products. What’s each parameter here?

Parameter In this scenario
Data subject The policyholder
Sender The policyholder, via the insurer’s app
Recipient The insurer, for pricing, not marketing
Information type Driving-behaviour data
Transmission principle For a quote, not for internal cross-selling

Nothing here was disclosed to anyone new. The violation is entirely in the fifth parameter changing.

Would the person who shared this data expect it to flow this way? That’s the test.

💬 Discuss (4 min)

Apply the contextual integrity test to a data-sharing practice you’ve encountered.

Would you have expected it to flow that way?

Break — 10 min

Part 4 — Case study: connected vehicles

From UBI telematics to connected vehicles

Traditional UBI (dongle, a small plug-in device, or app): customer opts in; data flows customer → insurer; visible, if intrusive.

Connected vehicles invert this entirely:

  • The vehicle is a networked sensor, and no opt-in is needed to generate data
  • GPS every few seconds, cabin mic, seat occupancy, charging patterns
  • Data flows: vehicle → OEM cloud (OEM = original equipment manufacturer, i.e. the vehicle maker) → brokers, insurers, advertisers
  • The driver is no longer the sender — the manufacturer is

Case Study: GM OnStar

What happened, 2022–2026

  1. GM enrolled millions of drivers in OnStar Smart Driver, described to customers as “personalised feedback” on their own driving
  2. Customers were enrolled by checking a box on the vehicle’s infotainment screen; data-sharing terms were buried in the fine print
  3. GM sold trip-level data, including hard-braking, acceleration, and speeding events, to the data broker LexisNexis Risk Solutions
  4. LexisNexis packaged this into driving behaviour scores and sold them on to insurers, including Allstate, Liberty Mutual, and State Farm
  5. Customers received premium increases based on scores they did not know existed and had never agreed to share with an insurer
  1. Following a New York Times / The Markup investigation, GM shut the programme down, March 2024
  2. The FTC (Federal Trade Commission, the US consumer-protection regulator) proposed a settlement in January 2025 and finalised a 20-year consent order in January 2026: affirmative express consent now required before collecting or sharing this data, plus a 5-year ban on sharing geolocation/driving data with consumer reporting agencies (Federal Trade Commission 2026)

💬 Discuss (4 min)

Before the formal analysis on the next two slides:

Using what you know about privacy, confidentiality, and consent, what specifically went wrong here — and at which step could it have been stopped?

Contextual integrity: every parameter violated

Parameter Customers understood What happened
Sender Me (via app) GM
Recipient GM, for coaching LexisNexis → insurers
Info type Feedback scores Raw trip-level records
Transmission Not for resale Sold commercially

Solove: surveillance (every trip) → aggregation (score) → secondary use + breach of confidentiality (sold) → exclusion (couldn’t see it) → decisional interference (premium hikes, no explanation).

💬 Discuss (4 min)

Every parameter in the OnStar case was violated — yet nothing about it was technically a data breach.

What does that tell you about the limits of security-based privacy protection?

Regulatory response

EU Data Act (2024): users get real-time access to connected-product data, and can direct it to a third party of their choice, breaking the OEM monopoly.

Australia: no equivalent yet. The ACCC (Australia’s competition regulator) has flagged OEM data access as a competition concern.

US: the FTC (Federal Trade Commission, the US consumer-protection regulator) opened an inquiry into GM/OnStar in 2024, and finalised a 20-year consent order in January 2026 (Federal Trade Commission 2026). Still no general federal law.

Part 5 — Privacy-Enhancing Techniques

Re-identification risk

Data that looks anonymised often isn’t.

  • Sweeney (2000) (Sweeney 2000): 87% of Americans uniquely identified by DOB + gender + 5-digit ZIP alone
  • Narayanan and Shmatikov (2008) (Narayanan and Shmatikov 2008): re-identified “anonymous” Netflix ratings via public IMDb reviews

Age + postcode + claims date is often enough to re-identify an insurance policyholder (a person who holds an insurance policy) in “aggregated” data.

k-Anonymity

Why: Sweeney (1997) re-identified Governor Weld’s hospital record from “anonymised” Massachusetts health data, using only ZIP + birth date + sex, cross-referenced with $20 of voter-roll data (Sweeney 2002).

|\{r' \in D : r'[Q] = r[Q]\}| \geq k

Every record indistinguishable from at least k-1 others on quasi-identifiers (age, postcode, gender, …).

Example: three policyholders sharing the same postcode, age band, and gender form one equivalence class of size 3, satisfying 3-anonymity.

Extensions: l-diversity and t-closeness

Technique Addresses Requirement
l-Diversity (Machanavajjhala et al. 2007) Homogeneity attack: all k records share the same sensitive value Class must contain at least l distinct sensitive values
t-Closeness (Li et al. 2007) Skewness attack: sensitive values cluster at one end Class’s sensitive-value distribution within t of the overall distribution

Example: five policyholders, all fraud-flagged, satisfy k-anonymity but fail l-diversity (a homogeneity attack). If two outcomes are present but 90% fraud vs. 5% portfolio-wide, that class fails t-closeness (a skewness attack) even though it is technically l-diverse.

Differential Privacy: why it was needed

2018: US Census Bureau researchers reconstructed and re-identified roughly one in six Americans’ 2010 Census records, using only published aggregate tables plus a commercial database (Abowd 2018). Traditional disclosure avoidance, including k-anonymity-style methods, was not safe against this.

\Pr[\mathcal{M}(D) \in S] \leq e^{\varepsilon} \cdot \Pr[\mathcal{M}(D') \in S]

In plain terms: the output barely changes whether or not any one person’s record is included, no matter what an attacker already knows.

The Laplace mechanism

\mathcal{M}(D) = f(D) + \text{Lap}\!\left(\frac{\Delta f}{\varepsilon}\right)

\Delta f = the worst-case swing one person’s record could cause in the true answer. \varepsilon = the privacy budget: smaller means more noise and stronger privacy.

Example: true average claim = $8,000, \Delta f = \$2{,}000, \varepsilon = 1. The released value might come out as $7,650 or $9,400, depending on the random draw.

\varepsilon Interpretation
\leq 1 Strong privacy, noisy output
1–10 Moderate, common in practice
> 10 Weak, close to the true value

Synthetic Data

Why: the IRS’s Statistics of Income division has published fully synthetic public-use tax files since the early 2000s (Internal Revenue Service, Statistics of Income Division 2018), generated via sequential regression, the same family of methods synthpop uses in Chapter 7.

Fit a generative model to the real data, then sample new, artificial records from it.

Example: a synthetic record for “34-year-old policyholder, mid-density postcode, Bonus level 20, no claims”, generated from learned relationships, not copied from any real policyholder.

Balance: fidelity (useful statistics) vs. privacy (no re-identification). Key risk without DP training: memorisation of rare or extreme records.

Three families of techniques

Technique Guarantee Preserves records Modelling
k-Anonymity No Yes (generalised) Reduced utility
Differential privacy Yes (\varepsilon-DP) No (adds noise) Statistics & ML
Synthetic data No (unless DP-trained) No (new records) Yes, if high fidelity

💬 Discuss (3 min)

Suppose you had to release a dataset for external research.

which of the three PET families would you reach for first, and what would you be trading away?

Federated learning: a different question

  • The three techniques above all assume data is centralised, even briefly, before protection is applied
  • Federated learning: the raw data never leaves the device or organisation that holds it
  • Central server sends a shared model → each participant trains it locally → only model updates (weights, not records) are sent back and aggregated

Example: three insurers jointly training a fraud-detection model each keep their claims data in-house, train locally overnight, and send back only updated weights, never a record, to be averaged into next day’s shared model.

Case Study: WeBank’s FATE

A real production deployment, in banking, not insurance

WeBank (Tencent’s Chinese digital bank) built FATE, an open-source federated learning platform donated to the Linux Foundation in 2019, and used it to build a federated credit-rating model. WeBank has since led China’s national standardisation effort for federated machine learning (Liu et al. 2021).

An analogous insurance use case, e.g. several insurers jointly training a fraud-detection model, remains mostly at the research/pilot stage — no comparably documented production deployment yet.

💬 Discuss (3 min)

Why would this matter for insurance specifically? What could several insurers gain by training a shared model this way that none of them could get alone — and what would stop them from just sharing the raw data instead?

Wrap-up

DPIA: the process

Data Protection Impact Assessment (DPIA, called a Privacy Impact Assessment or PIA in Australia), a structured risk review before deploying a new data system.

  1. Identify, map data flows
  2. Assess, risks against obligations
  3. Recommend, controls
  4. Respond, implement, document residual risk

Required under GDPR Art 35 for high-risk, large-scale, or systematic processing, and most consequential AI applications qualify.

Applied to a new UBI telematics app: identify (GPS pings, braking events, vendor → insurer → possible third parties), assess (location reveals home/work address; the same OnStar pattern could recur), recommend (trip summaries not raw GPS, purpose-limitation clauses, k-anonymity before external sharing), respond (document residual re-identification risk, get sign-off).

Privacy for large language models

LLMs introduce a failure mode tabular PETs weren’t built for. The model itself can memorise and regurgitate training data (Carlini et al. 2021; Nasr et al. 2023).

A real breach, not just a research finding: a March 2023 bug in ChatGPT’s Redis library exposed some users’ chat titles to other users, and the payment details of 1.2% of ChatGPT Plus subscribers active in a nine-hour window (name, email, billing address, partial card number) (OpenAI 2023).

Italy’s Garante temporarily halted ChatGPT’s processing of Italian users’ data (2023), citing this breach alongside a lack of legal basis and transparency (Garante per la protezione dei dati personali 2023).

Two questions, every LLM workflow: could the model leak its training data? And, more immediately, what happens to this prompt’s data at the provider?

Samsung learned the second one the hard way: in three separate 2023 incidents, engineers pasted proprietary source code, meeting notes, and confidential test data into ChatGPT, after which Samsung banned the tool for staff entirely (Ray 2023).

💬 Discuss (5 min) — wrap-up

Suppose your organisation started using a third-party LLM API for customer service.

what’s the first privacy question you’d want answered before sending real customer data into a prompt?

Summary: core principles

Lawfulness · Purpose limitation · Data minimisation · Transparency · Accuracy · Security · Accountability

All map to concrete practice: identify a legal basis, don’t repurpose silently, collect only what’s needed, explain plainly, keep data correct, protect it, document your decisions.

Next class

Chapter 7 — Privacy Practice

Bring a laptop with R installed. We’ll apply k-anonymity, differential privacy, and synthetic data generation to real insurance micro-data.

Abowd, John M. 2018. “The u.s. Census Bureau Adopts Differential Privacy.” Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2867.
California Privacy Protection Agency. 2023. California Privacy Rights Act (Amending the California Consumer Privacy Act). https://cppa.ca.gov/regulations/consumer_privacy_act.html.
Carlini, Nicholas, Florian Tramèr, Eric Wallace, et al. 2021. “Extracting Training Data from Large Language Models.” 30th USENIX Security Symposium, 2633–50.
Federal Trade Commission. 2026. “FTC Finalizes Order Settling Allegations That GM and OnStar Collected and Sold Geolocation Data Without Consumers’ Informed Consent.” https://www.ftc.gov/news-events/news/press-releases/2026/01/ftc-finalizes-order-settling-allegations-gm-onstar-collected-sold-geolocation-data-without-consumers.
Garante per la protezione dei dati personali. 2023. Provvedimento Del 30 Marzo 2023 [9870832]: Limitazione Provvisoria Del Trattamento Dei Dati Di ChatGPT. https://www.garanteprivacy.it/home/docweb/-/docweb-display/docweb/9870832.
Internal Revenue Service, Statistics of Income Division. 2018. Creating Homogeneous Synthetic Individual Tax Files for Public Use. Internal Revenue Service. https://www.irs.gov/pub/irs-soi/18rpsynthindtaxfiles.pdf.
Li, Ninghui, Tiancheng Li, and Suresh Venkatasubramanian. 2007. “T-Closeness: Privacy Beyond k-Anonymity and l-Diversity.” Proceedings of the 23rd IEEE International Conference on Data Engineering, 106–15.
Liu, Yang, Tao Fan, Tianjian Chen, Qian Xu, and Qiang Yang. 2021. “FATE: An Industrial Grade Platform for Collaborative Learning with Data Protection.” Journal of Machine Learning Research 22 (226): 1–6.
Machanavajjhala, Ashwin, Daniel Kifer, Johannes Gehrke, and Muthuramakrishnan Venkitasubramaniam. 2007. “\ell-Diversity: Privacy Beyond k-Anonymity.” ACM Transactions on Knowledge Discovery from Data 1: 3.
Narayanan, Arvind, and Vitaly Shmatikov. 2008. “Robust de-Anonymization of Large Sparse Datasets.” Proceedings of the 2008 IEEE Symposium on Security and Privacy, 111–25. https://doi.org/10.1109/SP.2008.33.
Nasr, Milad, Nicholas Carlini, Jonathan Hayase, et al. 2023. “Scalable Extraction of Training Data from (Production) Language Models.” arXiv Preprint arXiv:2311.17035.
National People’s Congress (China). 2021. Personal Information Protection Law of the People’s Republic of China (中华人民共和国个人信息保护法). http://www.npc.gov.cn/npc/c2/c30834/202108/t20210820_313088.html.
Nissenbaum, Helen. 2004. “Privacy as Contextual Integrity.” Washington Law Review 79 (1): 119–57.
Office of the Australian Information Commissioner. 2023. “Australian Privacy Principles.” https://www.oaic.gov.au/privacy/australian-privacy-principles.
OpenAI. 2023. “March 20 ChatGPT Outage: Here’s What Happened.” March. https://openai.com/index/march-20-chatgpt-outage/.
Personal Data Protection Commission Singapore. 2012. Personal Data Protection Act 2012. https://www.pdpc.gov.sg/overview-of-pdpa/the-legislation/personal-data-protection-act.
Ray, Siladitya. 2023. “Samsung Bans ChatGPT Among Employees After Sensitive Code Leak.” Forbes, May. https://www.forbes.com/sites/siladityaray/2023/05/02/samsung-bans-chatgpt-and-other-chatbots-for-employees-after-sensitive-code-leak/.
Regulation (EU) 2016/679 of the European Parliament and of the Council (General Data Protection Regulation) (2016).
Solove, Daniel J. 2006. “A Taxonomy of Privacy.” University of Pennsylvania Law Review 154 (3): 477–564.
Sweeney, Latanya. 2000. Simple Demographics Often Identify People Uniquely. Data Privacy Working Paper 3. Carnegie Mellon University.
Sweeney, Latanya. 2002. “K-Anonymity: A Model for Protecting Privacy.” International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems 10 (5): 557–70.