Unit 1: Introduction to Data Privacy & Privacy Regulations
1.1 What is Data Privacy?
Data privacy is the right of individuals to control how their personal information is collected, stored,
used, and shared by organizations. It ensures that sensitive data is protected from unauthorized
access and misuse. Data privacy is not just a technical concept — it is a fundamental human right
recognized globally.
Data privacy operates at the intersection of three key concerns:
• Legal compliance — organizations must follow data protection laws (GDPR, DPDPA,
HIPAA).
• Ethical responsibility — individuals deserve respect and control over their own information.
• Technical protection — systems must be designed to prevent unauthorized data access.
Key Definition: Data Privacy = Right of individuals to control personal information +
Protection from unauthorized collection, use or sharing of that information.
1.2 Historical Context of Data Privacy
The evolution of data privacy is closely linked to advances in communication technology and
changing societal norms. Below is a detailed timeline of key milestones:
Year Milestone
1890 Brandeis & Warren proposed the 'Right to Privacy' in Harvard Law Review. This
was triggered by concerns about press intrusion and new photographic
technology. It argued people have a right to be let alone.
1948 Universal Declaration of Human Rights (Article 12) recognized privacy as a
fundamental human right. 'No one shall be subjected to arbitrary interference
with his privacy, family, home or correspondence.'
1970s First formal data protection laws emerged: Sweden's Data Act (1973) was the
world's first national data protection law. US Privacy Act (1974) regulated federal
government use of personal data. US Fair Credit Reporting Act (1970) regulated
consumer credit information.
1980 OECD Guidelines on Privacy established foundational Fair Information Practices
(FIPPs): purpose specification, use limitation, data quality, security safeguards,
openness, individual participation, accountability.
1995 EU Data Protection Directive standardized privacy rules across Europe. Created
baseline protections for personal data processing.
2016/2018 GDPR (General Data Protection Regulation) enacted in 2016, came into full
force May 25, 2018. Landmark regulation with extra-territorial scope, strong
enforcement powers, and significant individual rights. Up to EUR 20 million or
4% of global annual revenue in fines.
2023 India passed the Digital Personal Data Protection Act (DPDPA 2023). Based on
GDPR principles, adapted for India's digital ecosystem. Establishes rights of
'Data Principals' and duties of 'Data Fiduciaries'.
1.3 Types of Sensitive Data
Data can be classified based on how directly it identifies an individual and how sensitive its
exposure would be. Understanding these categories is essential for data protection.
Term Definition / Explanation
Explicit Identifiers (EI) Data that directly and unambiguously identifies a specific person.
Examples: Full Name, National ID (Aadhar, SSN, Passport), Email
address, Phone number, Bank account number.
Quasi-Identifiers (QI) Data that does not directly identify someone but can be combined
with other data to identify them. Examples: Date of Birth, Gender,
Zip/Pin Code, Occupation. WARNING: DOB + Gender + Zip Code
can identify 83% of US population (Sweeney, 2002).
Sensitive Data (SD) Personal information that is confidential and whose exposure could
cause significant harm. Examples: Health/medical records, financial
status, salary, religion, caste, political opinions, sexual orientation,
biometric data.
PII (Personally Any data that can directly or indirectly identify a specific individual.
Identifiable Information) Includes both Explicit Identifiers and Quasi-Identifiers used in
combination.
Person-related Data Data relating to a person but not strictly PII. Examples: Interest
patterns, location data, browsing behavior, purchase history, social
media activity.
Proprietary / Confidential Business/organizational data protected by contracts, NDAs, or
internal policies. Examples: Trade secrets, business strategies,
financial projections, source code.
Important: Zip code + DOB + Gender alone can uniquely identify 83% of the US population
(Sweeney, 2002). This demonstrates why Quasi-Identifiers, though not directly identifying,
are extremely dangerous and must be handled carefully.
Data Categories under GDPR (Special Category Data):
• Racial or ethnic origin
• Political opinions
• Religious or philosophical beliefs
• Trade union membership
• Genetic and biometric data
• Health data
• Sexual orientation or sex life
These categories require EXPLICIT consent and stricter protection under GDPR Article 9.
1.4 Privacy Laws and Regulations
GDPR — General Data Protection Regulation (EU, 2018)
GDPR is the most comprehensive and widely influential data protection law in the world. It applies
to any organization worldwide that processes personal data of EU residents.
• Applies to ALL companies globally that process EU residents' personal data,
regardless of where the company is located (extra-territorial effect). Scope:
• Right to Access (know what data is held), Right to Erasure / 'Right to be Forgotten',
Right to Data Portability, Right to Rectification, Right to Object to processing, Right
not to be subject to automated decisions. Key Individual Rights:
• Lawfulness/fairness/transparency, Purpose Limitation, Data Minimization, Accuracy,
Storage Limitation, Integrity & Confidentiality, Accountability. Key Principles:
• Up to 4% of global annual revenue OR EUR 20 million (whichever is higher). Tier 1
violations (minor): up to 2% or EUR 10 million. Penalties:
• Organizations must report data breaches to the supervisory authority within 72 hours
of becoming aware. Must also notify affected individuals 'without undue delay' if high
risk. 72-hour rule:
• Mandatory for public authorities, large-scale processing, or processing of special
category data. Data Protection Officer (DPO):
India — Digital Personal Data Protection Act, 2023 (DPDPA)
• Governs collection, storage, and processing of digital personal data within India and for
Indian residents outside India.
• The individual whose data is being processed (equivalent to GDPR's 'data subject').
Data Principal:
• The organization that determines the purpose and means of processing personal data
(equivalent to GDPR's 'data controller'). Data Fiduciary:
• High-risk or large-scale processors subject to stricter obligations — must appoint
DPO, conduct Data Protection Impact Assessments (DPIAs). Significant Data
Fiduciaries:
• Right to information about processing, Right to correction and erasure, Right to
grievance redressal, Right to nominate another person to act on their behalf. Rights of
Data Principals:
• Adjudicatory body established to handle complaints and impose penalties (up to Rs.
250 crore for certain violations). Data Protection Board:
US Privacy Laws (Sectoral Approach)
Unlike GDPR, the US does not have a single comprehensive federal privacy law. Instead, it uses a
sectoral approach with separate laws for different industries:
Term Definition / Explanation
HIPAA (1996) Health Insurance Portability and Accountability Act. Protects
medical and health insurance records. Covered entities: hospitals,
insurance companies, healthcare providers.
FERPA (1974) Family Educational Rights and Privacy Act. Protects student
educational records. Students (or parents if under 18) control
access to records.
COPPA Children's Online Privacy Protection Act. Protects online data of
children under 13. Requires parental consent before collecting
children's data.
GLBA (1999) Gramm-Leach-Bliley Act. Protects consumer financial data held by
financial institutions (banks, insurers, investment companies).
CCPA / CPRA California Consumer Privacy Act / California Privacy Rights Act.
Strong state-level law similar to GDPR. Gives California residents
rights to access, delete, and opt-out of sale of personal data.
Exam Tip: GDPR = Most comprehensive global regulation with extra-territorial reach. India's
DPDPA 2023 is based on GDPR principles but adapted for India. USA has NO single federal
privacy law — uses sectoral approach. HIPAA = health, FERPA = education, COPPA =
children.
Unit 2: Data Privacy Attacks, Cryptography & Data Protection
2.1 Types of Attacks / Data Breaches
Security attacks are broadly classified into two categories based on how they interact with the target
system:
Passive Attacks (Eavesdropping — harder to detect)
Passive attacks involve monitoring or surveillance of transmissions. No data is altered. They are
hard to detect because they leave no trace.
• Intercepting and reading emails, file transfers, phone calls, or private conversations.
The attacker gains access to sensitive information. Release of Message Contents:
• Even if message content is encrypted, the attacker observes communication patterns
— frequency of messages, length, source/destination addresses. This metadata
reveals behavioral patterns. Traffic Analysis:
Active Attacks (Modify or disrupt data — harder to prevent)
Active attacks involve modification of data stream or creation of false streams. They are detectable
but difficult to prevent entirely.
Term Definition / Explanation
Masquerade Attacker pretends to be another entity (user, server, or system) to
gain unauthorized access. Example: phishing emails impersonating
a bank.
Replay Attack Captured data (e.g., login credentials, authentication tokens) is
retransmitted to reproduce an authorized session. Example:
replaying a captured session token.
Data Modification Legitimate message is altered, reordered, delayed, or duplicated by
the attacker. Example: changing the recipient's bank account in a
wire transfer.
Denial of Service (DoS) Prevents normal use of systems by flooding network resources or
disabling services. DDoS = Distributed DoS using multiple
machines. Targets availability.
Man-in-the-Middle (MitM) Attacker secretly intercepts and possibly alters communication
between two parties who believe they are communicating directly
with each other.
SQL Injection Malicious SQL code inserted into input fields to manipulate
database queries. Can expose, modify, or delete database
contents.
Phishing Social engineering attack using deceptive emails/websites to trick
users into revealing credentials or installing malware.
Ransomware Malware that encrypts victim's data and demands ransom payment
for decryption key. Major threat to organizational data.
2.2 Impact of Data Breaches
Data breaches have far-reaching consequences for both organizations and individuals:
Impact on Organizations:
• Legal and financial penalties — GDPR fines up to 4% of global annual revenue; DPDPA
fines up to Rs. 250 crore.
• Reputational damage — Loss of customer trust; long-term brand damage.
• Business disruption — Operational downtime, loss of business continuity.
• Legal liability — Class action lawsuits from affected individuals.
• Cost of remediation — Forensic investigation, system repair, notification costs.
Impact on Individuals:
• Identity theft — Stolen PII used to impersonate the victim for financial fraud.
• Financial loss — Unauthorized transactions, credit card fraud, loan fraud.
• Emotional distress — Anxiety, fear, loss of control over personal information.
• Discrimination — Exposed sensitive data (health, religion) used to discriminate.
• Reputational harm — Private information made public causing social/professional damage.
IDC Survey Finding: Data leakage is ranked as the #1 threat to organizations — higher
than viruses, worms, or physical theft. This highlights the critical importance of data privacy
controls.
2.3 Security Objectives — The CIA Triad
The CIA Triad represents the three core objectives that information security systems must achieve:
Term Definition / Explanation
Confidentiality Information must not be disclosed to unauthorized individuals,
processes, or programs. Loss of confidentiality = unauthorized
disclosure. Protected by: encryption, access controls, need-to-know
policies.
Integrity Data must be changed only in an authorized manner by authorized
entities. Loss of integrity = unauthorized modification or destruction
of data. Protected by: digital signatures, hashing, checksums.
Availability Systems and services must work promptly and service must not be
denied to authorized users. Loss of availability = disruption of
access to services. Protected by: redundancy, failover systems,
DDoS mitigation.
Authenticity Verifying that users and messages are genuine — that users are
who they claim to be. Ensures trust in message origin and identity.
Provided by: digital certificates, biometrics, MFA.
Accountability Actions of an entity can be traced uniquely to that entity. Supports
non-repudiation (cannot deny an action). Provided by: audit logs,
digital signatures.
2.4 Introduction to Cryptography
Cryptography is the discipline of transforming data to hide its semantic content, prevent
unauthorized use, or prevent undetected modification. It is the mathematical foundation of data
security.
Cryptography ensures four core security services:
• Confidentiality — Only authorized parties can read the data.
• Integrity — Data has not been altered in transit.
• Authentication — Verifying the identity of sender/receiver.
• Non-repudiation — Sender cannot deny sending a message.
3 Categories of Cryptographic Algorithms: 1. Keyless Algorithms (Hash functions —
SHA-256, MD5). 2. Single-Key / Symmetric Algorithms (AES, DES — same key for
encrypt/decrypt). 3. Two-Key / Asymmetric Algorithms (RSA — public/private key pair).
2.5 Symmetric Encryption (Single-Key / Secret-Key)
In symmetric encryption, the SAME secret key is used for both encryption and decryption. Both
sender and receiver must securely share and maintain the same key.
How it works:
1. Sender takes Plaintext + Secret Key as input to the encryption algorithm.
2. Algorithm produces Ciphertext (scrambled, unreadable output).
3. Ciphertext is transmitted over the network.
4. Receiver uses the SAME Secret Key + decryption algorithm to recover Plaintext.
Term Definition / Explanation
Plaintext Original readable message or data before encryption.
Ciphertext Scrambled, unreadable output after applying encryption.
Secret Key Shared key known ONLY to sender and receiver. Must be kept
secret.
Encryption Algorithm Mathematical function (substitutions/transformations) applied to
plaintext + key.
Decryption Algorithm Reverse of encryption algorithm; uses same key to recover
plaintext.
Types of Symmetric Ciphers:
• Encrypt one bit/byte at a time. Fast. Used for streaming data. Example: RC4. Stream
Ciphers:
• Encrypt fixed-size blocks of data (64-bit or 128-bit blocks). Example: AES (128-bit
blocks). Block Ciphers:
AES — Advanced Encryption Standard:
• Recommended by NIST (National Institute of Standards and Technology).
• Key lengths: 128-bit (AES-128), 192-bit (AES-192), or 256-bit (AES-256).
• AES-256 is considered quantum-resistant for near-term threats.
• Used in: WiFi (WPA2/WPA3), SSL/TLS, disk encryption (BitLocker), file encryption.
Attacks on Symmetric Encryption:
• Exploiting mathematical weaknesses in the algorithm to recover plaintext without the
key. Cryptanalysis:
• Systematically trying every possible key until the correct one is found. Infeasible for
AES-256 (2^256 possible keys). Brute-force Attack:
Advantage vs Disadvantage: ADVANTAGE: Very fast — suitable for encrypting large
amounts of data. DISADVANTAGE: Key distribution problem — how to securely share the
secret key with the other party? If key is intercepted, all security is lost.
2.6 Asymmetric Encryption (Two-Key / Public-Key Cryptography)
Asymmetric encryption uses a mathematically related KEY PAIR: a Public Key (shared openly with
everyone) and a Private Key (kept strictly secret by the owner). Data encrypted with one key can
only be decrypted with the other key of the pair.
Two modes of use:
• Sender encrypts with recipient's PUBLIC key → Only recipient's PRIVATE key can
decrypt. Only the intended recipient can read the message. Confidentiality mode:
• Sender encrypts hash with their PRIVATE key → Anyone with sender's PUBLIC key
can verify. Proves the message came from the sender (authentication + non-
repudiation). Authentication mode (Digital Signature):
RSA — Rivest-Shamir-Adleman:
• Most widely used asymmetric algorithm.
• Security based on difficulty of factoring very large numbers.
• Recommended key size: 2048 bits or higher (NIST recommendation).
• Used for: HTTPS/TLS handshakes, email encryption (PGP/S-MIME), digital certificates.
Term Definition / Explanation
Advantage Solves the key distribution problem — public key can be shared
openly without security risk. No need to securely exchange keys
beforehand.
Disadvantage Much SLOWER than symmetric encryption (100x to 1000x slower).
Only practical for small amounts of data (keys, hashes, digital
signatures).
Hybrid Encryption In practice, asymmetric encryption is used to securely exchange a
symmetric session key. Then fast symmetric encryption is used for
the actual data. This is how HTTPS works.
Key Difference: Symmetric = 1 shared secret key (fast, key distribution problem).
Asymmetric = 2 keys, public + private (slow, but solves key distribution). Real systems use
HYBRID: asymmetric to exchange key, symmetric to encrypt data.
2.7 Hashing
A hash function is a mathematical function that converts a variable-length input (message of any
size) into a fixed-length output called a hash value, digest, or fingerprint. It is a ONE-WAY function
— it is computationally infeasible to reverse.
Properties of a Cryptographic Hash Function:
• Always produces output of same length regardless of input size. SHA-256 always
produces 256-bit (64 hex characters) output. Fixed output size:
• Same input always produces exactly the same output. Deterministic:
• Fast to compute for any given input. Efficient:
• Given a hash value H, it is computationally infeasible to find the input M such that H =
hash(M). One-way property. Pre-image resistant:
• Given input M1, it is infeasible to find a different M2 such that hash(M1) = hash(M2).
Second pre-image resistant:
• It is computationally infeasible to find ANY two different inputs M1 and M2 such that
hash(M1) = hash(M2). Collision resistant:
• Even a tiny change in input (1 bit) causes a completely different hash output. Makes
patterns undetectable. Avalanche effect:
Hash Standards:
• Produces 256-bit hash. NIST recommended. Used in Bitcoin, SSL certificates, digital
signatures. SHA-256 (SHA-2 family):
• Latest NIST standard. Different internal design (Keccak). Also produces variable
output sizes. SHA-3:
• Older 128-bit hash. Considered cryptographically broken — do NOT use for security
purposes. MD5:
Applications of Hashing:
• Store hash of password, not plaintext. On login, hash the entered password and
compare. Even database breach doesn't reveal passwords. Password storage:
• Hash file before and after transmission; if hashes match, data was not modified. Data
integrity verification:
• Hash is computed first, then the hash (not the whole document) is digitally signed.
Digital signatures:
• Each block contains hash of previous block, making tampering detectable. Blockchain:
MAC — Message Authentication Code:
A MAC is computed using Hash + Secret Key. It provides both data integrity AND authentication
(proves message came from someone who knows the key). Unlike a digital signature, MAC requires
both parties to share the same secret key.
2.8 Digital Signatures
A digital signature is a cryptographic transformation of data that provides a mechanism for verifying
the origin of a message (authentication), data integrity, and non-repudiation (the signer cannot deny
signing).
How a Digital Signature works (Step by Step):
5. Signer computes the cryptographic hash of the document (e.g., SHA-256).
6. Hash is encrypted using the signer's PRIVATE key — this encrypted hash IS the digital
signature.
7. Original document + digital signature are sent to the receiver.
8. Receiver decrypts the signature using the signer's PUBLIC key to recover the original hash.
9. Receiver independently calculates the hash of the received document.
10. If both hashes MATCH → Signature is VALID (document is authentic and unmodified). If
hashes DON'T MATCH → Document was altered or signature is fake.
Why sign the HASH and not the whole document? Asymmetric encryption is very slow.
Hashing produces a short fixed-length summary (e.g., 256 bits). Signing only the hash is
much faster and still secures the entire document because hash = unique fingerprint of the
document.
Uses of Digital Signatures:
• Verifying the email sender's identity (S/MIME, PGP). Email authentication:
• Ensuring downloaded software has not been tampered with. Software code signing:
• Legal contracts, government documents, e-filing. Document integrity:
• Signer legally cannot deny having signed — provides legal accountability. Non-
repudiation:
• Web servers sign their certificates to prove identity to browsers. SSL/TLS certificates:
2.9 Public-Key Infrastructure (PKI)
PKI is a framework of policies, standards, hardware, software, and procedures needed to create,
manage, distribute, store, and revoke digital certificates that bind public keys to identities.
Term Definition / Explanation
Certificate Authority (CA) Trusted third party that issues, signs, and manages digital
certificates. Vouches for the identity of certificate holders. Example:
DigiCert, VeriSign, Let's Encrypt. Acts like a digital notary.
Registration Authority Handles identity verification and vetting before certificate issuance.
(RA) Offloads identity verification from CA. Does not issue certificates
directly.
Public-Key Certificate Digital document containing: owner's identity, owner's public key,
CA's digital signature, validity period, serial number. Proves 'this
public key really belongs to this entity.'
Certificate Revocation List of certificates that have been revoked before their expiry date
List (CRL) (e.g., key compromise, employee left). Parties must check CRL
before trusting a certificate.
OCSP Online Certificate Status Protocol — real-time alternative to CRL for
checking if a certificate is currently valid.
Unit 3: Data Collection, Use and Reuse
3.1 Harms Associated with Data Collection, Use and Reuse
When personal data is collected, shared, or reused without adequate protection or consent, it can
cause significant harm to individuals. These harms can be physical, financial, psychological, or
social.
Term Definition / Explanation
Identity Theft Using stolen PII to impersonate someone for financial gain.
Example: Opening credit cards, taking loans, filing tax returns in
victim's name. Affects millions globally each year.
Financial Loss Direct financial harm through fraud — unauthorized credit card
transactions, bank transfers, investment fraud, insurance fraud
using stolen data.
Discrimination Data used to unfairly deny credit, insurance, employment, housing
based on protected characteristics (race, health status, religion)
inferred from data profiles.
Surveillance Continuous, systematic tracking of individuals' location, behavior,
activities, and associations without meaningful consent. Can enable
authoritarian control.
Reputation Damage Exposure of private or embarrassing information causes harm to
social relationships, professional standing, and mental health.
Re-identification Combining 'anonymized' data with external datasets to re-identify
individuals. Even 'anonymous' data can be dangerous if proper
techniques aren't used.
Psychological Harm Stalking, harassment, cyberbullying enabled by data exposure.
Chilling effect on free expression when people know they are being
monitored.
Secondary Use Harms Data collected for one purpose (medical research) used for another
without consent (insurance risk assessment), causing unexpected
harm.
Privacy Paradox: People express strong concern about privacy but willingly share data for
convenience (social media, loyalty cards, apps). This contradiction arises from: lack of
transparency about actual data use, perceived benefits outweighing perceived risks, difficulty
understanding complex privacy policies, and normalized data sharing culture.
3.2 Introduction to Data Anonymization
Anonymization is the process of logically separating identifying information (PII) from sensitive data
so that individuals cannot be re-identified, even when combining the anonymized data with other
available datasets.
The fundamental goal of anonymization is to enable:
• Data sharing and analysis for research, public health, business intelligence.
• Statistical insights without compromising individual privacy.
• Compliance with data protection laws (GDPR allows processing of properly anonymized
data without consent).
The Privacy-Utility Trade-off: More privacy (stronger anonymization) = Less useful data
(reduced accuracy/granularity). Less privacy (weaker anonymization) = More useful data
(higher accuracy). The challenge of anonymization is finding the optimal balance for the
specific use case.
Privacy vs Anonymity — Important Distinction:
Term Definition / Explanation
Privacy Identity IS known, but the associated personal FACT is hidden. We
know WHO but not WHAT. Example: We know Alice's name but not
her medical condition.
Anonymity Personal FACT is known, but IDENTITY is hidden. We know WHAT
but not WHO. Example: We know someone has HIV but not who
that person is.
Two-Step Anonymization Process:
11. Systematically substitute, suppress, or scramble ALL Explicit Identifiers (EI) —
names, ID numbers, SSNs, account numbers, email addresses. These are fully
removed or replaced. Data Masking (Step 1):
12. Modify Quasi-Identifiers (QI) — DOB, gender, zip code — using transformation
function T so individuals cannot be re-identified even when data is combined with
external datasets. De-identification (Step 2):
Mathematical Formula: D' = T(D) where: D = original dataset, D' = anonymized dataset, T =
transformation function applied to QI fields. EI fields are fully masked/removed. SD (sensitive
data) is left unchanged for analysis.
3.3 Data Anonymization Techniques
Multiple techniques exist, each with different strengths and trade-offs. They can be combined for
stronger protection:
Term Definition / Explanation
Randomization / Adding statistical noise (random values) to data so individual values
Perturbation are distorted but aggregate patterns are preserved. Example: Add
±5% random noise to salary values. Used in differential privacy.
Generalization Replacing specific values with broader ranges or categories.
Example: Age 25 → Age group '20-30'. Zip code 560001 →
'560xxx' or just city name. Reduces precision but retains utility.
Suppression Completely removing or hiding identifying values that cannot be
generalized safely. Example: Replacing name with 'XXX'. Outliers
that are unique are suppressed.
k-Anonymity Each record in the dataset must be indistinguishable from at least k-
1 other records based on Quasi-Identifiers. If k=5, each record
looks identical to at least 4 others. Prevents identity disclosure.
l-Diversity Extension of k-anonymity. Each equivalence class (group of k
identical records) must have at least l DISTINCT values for the
sensitive attribute. Prevents attribute disclosure even when identity
is protected.
t-Closeness Extension of l-diversity. The distribution of sensitive attributes in
each equivalence class must closely match the distribution in the
overall dataset (within threshold t). Prevents inference attacks.
Tokenization Replace sensitive data with a non-sensitive placeholder (token) that
has no mathematical relationship to the original. The token maps to
real data only via a secure lookup table. Used in credit card
processing (PCI DSS).
Data Masking Substitute real data with realistic but fictitious data for
testing/development environments. Structure and format preserved
but values are fake. Example: Real customer database → Masked
test database.
Synthetic Data Generate entirely artificial datasets that statistically resemble real
Generation data — same distributions, correlations, and patterns — but contain
no real individual's information. Increasingly popular for AI/ML
training.
Differential Privacy Mathematical framework that adds carefully calibrated noise to
query results so that the presence or absence of any single
individual cannot be detected. Used by Apple, Google, US Census
Bureau.
Cryptography vs Anonymization: Cryptography is BINARY — data is either fully private
(encrypted) or fully accessible (decrypted). Anonymization works in SHADES OF GREY —
you can tune the level of privacy protection vs data utility. This makes anonymization more
flexible but also more complex.
k-Anonymity — Detailed Example:
If k=3, the following table shows anonymized records where no individual can be distinguished from
at least 2 others:
Age Group Gender Zip Code Disease
20-30 Female 560xxx Hypertension
20-30 Female 560xxx Hypertension
20-30 Female 560xxx Diabetes
30-40 Male 561xxx Asthma
30-40 Male 561xxx Asthma
30-40 Male 561xxx Cancer
3.4 Challenges in Anonymizing Different Data Types
Different data types have unique characteristics that make anonymization more complex:
Multidimensional / Relational Data:
• High dimensionality — many attributes make it hard to generalize without losing all utility.
• Difficulty determining boundary between QI and SD when attacker has background
knowledge.
• Clusters in sensitive datasets complicate k-anonymity implementation.
Transaction Data (Sparse, High-Dimensional):
• Very high dimensionality — thousands of products/items.
• Binary data (0 or 1 — bought or not bought) — unique purchase patterns identify individuals.
• Example: Netflix Prize dataset — researchers re-identified users from movie ratings.
• Conventional relational anonymization techniques not directly applicable.
Longitudinal Data (Healthcare Studies):
• Repeated measurements on same individual over time — data is correlated within person.
• Must prevent BOTH identity disclosure AND attribute disclosure simultaneously.
• Temporal patterns can uniquely identify individuals (e.g., consistent medication schedules).
Graph Data (Social Networks):
• Three types of disclosure: Identity disclosure, Link disclosure (friendship revealed),
Content/Attribute disclosure.
• Graph structure metrics (betweenness centrality, path length) can re-identify individuals.
• Modifying one node can disrupt the entire network structure.
• Example: AOL released 'anonymized' search queries — users re-identified from query
patterns.
Time Series Data:
• High dimensionality and continuously growing volume.
• Must retain statistical properties (mean, variance, autocorrelation) for utility.
• Must support range queries and pattern matching after anonymization.
Key Challenge: There is ALWAYS a trade-off between Privacy and Utility. Optimum
anonymization = maximum privacy protection with minimum utility loss. Perfect
anonymization = zero utility. Perfect utility = zero privacy. The goal is the best possible
balance for the specific use case.
Unit 4: Ethical Considerations in Data Privacy
4.1 Privacy and Surveillance
Modern digital surveillance systems are automated, data-driven decision-making systems that
continuously collect, aggregate, and analyze personal data to predict behavior and make or
influence decisions about individuals and groups.
5-Step Model of a Digital Surveillance System:
13. Identify the 'propensity of interest' and correlate it with observable attributes.
Example: Likelihood of buying a product based on browsing history; credit default
risk based on spending patterns; terrorism risk based on travel patterns. Define:
14. Collect attribute data across a population. Categorize individuals into behavioral
segments or risk classes using machine learning algorithms. Identify:
15. Systematically act on each identified segment. Show targeted advertisements,
recommend medical treatments, flag individuals for investigation, adjust prices,
determine insurance premiums. Intervene:
16. Collect data on actual outcomes to evaluate system accuracy. Did the prediction
match reality? Did the intervention achieve the desired result? Observe Outcomes:
17. Measure false positives (wrongly flagged), false negatives (missed cases), and overall
net benefit vs. harm of the surveillance system. Test Error Rate:
Types of Surveillance Concerns:
• Intelligence agencies collecting mass communication data (NSA PRISM program,
Snowden revelations). Raises concerns about civil liberties and chilling effects on
free speech. Government surveillance:
• Tech companies tracking online behavior, location, purchases to build detailed
profiles for targeted advertising. Business model of 'surveillance capitalism.'
Corporate surveillance:
• Government systems (e.g., China's Social Credit System) that score citizens based on
behavior and restrict access to services based on scores. Social credit systems:
• Monitoring employee emails, keystrokes, location, productivity metrics. Raises
concerns about dignity and trust. Workplace surveillance:
Term Definition / Explanation
Insecure Use Unauthorized or illegal use of data (identity theft, fraud, hacking).
GDPR and most privacy laws primarily address this type of harm.
Imprecise Use Legal but poor-quality or biased algorithms causing harm (wrong
medical diagnosis, biased credit scoring, false terrorism flags).
GDPR largely does NOT address this — a major gap in current
regulation.
Important Distinction: GDPR primarily regulates INSECURE use (unauthorized/illegal
processing). However, IMPRECISE use — legal processing with biased or inaccurate
algorithms — causes significant harm that existing regulations largely fail to address. This is
a critical gap in current data protection frameworks.
4.2 Ethics of Data Collection and Use
Data ethics provides the moral framework for evaluating whether data collection and use practices
are fair, just, and respectful of human dignity — beyond mere legal compliance.
Key Ethical Principles (GDPR / OECD Fair Information Practices):
Term Definition / Explanation
Purpose Limitation Data collected for a specific stated purpose must not be used for a
different, incompatible purpose without additional consent.
Example: Medical data collected for treatment cannot be sold to
insurance companies.
Data Minimization Collect only the minimum amount of data that is genuinely
necessary for the stated purpose. Do not collect 'just in case' data.
More data = more risk if breached.
Informed Consent Individuals must be given clear, plain-language information about
what data is collected, why, who has access, and for how long —
and they must freely agree without coercion.
Transparency Organizations must be open and honest about their data practices.
Privacy policies must be understandable, not buried in legal jargon.
Individuals should know what happens to their data.
Accountability Data controllers are legally responsible and can be held liable for
compliance with privacy principles. Senior leadership must
champion privacy as an organizational value.
Privacy by Design Privacy protection must be built into systems, products, and
services from the very beginning of the design process — not
added as an afterthought or patch.
Fairness Data processing must not unfairly harm individuals. Algorithms must
not produce outcomes that systematically disadvantage any group
or individual without justification.
Consent Issues in Practice:
• Privacy policy length and complexity: Average user would need 76 work days per year to
read all privacy policies for services they use.
• Bundled consent: Users forced to accept all-or-nothing terms — no granular choice over
which data is collected.
• Power imbalance: Organizations design consent mechanisms; individuals have little real
bargaining power.
• Privacy Paradox: People say they value privacy but share data freely for convenience, social
connection, or perceived benefits.
• IoT challenge: Devices (smart speakers, wearables) continuously collect data with no
meaningful moment of consent.
Privacy by Design — 7 Foundational Principles (Ann Cavoukian):
18. Prevent privacy violations before they occur, not just respond after. Proactive not
Reactive:
19. Maximum privacy is the default setting — users must actively choose to share less
privacy protection. Privacy as the Default:
20. Privacy is a core component, not a bolt-on feature. Privacy Embedded into Design:
21. Privacy AND functionality — not a zero-sum trade-off. Full Functionality:
22. Strong security throughout the entire lifecycle of data. End-to-End Security:
23. Verify that the system operates as promised. Visibility and Transparency:
24. Put the individual at the center — user-centric design. Respect for User Privacy:
4.3 Bias and Discrimination in Data Analysis
Data-driven algorithms can produce biased, discriminatory outcomes that harm individuals or
groups — even when the data collection was legal and consented to. Algorithmic bias is one of the
most significant ethical challenges in data privacy.
Types of Algorithmic Bias:
Term Definition / Explanation
Historical Bias Training data reflects past human discrimination and prejudice.
Algorithms learn and reproduce these patterns. Example: Hiring
algorithms trained on historically male-dominated datasets learn to
penalize female applicants.
Representation Bias Training data underrepresents certain groups (minorities, elderly,
disabled). Algorithm performs poorly and inaccurately for these
underrepresented groups. Example: Facial recognition fails more
often for dark-skinned women.
Measurement Bias Proxy variables used as substitutes for protected attributes (race,
gender, religion) that cannot legally be used. Proxies produce same
discriminatory outcomes indirectly. Example: ZIP code used as
proxy for race in lending decisions.
Aggregation Bias A single model applied to different population groups with different
underlying patterns. Ignores within-group variation. Example: One
diabetes risk model for all ethnicities when risk factors differ by
ethnicity.
Feedback Loops Biased predictions lead to actions that generate more biased data,
which reinforces the biased model. Self-reinforcing cycle of
discrimination. Example: Predictive policing sends more police to
already over-policed areas.
Evaluation Bias Model performance measured using benchmarks that don't reflect
real-world population diversity. Model appears accurate but fails for
minority groups.
Real-World Examples of Algorithmic Discrimination:
• Used in US courts to predict likelihood of reoffending. Found to be twice as likely to
falsely flag Black defendants as high-risk compared to white defendants. COMPAS
recidivism algorithm:
• Penalized resumes containing the word 'women's' (as in 'women's chess club').
Scrapped after gender bias was discovered. Amazon's hiring algorithm:
• Algorithm systematically assigned lower risk scores to Black patients despite equal
illness severity, resulting in less access to care programs. Healthcare risk scoring:
• Traditional credit scoring penalizes those without credit history — disproportionately
affecting immigrants and low-income communities. Credit scoring:
• NIST study found false positive rates 10-100x higher for African and Asian faces
compared to Caucasian faces in commercial systems. Facial recognition:
Data Ethics Framework — Key Principles for Responsible AI/Data:
Term Definition / Explanation
Fairness Algorithms must not produce outcomes that unfairly disadvantage
any group. Requires defining and measuring fairness carefully
(statistical parity, equal opportunity, calibration).
Transparency Data subjects should know when automated decisions affect them,
what data was used, and how the decision was made. GDPR
Article 13-15 requires this.
Explainability Algorithmic decisions should be understandable to the people
affected. 'Black box' systems are ethically problematic. Right to
explanation under GDPR Article 22.
Accountability Organizations are legally and morally responsible for the outcomes
of their algorithms, including unintended discriminatory effects.
Human Oversight Humans must remain in control of consequential automated
decisions. Pure automation without human review is ethically
problematic for high-stakes decisions.
GDPR Article 22: Data subjects have the RIGHT not to be subject to a decision based
SOLELY on automated processing (including profiling) that produces significant effects on
them (legal or similarly significant). Controller must implement safeguards and provide
means for human intervention and the right to contest the decision.