DATA PRIVACY - COMPREHENSIVE STUDY
NOTES
[Link]. (H) Computer Science | DSE/GE6d | 2024-25
UNIT 1: INTRODUCTION TO DATA PRIVACY AND
PRIVACY REGULATIONS (20 hours)
1.1 NOTION OF DATA PRIVACY
Definition: Data privacy refers to the right of individuals to control what
information is collected about them, how it is used, and who has access to it.
It encompasses the principles and practices that protect personal data from
unauthorized access, misuse, and disclosure.
Key Concepts: - Personally Identifiable Information (PII): Information
that can directly or indirectly identify an individual (name, email, phone, SSN,
IP address, biometric data) - Personal Data: Any information relating to an
identified or identifiable natural person - Data Subject: The person to whom
personal data relates - Data Controller: The entity determining the purposes
and means of data processing - Data Processor: The entity processing data
on behalf of the controller - Privacy: The right to be left alone; freedom from
unauthorized intrusion into personal affairs
Privacy vs Security: - Privacy: Control over access to and use of personal
information - Security: Technical and organizational measures to protect data
from unauthorized access - Both are interconnected but distinct concepts
Why Data Privacy Matters: 1. Protection from identity theft and fraud 2.
Prevention of unauthorized profiling and discrimination 3. Control over personal
information in digital age 4. Prevention of surveillance and manipulation 5.
Maintaining human dignity and autonomy
1.2 HISTORICAL CONTEXT OF DATA PRIVACY
Early Evolution: - 1890s: Warren and Brandeis article on “Right to Privacy”
in US - 1948: Universal Declaration of Human Rights (Article 12) - “No one
shall be subjected to arbitrary interference with his privacy” - 1950: European
Convention on Human Rights - Protection of private and family life - 1960s-
1970s: Rise of computerized data collection; public concern over surveillance
Landmark Legislation: - 1974 - Privacy Act (USA): First comprehensive
federal privacy law; established Privacy Principles - 1981 - OECD Privacy
Guidelines: Established 8 principles: - Collection Limitation Principle - Data
Quality Principle - Purpose Specification Principle - Use Limitation Principle
1
- Security Safeguards Principle - Openness Principle - Individual Participation
Principle - Accountability Principle
Digital Age Developments: - 1995 - EU Data Protection Directive
(95/46/EC): Gold standard for privacy; introduced concepts of GDPR -
2000s: Rise of internet, social media, and big data challenges - 2018 - GDPR
Implementation (EU): Most comprehensive privacy regulation; influenced
global standards - 2023 - Digital Personal Data Protection Act (India):
India’s comprehensive privacy law
Key Events: - Edward Snowden revelations (2013): Mass surveillance pro-
grams - Facebook-Cambridge Analytica scandal (2018): Unauthorized data use
- GDPR enforcement (2018): €20 million+ fines for violations - Rise of AI/ML:
Privacy concerns with automated decision-making
1.3 TYPES OF SENSITIVE DATA
1. Personally Identifiable Information (PII): - Name, surname, identifica-
tion number - Location data, online identifiers - Email address, phone number
- Postal address, IP address
2. Biometric Data: - Fingerprints, facial recognition data - Iris/retina scans
- Voice patterns - DNA information - Gait recognition
3. Health/Medical Data: - Medical history, diagnoses, prescriptions - Mental
health records - Genetic information - COVID-19 vaccination status - Fitness
and health metrics
4. Financial Data: - Bank account numbers, routing numbers - Credit card
information - SSN/Tax ID numbers - Income information - Credit score - Loan
history
5. Behavioral Data: - Web browsing history - Purchase history - Search
queries - Social media activity - Location movement patterns - App usage data
6. Special Category Data (EU GDPR): - Racial or ethnic origin - Political
opinions - Religious beliefs - Trade union membership - Genetic data - Biometric
data (for identification) - Health data - Sex life or sexual orientation data -
Criminal records
7. Children’s Data: - Information about minors (typically <13-16 years) -
Educational records - Parental consent requirements
8. Government/Official Data: - Passport information - Driver’s license -
National identification numbers - Immigration status
Sensitivity Classification:
2
Classification Examples Risk Level
Public Aggregate statistics, Low
published research
Internal Employee directory, Medium
general business info
Confidential Financial records, High
health data, credentials
Restricted National security data, Critical
criminal records,
biometric
1.4 PRIVACY LAWS AND REGULATIONS
1.4.1 European General Data Protection Regulation (GDPR) - 2018
Scope: Applies to any organization processing personal data of EU residents
Key Principles (7 Principles): 1. Lawfulness, Fairness, Transparency:
Processing must be lawful, fair, and transparent to data subject 2. Purpose
Limitation: Data collected for specific purposes; cannot be repurposed with-
out consent 3. Data Minimization: Collect only necessary data (“privacy by
design”) 4. Accuracy: Keep data accurate and up-to-date 5. Storage Limita-
tion: Keep data only as long as necessary 6. Integrity and Confidentiality:
Secure data against unauthorized access 7. Accountability: Document and
demonstrate compliance
Rights of Data Subjects: - Right to be informed: Clear privacy notices
- Right of access: Access own data (Data Subject Access Request) - Right
to rectification: Correct inaccurate data - Right to erasure: “Right to be
forgotten” - Right to restrict processing: Limit how data is used - Right to
data portability: Receive data in structured format - Right to object: Ob-
ject to processing - Rights related to automated decision-making: Chal-
lenge automated decisions
Obligations of Organizations: - Data Protection Impact Assessments
(DPIA) - Privacy policies and notices - Consent management - Data breach
notification (within 72 hours) - Designate Data Protection Officer (DPO) -
Data Protection by Design and Default
Penalties: Up to €20 million or 4% of global annual revenue (whichever is
higher)
1.4.2 Digital Personal Data Protection Act (DPDP) 2023 - INDIA
Scope: Applies to processing of digital personal data of Indian residents; indi-
viduals and organizations in India or accessing Indian residents’ data
3
Key Principles:
1. Collection and Processing:
• Data must be collected for lawful purpose
• Minimal collection (minimum necessary)
• Clear consent for processing
• Different rules for children (under 18) and sensitive data
2. Purpose and Use:
• Data used only for stated purpose
• No secondary use without consent
• Purpose must be communicated to data subject
3. Data Quality:
• Keep data accurate and complete
• Remove/correct inaccurate data
• Store only for necessary period
4. Security and Confidentiality:
• Implement security measures
• Protect against unauthorized access
• Maintain confidentiality
Individual Rights (Data Principals): - Right to information: Know
what data is collected - Right of access: Access own data - Right to cor-
rection: Correct inaccurate data - Right to erasure: Delete data (with ex-
ceptions) - Right to data portability: Receive data in structured format -
Right to lodge complaint: File complaints with Data Protection Board -
Right to restrict processing: Limit use of data
Exempt Categories (No consent required): - Government functions and
national security - Legal obligations - Data already in public domain - Employer-
employee relationship (with restrictions)
Penalties: - Up to �500 crores for contraventions affecting many data principals
- Up to �250 crores for other contraventions - Up to �100 crores for failure to
cooperate with Board
Data Protection Board: Independent regulatory authority to enforce the Act
1.4.3 Other Important Regulations California Consumer Privacy
Act (CCPA) - USA (2018) - Gives Californians rights over personal informa-
tion - Right to know, delete, and opt-out of sale - Right to non-discrimination -
Applies to companies with gross annual revenue >$25M or buying/selling data
of 100,000+ consumers
Health Insurance Portability and Accountability Act (HIPAA) - USA
- Protects health information privacy - Applies to healthcare providers, insurers,
and clearinghouses - Requires encrypted patient records - Breach notification
within 60 days
4
Children’s Online Privacy Protection Act (COPPA) - USA - Protects
children under 13 - Requires verifiable parental consent before collecting data -
No behavioral advertising to children
Payment Card Industry Data Security Standard (PCI DSS) - Techni-
cal standard for payment card data protection - Required encryption, access
controls, monitoring - 6 levels of compliance
Singapore Personal Data Protection Act (PDPA) - Consent-based per-
sonal data protection - Similar to GDPR but with different enforcement
1.5 KEY PRIVACY PRINCIPLES (CORE CONCEPTS)
1. Transparency: - Organizations must clearly communicate data collection
and use - Privacy policies must be accessible, understandable - Data subjects
should know what happens to their data
2. Control: - Individuals should control their personal data - Meaningful
choice in data collection - Ability to access, correct, delete data
3. Consent: - Free, specific, informed, unambiguous consent required - Opt-
in (affirmative action) preferred over opt-out - Separate consent for different
purposes - Easy withdrawal of consent
4. Purpose Limitation: - Data collected for specific, legitimate purpose -
Cannot be used for incompatible purposes - Secondary use requires new consent
5. Data Minimization: - Collect only necessary data - Proportionate to
stated purpose - Don’t collect “just in case”
6. Security: - Implement appropriate technical and organizational measures
- Protect against unauthorized access, loss, theft, damage - Regular security
assessments and updates
7. Accountability: - Organizations responsible for compliance - Document
processing activities - Demonstrate compliance to authorities
UNIT 2: DATA PRIVACY ATTACKS, CRYPTOGRA-
PHY, AND DATA PROTECTION (12 hours)
2.1 TYPES OF ATTACKS AND DATA BREACHES
Definition: A data breach is an unauthorized access, use, disclosure, modifica-
tion, loss, or destruction of personal data.
5
2.1.1 Attack Vectors 1. External Attacks:
a) Hacking/Unauthorized Access: - Exploitation of software vulnerabili-
ties - Weak credentials/password cracking - Phishing emails to gain credentials
- SQL injection attacks - Cross-site scripting (XSS) - Man-in-the-middle (MITM)
attacks - Network sniffing - Example: Equifax breach (2017) - Apache vulnera-
bility
b) Malware: - Viruses, trojans, ransomware - Spyware and keyloggers - Cryp-
tojacking - Example: WannaCry ransomware (2017)
c) Distributed Denial of Service (DDoS): - Overload systems with traffic
- Take down services - Create opportunities for data theft during chaos
d) Social Engineering: - Impersonation of trusted entities - Pretexting (cre-
ating false scenarios) - Baiting (offering something to gain access) - Tailgating
(physical access) - Phishing and spear-phishing
2. Internal Attacks:
a) Insider Threats: - Disgruntled employees stealing data - Unauthorized ac-
cess by legitimate users - Privilege escalation abuse - Example: Edward Snowden
(NSA contractor accessing classified data)
b) Accidental Disclosure: - Misconfigured cloud storage/databases - Sending
data to wrong recipient - Leaving devices unattended - Example: Facebook
exposed 267 million user records via exposed database
3. Supply Chain Attacks: - Compromising third-party vendors - Supplier
data breaches affecting customers - Example: Target breach (2013) via HVAC
supplier
2.1.2 Common Breach Scenarios
Attack Type Method Example
Credential Breach Stolen login credentials LinkedIn 700M accounts
(2021)
Database Breach Direct database access Yahoo 3B accounts
(2014)
Payment Card Theft Point-of-sale malware Home Depot 56M cards
(2014)
Cloud Exposed S3 buckets Millions of exposed
Misconfiguration records
Ransomware Encryption + extortion Colonial Pipeline (2021)
API Abuse Unauthorized API Facebook Cambridge
access Analytica
Device Loss Lost/stolen laptops, Unencrypted portable
phones devices
6
2.2 IMPACT OF DATA BREACHES AND ATTACKS
2.2.1 Individual/Personal Impact:
a) Identity Theft: - Fraudulent credit accounts opened - Tax fraud (filing
false returns) - Medical identity theft - Cost to victim: Thousands in fraudulent
charges, years to restore credit
b) Financial Loss: - Direct theft from accounts - Fraudulent transactions -
Insurance costs - Legal fees - Average individual loss: $100-$1,000+
c) Emotional/Psychological Impact: - Loss of trust in organizations - Anx-
iety and stress - Embarrassment (especially with sensitive data leaks) - Long-
term privacy concerns
d) Reputation Damage: - Reputation harm especially with inti-
mate/sensitive data - Cyberbullying and harassment - Social stigma -
Example: Ashley Madison breach (2015) - infidelity site hack affecting users’
families
e) Misuse of Personal Information: - Targeted scams/fraud - Discrimina-
tory treatment - Targeted advertising manipulation - Political manipulation via
personal data
2.2.2 Organizational Impact:
a) Financial Costs: - Cost of breach response (investigation, forensics, legal)
- Customer notification costs - Regulatory fines and penalties - Regulatory re-
mediation requirements - Litigation and settlement costs - Average data breach
cost: $4.45 million (2023, IBM report)
b) Reputation and Customer Harm: - Loss of customer trust - Customer
churn (defection) - Brand reputation damage - Stock price decline - Negative me-
dia coverage - Example: Equifax (2017) - $700M settlement, massive reputation
damage
c) Operational Impact: - System downtime and recovery - Diversion of IT
resources to incident response - Disruption of business operations - Service in-
terruptions to customers
d) Legal and Regulatory Consequences: - GDPR fines (up to 4% of global
revenue) - Class action lawsuits - Regulatory investigations - Compliance re-
quirements (e.g., credit monitoring for affected customers) - Mandatory breach
disclosures
e) Competitive Disadvantage: - Loss of intellectual property - Competitive
information stolen - Customer lists stolen - Product designs compromised
7
2.2.3 Societal/Systemic Impact:
a) Erosion of Trust: - Public skepticism about data security - Reduced will-
ingness to share data online - Decreased digital adoption
b) Systemic Risk: - Critical infrastructure attacks (power grids, hospitals) -
National security concerns - Economic disruption
c) Widening Inequality: - Marginalized communities more vulnerable to
misuse - Socioeconomic discrimination through data - Health disparities through
biased health data
2.3 INTRODUCTION TO CRYPTOGRAPHY
Definition: Cryptography is the practice of converting information into a code
to prevent unauthorized access. It involves mathematical algorithms to encrypt
(encode) and decrypt (decode) data.
Fundamental Concepts:
a) Plaintext: Original readable information
b) Ciphertext: Encrypted, unreadable information
c) Encryption: Process of converting plaintext to ciphertext
d) Decryption: Process of converting ciphertext back to plaintext
e) Key: Secret information used in encryption/decryption process
f) Algorithm: Mathematical procedure for encryption/decryption
2.4 SYMMETRIC ENCRYPTION
Definition: Uses same key for both encryption and decryption. Key must be
kept secret and shared securely between parties.
Advantages: - Fast processing speed - Computationally efficient - Suitable for
large data volumes - Established standards
Disadvantages: - Key distribution problem (how to securely share key?) -
Doesn’t provide non-repudiation (sender can deny sending) - Not scalable for
many-to-many communication - Single compromised key affects all encrypted
data
8
2.4.1 Common Symmetric Algorithms 1. DES (Data Encryption
Standard) - OUTDATED - Block size: 64 bits - Key size: 56 bits - Status:
DEPRECATED (too weak for modern use) - Vulnerable to brute-force attacks
2. Triple DES (3DES) - LEGACY - Applies DES encryption three times
- Effective key size: 112-168 bits - Status: Still used but being phased out -
Slower than modern alternatives
3. AES (Advanced Encryption Standard) - STANDARD - Block size:
128 bits - Key sizes: 128, 192, or 256 bits - Status: CURRENT STANDARD
(NIST approved) - Fast and secure - Used in: Government (NSA Suite B),
corporate security, SSL/TLS, disk encryption
Block Cipher Modes (how blocks are encrypted): - ECB (Electronic
Codebook): Simple but weak; same plaintext block = same ciphertext block
- CBC (Cipher Block Chaining): Each block depends on previous; more
secure - CTR (Counter): Uses counter; allows parallel processing - GCM
(Galois/Counter Mode): Provides authentication; modern standard
4. ChaCha20 - MODERN - Stream cipher - Key size: 256 bits - Designed
for mobile devices - Used in: TLS, WhatsApp, Signal
2.5 ASYMMETRIC ENCRYPTION (PUBLIC KEY CRYPTOGRA-
PHY)
Definition: Uses pair of mathematically related keys: - Public Key: Shared
openly; used to encrypt data - Private Key: Kept secret; used to decrypt data
Key Property: Data encrypted with public key can only be decrypted with
corresponding private key (and vice versa).
Advantages: - Solves key distribution problem (no need to share secret keys)
- Enables digital signatures (authentication and non-repudiation) - Scalable for
many-to-many communication - Each user has unique key pair
Disadvantages: - Computationally slower than symmetric encryption - Not
practical for large data volumes - Requires robust key management infrastruc-
ture - Larger ciphertext size
2.5.1 Common Asymmetric Algorithms 1. RSA (Rivest-Shamir-
Adleman) - Based on difficulty of factoring large prime numbers - Key sizes:
1024, 2048, 4096 bits (2048+ recommended) - Usage: Email encryption (PGP),
TLS/SSL, digital signatures, key exchange - Secure: 2048-bit RSA considered
secure until 2030
How RSA Works (Simplified): 1. Choose two large prime numbers (p, q)
2. Calculate n = p × q 3. Calculate �(n) = (p-1)(q-1) 4. Choose e (public
exponent) where 1 < e < �(n) 5. Calculate d (private exponent) where (e × d)
9
mod �(n) = 1 6. Public key = (e, n); Private key = (d, n) 7. Encryption: C =
M^e mod n 8. Decryption: M = C^d mod n
2. ECC (Elliptic Curve Cryptography) - Based on elliptic curve mathe-
matics - Smaller key sizes for equivalent security (256-bit ECC � 3072-bit RSA) -
Faster processing than RSA - Usage: TLS 1.3, cryptocurrency (Bitcoin), mobile
devices - Emerging standard for high-security applications
3. Diffie-Hellman (DH) - Key exchange protocol, not encryption - Allows
two parties to establish shared secret over insecure channel - Foundation for
many secure protocols - Variant: ECDH (using elliptic curves)
2.6 HASHING AND DIGITAL SIGNATURES
2.6.1 Hashing Definition: Cryptographic hash function converts input data
to fixed-size string of bytes (hash/digest). One-way function (cannot reverse).
Properties of Cryptographic Hash Functions: 1. Deterministic: Same
input always produces same output 2. One-way: Computationally infeasible
to reverse (find input from hash) 3. Collision-resistant: Two different inputs
should not produce same hash 4. Avalanche effect: Small change in input
completely changes hash 5. Fast: Quick to compute 6. Fixed output: Same
size regardless of input size
Common Hash Algorithms:
Algorithm Output Size Status Use Case
MD5 128 bits BROKEN Legacy only; NOT secure
SHA-1 160 bits DEPRECATED Phased out; no new use
SHA-256 256 bits STANDARD Current standard; blockchain
SHA-512 512 bits STANDARD Higher security
SHA-3 Variable STANDARD Latest (2015); alternative
BLAKE2 256-512 bits MODERN Fast; for modern applications
Hashing Use Cases:
a) Password Storage: - Never store plaintext passwords - Store hash of pass-
word - During login: hash entered password, compare with stored hash - At-
tacker cannot derive password from hash
b) Data Integrity Verification: - Compute hash of data - Send hash through
secure channel - Recipient recomputes hash, verifies match - Detects if data
corrupted/modified in transit
c) Digital Signatures (discussed below): - Hash data, then encrypt hash
with private key - Provides authentication and non-repudiation
10
d) Blockchain: - Hash of previous block included in current block - Creates
chain; any modification detected
e) File Downloads: - Websites provide hash of downloaded files - Users verify
hash to ensure authentic/uncorrupted download
2.6.2 Digital Signatures
Definition: Cryptographic technique that verifies authenticity, integrity, and
non-repudiation of digital messages/documents.
How Digital Signatures Work:
Signing Process:
1. Sender hashes the message (M)
2. Sender encrypts hash with their PRIVATE key
3. Sender sends: Original message (M) + Encrypted hash (signature)
Verification Process:
1. Recipient receives: Message + Signature
2. Recipient hashes the received message
3. Recipient decrypts signature using sender's PUBLIC key
4. Recipient compares:
- Hash computed from message
- Hash from decrypted signature
5. If match → Message authentic & unmodified; Sender identity verified
6. If no match → Message tampered or not from sender
Digital Signature Properties:
1. Authentication: Verifies sender’s identity (only sender has private key)
2. Integrity: Proves message wasn’t altered (hash would change)
3. Non-repudiation: Sender cannot deny signing (private key involved)
4. Timestamp (optional): Proves when document was signed
Use Cases:
a) Email Authentication: - Sign emails to prove sender identity - Verify
email not spoofed/modified - PGP/GPG email signatures
b) Legal/Contract Signing: - Legally binding digital signatures - Legal en-
forceability in many jurisdictions - E-signatures for contracts
c) Software Code Signing: - Developers sign code/executables - Users verify
code authentic (not malware) - Builds trust in software
d) Digital Certificates: - CAs sign certificates binding public key to identity
- SSL/TLS certificates for websites - Email certificates
e) Blockchain Transactions: - Users sign transactions with private key -
Network verifies with public key - Prevents transaction tampering
11
2.7 ENCRYPTION IN PRACTICE: TRANSPORT AND STORAGE
Transport Security (Encrypting Data in Transit):
SSL/TLS (Secure Sockets Layer/Transport Layer Security): - Encrypts
data between client and server (e.g., HTTPS) - Uses combination of asymmetric
(key exchange) and symmetric (data encryption) encryption - Current standard:
TLS 1.3 - Certificates verify server identity
VPN (Virtual Private Network): - Encrypts all traffic through secure tun-
nel - Hides user IP and location - Protects on public WiFi
Email Encryption: - S/MIME or PGP/GPG - Encrypts email messages and
attachments
Storage Security (Encrypting Data at Rest):
Full Disk Encryption: - BitLocker (Windows), FileVault (macOS), LUKS
(Linux) - Entire disk encrypted - Decrypted on boot with password/key
File/Folder Encryption: - Encrypt specific files/folders - EFS (Windows),
eCryptfs (Linux) - Fine-grained control
Database Encryption: - Encrypt sensitive columns - Transparent Data En-
cryption (TDE) - Application-level encryption
Cloud Storage Encryption: - Encrypt data before uploading - Client-side
encryption (user controls key) - Provider-side encryption (provider controls key)
UNIT 3: DATA COLLECTION, USE AND REUSE (8
hours)
3.1 HARMS ASSOCIATED WITH DATA COLLECTION, USE,
AND REUSE
Definition: Data collection harms occur when gathering personal data violates
privacy, autonomy, dignity, or creates risks for individuals.
3.1.1 Direct Harms from Collection 1. Privacy Intrusion: - Unautho-
rized collection of sensitive personal information - Surveillance without knowl-
edge/consent - Example: Phone location tracking without consent
2. Unauthorized Disclosure: - Collection leading to data breach exposure -
Sensitive information made public - Example: Medical records leaked online
12
3. Identity Theft Risk: - Collected information used for fraud - Financial
accounts opened in person’s name - Example: Social Security number harvested
and misused
4. Discrimination Based on Collected Data: - Collected data used to
discriminate - Insurance rates based on genetic data - Job discrimination based
on financial history
5. Manipulation and Targeting: - Collection enables targeted manipulation
- Microtargeted political ads - Behavioral advertising exploiting vulnerabilities
3.1.2 Harms from Unauthorized Use/Reuse 1. Purpose Creep: -
Data collected for one purpose used for different purpose - Example: Location
data collected for GPS navigation used for ad targeting - No new consent ob-
tained
2. Secondary Use Harms: - Original purpose harmless; secondary use harm-
ful - DNA collected for paternity testing used for criminal investigation - Health
data collected for treatment used for insurance discrimination
3. Re-identification Risks: - Anonymized data re-identified when combined
with other datasets - Example: Netflix Prize dataset re-identified via IMDb -
Privacy of anonymized data compromised
4. Algorithmic Harm: - Data used in automated decision-making with biased
outcomes - Credit scoring algorithms discriminating against minorities - Hiring
algorithms screening out qualified candidates
5. Chilling Effect: - People avoiding activities/services to prevent collection -
Avoiding healthcare due to privacy concerns - Suppressing free speech to avoid
surveillance
3.1.3 Aggregate/Systemic Harms 1. Surveillance Chilling Effect: -
Knowing data collected and analyzed suppresses behavior - People self-censor
due to surveillance awareness - Reduced autonomy and freedom
2. Discrimination at Scale: - Discriminatory patterns in automated systems -
Affect millions via algorithms - Example: COMPAS recidivism algorithm biased
against Black defendants
3. Inequality Amplification: - Data collection concentrates power with
tech companies - Creates information asymmetry - Marginalizes less tech-savvy
populations
4. Democratic Threats: - Targeted political manipulation via collected data -
Micro-targeted disinformation campaigns - Undermines democratic deliberation
13
5. Loss of Anonymity/Public Space: - Ubiquitous data collection elimi-
nates anonymous spaces - Cannot move through public without being tracked -
Loss of freedom to be alone in public
3.2 INTRODUCTION TO DATA ANONYMIZATION
Definition: Data anonymization is the process of removing or encrypting per-
sonally identifiable information from data to prevent identification of individuals
while preserving data utility for analysis.
Key Distinction: - Anonymization: Data cannot be traced back to individ-
ual (permanent, irreversible) - Pseudonymization: Data can be traced back
with additional information (reversible) - De-identification: Removing direct
identifiers (may be reversible if combined with other data)
Goal of Anonymization: - Enable data analysis/research while protecting
privacy - Balance: Privacy protection vs. Data utility - Data should be useful
but individual identity hidden
3.3 DATA ANONYMIZATION TECHNIQUES
3.3.1 Direct Identifier Removal Definition: Removing directly identify-
ing information
Direct Identifiers to Remove: - Name, address, email, phone number -
Social security number, passport number, driver’s license - Date of birth (if
combined with other data, can identify) - IP address, device ID, cookie identifiers
- Biometric data, photos/video
Process: 1. Identify all direct identifiers in dataset 2. Remove or encrypt them
3. Replace with random codes/pseudonyms 4. Retain non-identifying attributes
for analysis
Example:
Original:
Name: John Smith, DOB: 1980-05-15, Address: 123 Main St, Age: 42, Income: $75,000
Anonymized:
PersonID: 47284, Age_Group: 40-50, Income_Range: $75K-$100K, Region: Northeast
Limitations: - Removing data reduces analysis capability - Quasi-identifiers
still present (age, gender, location can re-identify) - Different jurisdictions have
different definitions of direct identifiers
14
3.3.2 Generalization (Coarsening) Definition: Reducing precision of
data by grouping values into ranges
Techniques:
a) Age Generalization:
Original: 23 years old
Generalized: 20-29 age group
Further: Adult
Provides analysis of age patterns without exposing exact age
b) Location Generalization:
Original: Latitude 40.7128, Longitude -74.0060
Generalized: New York City
Further: New York State
Further: Northeast USA
c) Temporal Generalization:
Original: 2024-12-18 11:02 PM
Generalized: December 2024
Further: Q4 2024
d) Income Generalization:
Original: $73,452
Generalized: $50K-$100K
Further: Middle Income
Advantages: - Reduces re-identification risk - Retains useful information for
analysis - Flexible levels of anonymization
Disadvantages: - Reduces data granularity - May lose important patterns -
Trade-off between privacy and utility
3.3.3 Suppression/Masking Definition: Removing or masking specific val-
ues or records
Techniques:
a) Value Suppression: - Replace specific values with asterisks or blanks -
Example: Phone: 555-***-1234 (only show partial)
b) Record Suppression: - Remove entire records if too identifying - Remove
rare values that could identify - Example: If only one person in dataset with
certain profession + location
15
c) Cell Masking: - Replace cell values with symbols: *, [blank], [masked] -
Indicates data suppressed without showing original
Example:
Original:
John Smith, Rare Disease Diagnosis, Employer: XYZ Inc, Salary: $200K
Suppressed:
ID: 8847, Disease: [MASKED], Employer: [MASKED], Salary Range: $150K-$300K
Rare values (disease, high salary with rare employer) removed to prevent re-identification
Disadvantages: - Loses information completely - May not provide enough
context - Over-suppression makes data useless
3.3.4 Perturbation (Adding Noise) Definition: Adding random noise to
data values to obscure true values while maintaining statistical properties
Techniques:
a) Additive Noise:
Original: Age 42
Add random noise: +3
Noisy: Age 45
Analysis still valid (statistical properties preserved) but true value hidden
b) Multiplicative Noise:
Original: Salary $80,000
Multiply by 1.05 (5% noise)
Noisy: Salary $84,000
c) Differential Privacy: - Mathematical framework ensuring privacy while
enabling analysis - Adds calibrated noise to query results - Mathematical guar-
antee: adversary cannot determine if individual in dataset
Advantages: - Retains statistical accuracy - Enables statistical analysis - Pri-
vacy guarantee with differential privacy
Disadvantages: - Utility reduced if too much noise - Calibrating noise level is
complex - Not suitable for identifying specific records
3.3.5 Aggregation (Grouping) Definition: Combining individual records
into groups and reporting only group-level statistics
16
Example:
Individual level data:
[Name: John, Age: 42, Salary: $80K, Department: Sales]
[Name: Jane, Age: 38, Salary: $75K, Department: Sales]
[Name: Bob, Age: 45, Salary: $95K, Department: Sales]
Aggregated:
Sales Department: 3 employees, Avg Age: 41.7, Avg Salary: $83.3K
Individual identities hidden; only group statistics visible
Use Cases: - Census data reporting - Public health statistics - Departmental
statistics in organizations
Limitations: - Not suitable for individual-level analysis - May not be useful
for small groups - Data utility greatly reduced
3.3.6 Differential Privacy (Advanced) Definition: Mathematical frame-
work providing privacy guarantee: results of analysis should be almost same
whether or not specific individual’s data included
How It Works: 1. Query database for statistical result (e.g., average salary) 2.
Add calibrated noise to result 3. Output noisy result 4. Mathematical guarantee:
attacker cannot distinguish if individual included
Epsilon (�) Parameter: - Smaller � = more privacy (more noise) - Larger � =
more utility (less noise) - � = 1 considered good privacy, � = 0.1 strong privacy
Applications: - Census data (US Census Bureau uses differential privacy) -
Location data (Google Federated Learning uses DP) - Medical research - Social
media analytics
Advantages: - Rigorous mathematical privacy guarantee - Proven security
against powerful adversaries - Enables analysis on sensitive data
Limitations: - Complex to implement correctly - Requires tuning � parameter
- Reduces query accuracy for very specific questions
3.4 CHALLENGES IN DATA ANONYMIZATION
3.4.1 Re-identification Risks 1. Linkage/Record Linkage Attack: -
Combining anonymized data with other public datasets - Example: Netflix Prize
anonymized data + IMDb data → Users re-identified - Quasi-identifiers (age,
ZIP code, gender) used for linking
17
2. Attribute Disclosure: - Cannot determine individual identity but infer
sensitive attribute - Example: Knowing someone in dataset has age 42, lives in
small town → Likely disease from aggregated health data - Attribute disclosed
without identifying individual
3. Membership Disclosure: - Attackers determine whether individual’s data
in dataset - Privacy breach even if identity not revealed - Example: “Person with
condition X in hospital database”
4. Compositional Attacks: - Multiple anonymous queries reveal information
together - Individual queries safe but combination leaks info - Example: Average
age 40, females average age 42 → Can infer male’s approx age
3.4.2 Balancing Privacy and Utility Challenge: More aggressive
anonymization = more privacy but less useful data
Trade-off: - Heavy anonymization: Highly private but limited analysis possible
- Light anonymization: More useful but greater re-identification risk - Middle
ground: Achieve both but imperfect
Example:
Data Protection Level vs. Usability:
Highly Anonymized:
ID, Age_Group, Region
Privacy: High �
Utility: Low �
Analysis: Only broad patterns possible
Moderately Anonymized:
ID, Age (range), City, Employment_General
Privacy: Medium �
Utility: Medium �
Analysis: Reasonable patterns discoverable
Minimally Anonymized:
ID, DOB, Exact Address, Exact Job Title
Privacy: Low �
Utility: High �
Analysis: Detailed analysis possible but easy to re-identify
3.4.3 Emerging Data Sources 1. Biometric Data: - Fingerprints, facial
recognition cannot be anonymized (unique, irreversible) - Even if encrypted,
same biometric always looks same (tracking risk) - Challenge: Not truly
anonymizable
2. Behavioral Data/Clickstreams: - Mouse movements, typing patterns,
18
browsing history create unique fingerprints - Can re-identify even if demographic
data removed - Growing tracking capability
3. IoT and Location Data: - Precise location history creates unique move-
ment patterns - Few people follow same exact path at same times - Highly
re-identifiable
4. Genetic Data: - Close relatives’ genetic data can reveal personal genetic
information - Cannot be completely anonymized (genetic link permanent) - Chal-
lenges shared by family
3.4.4 Technical Challenges 1. Data Heterogeneity: - Different data
types (text, numbers, images, video) have different anonymization approaches
- Unstructured data (text, images) harder to anonymize - Challenge: Develop
techniques for diverse data types
2. Evolving Adversary Knowledge: - Background knowledge improves over
time - Today’s safe anonymization may be broken with future data/techniques
- Example: Anonymization 2010 broken by 2015 with new datasets
3. Scalability: - Anonymizing large datasets is computationally expensive -
Ensuring consistency across distributed data challenging - Challenge: Efficient
anonymization of big data
4. Real-time Data Challenges: - Streaming data makes batch anonymiza-
tion impossible - Need to anonymize at collection point - Challenge: Real-time
anonymization maintaining utility
3.4.5 Governance and Regulatory Challenges 1. Legal Definitions: -
Different jurisdictions define “anonymized” differently - GDPR considers data
anonymized only if non-reversible - Challenge: Meeting different legal standards
2. Responsibility: - When anonymization insufficient, who liable for re-
identification? - Data creator’s responsibility vs. data user’s responsibility? -
Not clearly defined in regulations
3. Consent Requirements: - If anonymization fails, secondary use requires
consent - But keeping consent records defeats anonymization - Challenge: Bal-
ancing consent and anonymization
19
UNIT 4: ETHICAL CONSIDERATIONS IN DATA PRI-
VACY (5 hours)
4.1 PRIVACY AND SURVEILLANCE
Definition of Surveillance: Systematic monitoring and observation of indi-
viduals’ activities, communications, and movements.
4.1.1 Types of Surveillance 1. Government Surveillance: - Mass
surveillance: Monitor entire populations - Targeted surveillance: Mon-
itor specific individuals/groups - Signal intelligence (SIGINT): Intercept
communications - Law enforcement: Investigation and crime prevention -
National security: Counter-terrorism and counter-intelligence
Examples: - NSA PRISM program (Snowden revelations, 2013): Bulk collec-
tion of internet communications - FBI backdoor demands in encrypted devices
- China’s social credit system: Comprehensive personal monitoring
2. Corporate/Commercial Surveillance: - Data collection by companies
for profit - Behavioral tracking via cookies, pixels, analytics - Location tracking
via mobile apps - Financial profiling via purchases and browsing - Predictive
analytics to anticipate behavior
Examples: - Facebook’s data collection (tracked off-Facebook via pixel) - Ama-
zon tracking across websites via advertising - Google’s location tracking even
when location services off - Telecom companies selling location data
3. Workplace Surveillance: - Employee monitoring during work hours -
Email monitoring and content inspection - Keystroke logging and screen record-
ing - GPS tracking of company vehicles - Monitoring of personal device usage
on company network
4. Financial Surveillance: - Banking transaction monitoring - Credit score
tracking and reporting - Cryptocurrency transaction tracking - Insurance com-
pany data collection - Loan application data profiling
5. Health Surveillance: - Electronic health record systems - Insurance com-
pany monitoring of patient data - Genetic testing and genetic database creation
- Fertility clinic data collection - Mental health record accessibility
6. Biometric Surveillance: - Facial recognition systems in public spaces -
Fingerprint scanning at borders and workplaces - Iris/retina scanning - Gait
recognition - Thermal imaging
Examples: - China’s facial recognition cameras monitoring Uyghurs -
Clearview AI scraping billions of photos for facial recognition - TSA facial
recognition at airports
20
4.1.2 Privacy vs. Security Trade-off The Argument for More Surveil-
lance (Security): - Detect and prevent crimes proactively - Identify and stop
terrorists before attacks - Enhance public safety in shared spaces - Help law
enforcement solve crimes - National security against external threats
The Argument for Less Surveillance (Privacy): - People have fundamen-
tal right to privacy - Surveillance chills free speech and expression - Surveillance
enables government/corporate abuse - Innocent people harmed by surveillance
(mistakes, misuse) - Power imbalance: Observed cannot see observers - Surveil-
lance not proven as crime deterrent - False sense of security vs. actual security
Historical Examples of Surveillance Abuse: - COINTELPRO (1956-
1971): FBI illegally surveilled civil rights activists, used information to discredit
leaders - Stasi (East Germany, 1950-1990): Secret police mass surveillance
of citizens, repression of dissidents - Watergate (1972): Government surveil-
lance used for political purposes - Myanmar military (2021): Surveillance
used to suppress opposition to coup
Finding Balance: - Transparent surveillance: Rules, oversight, public
knowledge - Limited surveillance: Narrow scope, specific purposes only -
Oversight mechanisms: Independent review of surveillance programs - Con-
sent and opt-out: Individuals informed and able to refuse - Proportionality:
Surveillance justified by actual benefit - Accountability: Those conducting
surveillance held responsible
4.2 ETHICS OF DATA COLLECTION AND USE
4.2.1 Fundamental Ethical Principles 1. Autonomy and Consent: -
Ethical principle: Individuals have right to control their personal information -
Practical requirement: Informed, voluntary, specific consent - Challenge: Most
people don’t read privacy policies; consent is assumed rather than meaningful -
Ethical concern: “Forced” consent if only way to access service
Ethical Requirements: - Clear, understandable information about data col-
lection - Genuine choice (not all-or-nothing) - Separate consent for different
purposes - Easy withdrawal of consent - No penalty for refusing
2. Respect for Persons: - Treat people as ends, not means - Don’t use people
as merely tools for your purposes - Respect dignity and autonomy
Ethical Violations: - Manipulative data collection (deceptive collection meth-
ods) - Using personal data to manipulate behavior - Profiling without individ-
ual’s knowledge - Using vulnerable people’s data without protection
3. Beneficence and Non-maleficence: - Beneficence: Maximize benefits
from data use - Non-maleficence: Minimize harms from data use
Ethical Balance: - Data collection should benefit individual or society more
21
than harm - Health research using patient data: Possible cures (benefit) vs. pri-
vacy intrusion (harm) - Should err on side of minimizing harm
4. Justice and Fairness: - Equal treatment in data-driven decisions - No
discrimination in algorithmic decision-making - Fair distribution of benefits and
burdens - Special protection for vulnerable populations
Ethical Concerns: - Algorithmic bias affecting marginalized groups - Data
deserts: Some groups have no data, fall outside analysis - Winner-take-all dy-
namics in digital economy
4.2.2 Contextual Integrity and Norms Concept: Privacy depends on
“appropriate information flow” in different contexts
Context-Based Norms: - Medical context: Doctors-patients share health
information; norm is confidentiality - Financial context: Banks-customers
share financial information; norm is security - Social context: Friends share
personal information; norm is not spreading it publicly - Commercial context:
Retailers-customers exchange money for goods; norm is limited data collection
Ethical Violations - Norm Violation: - Health data (medical context) sold
to insurers (commercial context) violates norms - Social data (friend context)
used by employers (employment context) violates norms - Location data (per-
sonal context) used for mass surveillance (government context) violates norms
Principle: - Personal data can be shared in original context (medical privacy
with doctor) - Moving data to new context without consent violates privacy
norms - Even if anonymized, context violation raises ethical concerns
4.2.3 Ethical Problems in Data Collection 1. Deceptive Collection:
- Collecting data without user knowledge - Hidden tracking (cookies, pixels,
location) - Misrepresenting what data collected/used for
Ethical Problem: Violates autonomy; people cannot consent to what they
don’t know about
2. Collection from Vulnerable Populations: - Children (cannot properly
consent) - Low-income individuals (forced choice due to financial need) - Un-
documented immigrants (fear of authorities) - Mental health patients (reduced
autonomy during crisis)
Ethical Problem: Vulnerable people cannot provide meaningful consent; their
data misused more easily
3. Aggregate Harm: - Individual data collection seems harmless (one person’s
location) - Combined data collected from millions reveals patterns - Aggregate
creates harm beyond individual perception
22
Ethical Problem: Individuals cannot assess harm from their data alone; ag-
gregate harm not apparent
4. Purpose Creep: - Data collected for one purpose (navigation) - Used for
different purposes (targeting ads, law enforcement)
Ethical Problem: Even if consent given originally, secondary uses not con-
sented to; context changes
5. Commodification of Data: - Personal data treated as product to buy/sell
- People’s information used to profit without compensation - Data about person
sold to others
Ethical Problem: Body/information exploited for corporate profit; individu-
als have no benefit
4.3 BIAS AND DISCRIMINATION IN DATA ANALYSIS
4.3.1 Sources of Bias 1. Data Bias:
a) Historical Bias: - Historical discrimination encoded in training data - Ex-
ample: Hiring algorithm trained on past hiring decisions favoring men → Dis-
criminates against women in future - Historical injustice perpetuated in future
decisions
b) Representation Bias: - Some groups underrepresented in data - Algo-
rithm performs worse on underrepresented groups - Example: Facial recognition
trained mostly on white faces → Less accurate on dark-skinned faces
c) Measurement Bias: - How variables measured introduces bias - Example:
Income data from loan applications biased toward those who apply for loans
(excludes low-income who don’t apply) - Example: Health data from wealthy
neighborhoods more complete than poor neighborhoods
d) Sampling Bias: - Data collected from biased sample of population - Ex-
ample: Opinion surveys done online exclude those without internet access -
Example: Clinical trials mostly white participants → Drug efficacy unclear for
other races
2. Algorithmic Bias:
a) Proxy Variables: - Algorithm uses variable as proxy for protected char-
acteristic - Example: ZIP code used in algorithm; ZIP code is proxy for race -
Algorithm then indirectly discriminates based on race
b) Correlation vs. Causation: - Algorithm identifies correlation that doesn’t
imply causation - Example: Algorithm finds correlation between name and cred-
itworthiness; correlation due to discrimination, not causation - Algorithm per-
petuates discrimination
23
c) Threshold Bias: - Different thresholds applied to different groups - Ex-
ample: Job interview system calls top candidates; calls more candidates from
favored group to reach quota - Different standards applied to different groups
3. Human Bias:
a) Confirmation Bias: - Model builders confirm their existing biases in model
- Example: Assuming certain group “naturally” less qualified; model reflects this
bias
b) Apophenia: - Seeing meaningful patterns in random data - Example: As-
suming algorithm’s output is objective when it reflects human bias
4.3.2 Algorithmic Discrimination - Real-World Examples 1. COM-
PAS (Correctional Offender Management Profiling for Alternative
Sanctions): - Recidivism prediction algorithm used in criminal justice - Pur-
pose: Predict who likely to re-offend - Issue: Biased against Black defendants
- Study (ProPublica, 2016): Black defendants marked high-risk at rate 45%
vs. white defendants 23% - Even when controlling for criminal history, racial
bias present - Consequence: Black defendants receive harsher sentences
2. Amazon Recruiting Algorithm: - Machine learning algorithm to screen
job resumes - Trained on 10 years of hiring history (male-dominated tech) -
Issue: Systematically downgraded female applicants - Algorithm learned male
= better for tech jobs (historical bias) - Consequence: Fewer women invited to
interviews
3. STEM Hiring Algorithms: - Algorithm analyzing resumes and determin-
ing fit for STEM roles - Issue: Female-coded words (e.g., “women’s”) associated
with lower scores - Algorithm discriminated based on gender signals
4. Healthcare Algorithms: - Algorithm allocating healthcare resources based
on cost (proxy for need) - Issue: Black patients use fewer healthcare resources
on average (access issues) - Algorithm thought Black patients needed less care
- Consequence: Black patients received lower-quality care
5. Mortgage/Credit Algorithms: - Algorithm determines loan approval
and interest rates - Issue: Minority groups offered worse terms at higher rates -
Algorithm learned from historical redlining data
4.3.3 Impact of Algorithmic Discrimination On Individuals: - Denied
loans, jobs, housing, education due to algorithm - No understanding why de-
nied (algorithm is “black box”) - No recourse or appeal possible - Perpetuates
disadvantage (one denial → harder to get future opportunity)
24
On Groups: - Entire demographic groups systematically disadvantaged - Re-
inforces historical stereotypes and prejudices - Widening inequality gaps - Con-
centration of disadvantage
On Society: - Discrimination at scale and speed (affects millions) - Legitimacy
given through “objectivity” of algorithm - Discrimination harder to detect and
challenge (hidden in algorithm) - Erosion of equal opportunity principle
4.3.4 Fairness and Mitigation Strategies Fairness Definitions (Dif-
ferent Perspectives):
1. Statistical Parity: - Algorithm should have equal outcomes across groups -
Same percentage selected from each group - Issue: Doesn’t account for different
group qualifications
2. Equalized Odds: - Algorithm should have equal false positive and false
negative rates across groups - Equal type 1 and type 2 errors for all groups -
Issue: May require different thresholds for different groups
3. Equal Opportunity: - Algorithm should not discriminate against protected
groups - Same probability of positive outcome for qualified individuals across
groups - Issue: Difficult to implement with imperfect data
4. Calibration: - Predictions equally accurate across groups - Algorithm
equally confident in predictions for all groups - Issue: May not eliminate dis-
parate impact
Mitigation Strategies:
1. Data Level Mitigation: - Collect representative data (include underrep-
resented groups) - Reduce bias in training data - Augment data for underrepre-
sented groups - Use de-biasing techniques
2. Algorithm Level Mitigation: - Use fairness-aware algorithms - Add
fairness constraints to optimization - Use multiple fairness metrics in testing -
Conduct disparate impact analysis - Set different thresholds for different groups
(if legally permissible) - Feature selection to remove proxy variables
3. Process Level Mitigation: - Human oversight of algorithmic decisions -
Explainability: Understand why algorithm made decision - Transparency: Dis-
close use of algorithm - Auditability: Regular testing for bias and discrimination
- Appeal process: Allow people to challenge decisions - External audits: Inde-
pendent fairness audits
4. Governance Level Mitigation: - Regulate high-risk algorithmic systems
- Require fairness impact assessments - Mandate human review for critical de-
25
cisions - Transparency and explainability requirements - Right to explanation
and recourse - Liability for discriminatory algorithmic outcomes
4.3.5 Challenges in Achieving Fairness 1. Fairness Tradeoffs: - Cannot
simultaneously satisfy all fairness metrics - Improving one metric may worsen
another - No single “correct” fairness definition
2. Context Dependency: - Appropriate fairness definition depends on con-
text - Criminal justice: May want lower false positive (don’t wrongly imprison)
- Healthcare: May want lower false negative (don’t miss people needing care)
3. Data Limitations: - Cannot fix bias in data with algorithm alone - If
historical discrimination in data, algorithm reflects it - Need to address root
causes (historical injustice)
4. Defining Protected Groups: - What characteristics are “protected”? -
Different in different jurisdictions - Some characteristics (ZIP code as proxy for
race) protected indirectly
5. Legal Ambiguity: - Disparate impact (outcome discrimination) legal status
unclear - Different standards in different jurisdictions - Companies uncertain
what’s legally required
COMPREHENSIVE SUMMARY TABLE
Topic Key Concepts Application
Privacy Notion Control over personal Individual right to
data, PII, autonomy manage information
about self
Historical Context From Warren-Brandeis Understanding
to GDPR/DPDP evolution of privacy
protections
Sensitive Data Biometric, health, Classification and
financial, behavioral protection levels
Privacy Laws GDPR, DPDP 2023, Compliance and legal
CCPA, HIPAA obligations
Privacy Principles Transparency, consent, Framework for
minimization, security privacy-respecting
practices
Attacks/Breaches Hacking, malware, Defense and prevention
social engineering measures
Breach Impacts Identity theft, financial Risk assessment and
loss, discrimination mitigation
26
Topic Key Concepts Application
Cryptography Basics Encryption, keys, Data protection
algorithms fundamentals
Symmetric AES, 3DES, ChaCha20, Efficient large-data
Encryption block modes encryption
Asymmetric RSA, ECC, key Scalable secure
Encryption exchange communication
Hashing SHA-256, properties, Data integrity and
collision-resistance password security
Digital Signatures Authentication, Authenticity and
non-repudiation, accountability
verification
Collection Harms Privacy intrusion, Minimizing data
identity theft, collection risks
manipulation
Anonymization Removal, generalization, Privacy-preserving data
suppression, analysis
perturbation
Re-identification Linkage attacks, Limitations of
Risks quasi-identifiers anonymization
Privacy Transparency, oversight, Balancing security and
vs. Surveillance proportionality privacy
Ethical Principles Autonomy, respect, Ethical framework for
beneficence, justice data practices
Bias in Algorithms Historical bias, Fairness and mitigation
representation bias, strategies
discrimination
KEY FORMULAS AND DEFINITIONS
Cryptography: - RSA: C = M^e mod n; M = C^d mod n - Hash property:
H(x) = H(y) only if x = y (collision-resistant) - Symmetric key requirement:
Same key for encryption and decryption - Asymmetric key property: Private
key pair cannot be derived from public key (computational difficulty)
Anonymization Quality: - Information loss vs. Privacy gain trade-off - Re-
identification risk = f(quasi-identifiers, background knowledge, time) - Differen-
tial privacy: � parameter controls privacy-utility trade-off
27
EXAM TIPS AND STRATEGIES
High-Frequency Topics (Likely Exam Questions):
1. Definition and distinction: Privacy vs. Security, Anonymization
vs. Pseudonymization
2. Regulations: GDPR principles, DPDP 2023 key features, Indian privacy
framework
3. Cryptography: AES vs. RSA differences, hashing properties, digital
signature process
4. Anonymization techniques: When to use which technique, limitations
5. Real-world examples: Breaches, biased algorithms, surveillance cases
6. Ethical principles: Autonomy, consent, fairness in algorithms
7. Impact assessment: Consequences of breaches, algorithmic discrimina-
tion harms
Study Strategy for Full Marks:
� Understand concepts deeply - Not just definitions; understand WHY and
HOW
� Know applications - How concepts applied in real scenarios
� Learn with examples - Every concept has real-world example; use them to
illustrate understanding
� Compare and contrast - GDPR vs. DPDP, Symmetric vs. Asymmetric
encryption, etc.
� Remember case studies - COMPAS algorithm, Amazon hiring algorithm,
Equifax breach - these appear in exams
� Understand ethical dimensions - Exams test moral reasoning, not just
technical knowledge
� Balance details - Know technical details AND broader implications
� Practice definitions - Precise definitions score points; sloppy definitions lose
marks
Last Minute Checklist Before Exam:
□ Can explain notion of privacy and PII clearly
□ Know GDPR and DPDP principles and differences
□ Understand AES and RSA encryption (when to use each)
□ Can explain hash functions and digital signatures
□ Know 4+ anonymization techniques with examples
□ Understand re-identification risks and limitations of anonymization
□ Can explain surveillance trade-offs with arguments for both sides
□ Know major algorithmic bias cases (COMPAS, Amazon hiring, etc.)
28
□ Understand fairness definitions and mitigation strategies
□ Can trace ethical concerns through data lifecycle (collection → use →
reuse)
GOOD LUCK! This is comprehensive material covering all exam top-
ics. Study methodically, understand not memorize, and apply con-
cepts to examples. You’re prepared for full marks!
29