BIG DATA ANALYSIS PYQ AIDS
Module 1: Introduction to Big Data & Hadoop
• 1.1 Introduction to Big Data, 1.2 Big Data characteristics
o Q1 a) What is Hadoop and Why it Matters. - Dec 2023
o Q1 b) Compare traditional database and big data. - Dec 2023
o Q1 (a) Explain 5 V’s of big data. - Nov 2024
o Q1 a) Big Data and its characteristics - May 2024
o Q1 (a) Compare traditional data and big data. - Jun 2025
o Q2 (b) Explain big data enabling technologies. - Nov 2024 (Covers Hadoop
Ecosystem from 1.6)
• 1.5 Concept of Hadoop, 1.6 Core Hadoop Components; Hadoop Ecosystem
o Q2 a) Draw Hadoop Ecosystem and briefly explain its components. - Dec 2023
o Q1 (c) Write the limitations of Hadoop. - Nov 2024 (Covers 2.4 Hadoop
Limitations)
o Q4 (a) Explain Hadoop Architectural Model with both components in
detail - Nov 2024
o Q2 a) Explain HDFS architecture. - May 2024 (Covers 2.1)
o Q1 (b) What are the advantages and limitations of Hadoop - Jun 2025
o Q2 a) Draw Hadoop Ecosystem and briefly explain its components - Jun 2025
REPEATED QUESTIONS LIST
High Priority (Repeated 3-4 times)
• Hadoop Ecosystem & Components - Explain/Draw and explain the Hadoop
Ecosystem and its components. (Jun 2025, Dec 2023, Nov 2024)
• Big Data Characteristics & Vs. Traditional Data - Explain 5 V's, compare Big Data with
traditional data/database. (May 2024, Jun 2025, Dec 2023, Nov 2024)
Medium Priority (Repeated 2 times)
• Hadoop Limitations - List/Explain the advantages and limitations of Hadoop. (Jun
2025, Nov 2024)
Low Priority (Asked Once)
• Introduction to Hadoop - What is Hadoop and why it matters? (Dec 2023)
• Big Data Enabling Technologies - Explain big data enabling technologies. (Nov 2024)
• HDFS Architecture - Explain HDFS architecture. (May 2024)
Module 2: Hadoop HDFS and Map Reduce
• 2.1 Distributed File Systems (HDFS)
o Q2 a) Explain HDFS architecture. - May 2024
• 2.2 MapReduce: The Map and Reduce Tasks, Execution, Failures
o Q1 d) Compare DBMS VS DSMS. - Dec 2023 (Related to stream processing
context)
o Q1 (d) Explain how failures are handled in Map Reduce job - Nov 2024
o Q1 c) The Map and Reduce Tasks - May 2024
o Q4 a) List the main components of Mapreduce execution pipeline. - Dec 2023
o Q4 (b) Write the functions of the components and execution steps in Map
Reduce - Nov 2024
o Q4 a) List the main components of Map reduce execution pipeline. - May
2024
o Q2 (b) Write the functions of the components and execution steps in Map
Reduce - Jun 2025
• 2.3 Algorithms Using MapReduce: Relational-Algebra Operations, Matrix
Multiplication
o Q2 (a) Illustrate relational algebra operations with example. - Nov 2024
o Q3 a) Write a Map reduce pseudo code to multiply two matrices. - May 2024
o Q3 (a) Explain Selection and Projection algebraic operation using
MapReduce. - Jun 2025
• 2.4 Hadoop Limitations
o Q1 (c) Write the limitations of Hadoop. - Nov 2024
o Q1 (b) What are the advantages and limitations of Hadoop - Jun 2025
REPEATED QUESTIONS LIST:
High Priority (Repeated 3-4 times)
• MapReduce Execution & Steps - List components, explain functions, and detail the
execution steps of MapReduce. (May 2024, Dec 2023, Jun 2025, Nov 2024)
Medium Priority (Repeated 2 times)
• MapReduce Algorithms (Relational Algebra) - Explain/Illustrate Selection, Projection,
and other relational algebra operations using MapReduce.
o (Nov 2024, Jun 2025)
• Handling Failures in MapReduce - Explain how failures are handled in a MapReduce
job.(Nov 2024) [Implied in execution steps of other papers]
Low Priority (Asked Once)
• Matrix Multiplication with MapReduce - Write a MapReduce pseudo code to
multiply two matrices.(May 2024)
• DBMS vs. DSMS - Compare DBMS and DSMS (Data Stream Management System).
(Dec 2023)
Module 3: NoSQL
• 3.1 Introduction to NoSQL, SQL vs NoSQL
o Q1 c) Explain CAP theorem. State how it is different from ACID
properties. - Dec 2023
o Q1 (b) Differentiate between SQL vs NoSQL - Nov 2024
o Q1 (c) Differentiate between SQL vs NoSQL - Jun 2025
• 3.2 NoSQL Data Architecture Patterns
o Q2 b) Explain the four types of NoSQL database. - Dec 2023
o Q3 (b) Compare different types of NoSQL architectural pattern - Nov 2024
o Q2 b) Explain Column family store and Graph Store NoSQL architectural
pattern - May 2024
o Q3 (b) Explain Key-value store and Document Store NoSQL architectural
pattern with example. - Jun 2025
o Q6 (d) Four ways that NoSQL systems handle big data problems. - Jun
2025 (Covers 3.3)
• 3.3 NoSQL solution for big data
o Q6 (d) Four ways that NoSQL systems handle big data problems. - Jun 2025
REPEATED QUESTIONS LIST:
High Priority (Repeated 4 times)
• SQL vs. NoSQL - Differentiate between SQL and NoSQL databases.
o (Jun 2025, Nov 2024, Dec 2023 via CAP Theorem, May 2024 via architectural
patterns)
Medium Priority (Repeated 2-3 times)
• NoSQL Architectural Patterns / Types - Explain and compare the four
types/architectural patterns of NoSQL databases (Key-Value, Document, Column-
Family, Graph).
o (Dec 2023, Jun 2025, May 2024)
Low Priority (Asked Once)
• CAP Theorem - Explain the CAP theorem and how it differs from ACID properties.
o (Dec 2023)
• NoSQL for Big Data - Explain the four ways that NoSQL systems handle big data
problems.
o (Jun 2025)
Module 4: Mining Data Streams
• 4.1 The Stream Data Model: Issues, DSMS
o Q1 d) Compare DBMS VS DSMS. - Dec 2023
o Q5 (a) Write issues in data stream queries. Explain the issues in data
streaming - Nov 2024
o Q3 b) Explain Issues in Data stream query processing - May 2024
o Q4 (a) Draw a neat sketch, explain the architecture of the data-stream
management system - Jun 2025
• 4.3 Filtering Streams: Bloom Filter
o Q1 d) Bloom filter for stream data mining - May 2024
o Q6 (a) Bloom Filter with analysis - Jun 2025
• 4.5/4.6 Counting Frequent Items & Counting Ones (DGIM Algorithm)
o Q3 a) Explain PCY algorithm and its types with neat labeled diagram - Nov
2024
o Q3 b) Explain DGIM algorithm. - Dec 2023
o Q4 b) Explain DGIM algorithm. - May 2024
o Q4 (b) Explain DGIM algorithm for counting ones in a stream with
example - Jun 2025
REPEATED QUESTIONS LIST:
High Priority (Repeated 3 times)
• DGIM Algorithm - Explain the DGIM algorithm for counting ones in a stream.
o (May 2024, Dec 2023, Jun 2025)
Medium Priority (Repeated 2 times)
• Stream Data Model & Issues - Explain the issues in data stream query processing and
the architecture of a Data-Stream Management System (DSMS).
o (May 2024, Jun 2025, Nov 2024)
Low Priority (Asked Once)
• Bloom Filter - Explain the Bloom filter with analysis for stream data mining.
o (May 2024, Jun 2025)
• PCY Algorithm - Explain the PCY algorithm and its types.
o (Nov 2024)
Module 5: Finding Similar Items and Clustering
• 5.1 Distance Measures
o Q6 a) Explain with example two major classes of distance measures. - Dec
2023
o Q1 b) Distance measures for Big Data - May 2024
o Q1 (d) List and explain Distance measures for Big Data - Jun 2025
• 5.2 CURE Algorithm, Stream-Clustering
o Q4 b) Explain cure algorithm. - Dec 2023
o Q6 a) Explain CURE algorithm with its advantages over traditional clustering
algorithm - Nov 2024
o Q6 b) Explain CURE algorithm. - May 2024
o Q6 (b) Cure Algorithm - Jun 2025
REPEATED QUESTIONS LIST:
High Priority (Repeated 3 times)
• CURE Algorithm - Explain the CURE algorithm, its advantages, and how it works.
o (Dec 2023, May 2024, Nov 2024, Jun 2025)
Medium Priority (Repeated 2 times)
• Distance Measures - List, explain, and provide examples of major classes of distance
measures (e.g., Jaccard, Cosine, Euclidean). (May 2024, Dec 2023, Jun 2025)
Module 6: Real-Time Big Data Models
• 6.1 PageRank
o Q6 b) Explain the structure of web with suitable diagram. - Dec 2023 (Leads
to PageRank)
o Q5 (b) Explain Page rank using Map reduce, also explain spider traps and dead
ends - Nov 2024
o Q6 a) Explain PageRank algorithm. - May 2024
o Q5 (a) Explain Page rank using Map reduce, also explain spider traps and dead
ends - Jun 2025
• 6.2 A Model for Recommendation Systems
o Q5 a) What is Recommender System? Explain Types of recommender
system. - Dec 2023
o Q5 a) Explain Collaborative filtering system. How is it different from content
based system. - May 2024
o Q6 (b) Explain Movie recommendation using Collaborative -based
filtering. - Nov 2024
o Q5 (b) Explain Movie recommendation using Content -based filtering. - Jun
2025
• 6.3 Social Networks as Graphs, Clustering
o Q5 b) What is a Social Network? Give Varieties of Social Networks and the
need for social network graph. - Dec 2023
o Q5 b) What is clique percolation method Write an algorithm on (CPM). - May
2024
o Q6 (c) Clustering of Social-Network Graphs. - Jun 2025
REPEATED QUESTIONS LIST:
High Priority (Repeated 3 times)
• PageRank - Explain the PageRank algorithm, its computation using MapReduce, and
challenges like spider traps and dead ends.
o (May 2024, Nov 2024, Jun 2025)
• Recommendation Systems - Explain types of recommender systems, specifically
Content-based and Collaborative Filtering, and how they differ.
o (Dec 2023, May 2024, Nov 2024, Jun 2025)
Low Priority (Asked Once)
• Social Network Graphs - Clustering of social networks, the Clique Percolation
Method (CPM), and the structure of the web.
o (Dec 2023, May 2024, Jun 2025)
Based on the analysis of these 4 exam papers, here is a breakdown of high-probability
questions and a strategic preparation plan.
Overall Analysis & Strategic Suggestions
1. Heavy Weightage on Core Concepts: Modules 1 (Intro & Hadoop), 2 (MapReduce),
and 3 (NoSQL) form the foundation and are consistently tested in every paper, often
in the compulsory Q1.
2. Pattern of Repetition: The university question paper setters have a clear pattern of
repeating questions, sometimes verbatim, from the last 2-3 years. Focus intensely on
the 2024 and 2025 papers.
3. Question 1 is Crucial: Since Q1 is compulsory, mastering the frequently repeated
topics from this section is non-negotiable for a high score.
4. "Explain with Example" is Key: Many questions demand examples. Prepare concrete,
simple examples for algorithms, architectural patterns, and systems.
High-Probability Questions (Must Prepare)
Here are the questions with the highest chance of appearing, based on repetition and
consistent patterns.
1. Module 1: Introduction to Big Data & Hadoop
• Very High Chance: Compare Traditional Data vs. Big Data (and its 5 V's
characteristics).
• Very High Chance: Explain the Hadoop Ecosystem with a diagram and its core
components.
• High Chance: What are the limitations of Hadoop?
2. Module 2: Hadoop HDFS and MapReduce
• Very High Chance: Explain the MapReduce execution pipeline/job execution steps
with all its components and functions.
• High Chance: Explain how Selection and Projection operations are performed using
MapReduce.
3. Module 3: NoSQL
• Very High Chance: Differentiate between SQL and NoSQL databases.
• High Chance: Explain the four types of NoSQL databases / NoSQL architectural
patterns (Key-Value, Document, Column-Family, Graph) with examples.
4. Module 4: Mining Data Streams
• Very High Chance: Explain the DGIM algorithm for counting ones in a window with
an example.
• High Chance: Explain the Bloom Filter and its analysis.
5. Module 5: Finding Similar Items and Clustering
• Very High Chance: Explain the CURE algorithm and its advantages over traditional
clustering algorithms.
• High Chance: List and explain different Distance Measures (e.g., Jaccard, Cosine,
Euclidean).
6. Module 6: Real-Time Big Data Models
• Very High Chance: Explain the PageRank algorithm. (Be prepared to explain its
computation using MapReduce and concepts like Spider Traps).
• Very High Chance: Explain Recommendation Systems, specifically differentiating
between Content-based and Collaborative Filtering.
Final Preparation Strategy & To-Do List
Based on this analysis, here is a concrete action plan for you:
1. Create a "Super List" of Definitions & Differences:
o Prepare a single document with crisp, clear answers for all Q1-type short
notes, especially:
▪ Big Data vs. Traditional Data
▪ SQL vs. NoSQL
▪ Hadoop Limitations
▪ Distance Measures
▪ Map & Reduce Tasks
▪ CAP Theorem
▪ DBMS vs. DSMS
2. Master the Diagrams:
o Hadoop Ecosystem: Be able to draw this from memory and label all
components.
o MapReduce Execution Steps: Draw a flow chart of the pipeline (Input -> Split
-> Map -> Shuffle & Sort -> Reduce -> Output).
o HDFS Architecture: Know the NameNode and DataNode interaction.
o Data-Stream Management System (DSMS) Architecture: As it appeared in
the 2025 paper.
3. Practice Algorithm Explanations with Examples:
o For each algorithm below, write down a step-by-step explanation and a
simple numerical example.
▪ DGIM Algorithm: Create a sample bit stream (e.g., 101011001110)
and show how DGIM would count the 1s in a window.
▪ PageRank: Use a simple web graph with 3-4 pages and demonstrate
one iteration of the rank calculation.
▪ CURE Algorithm: Explain how it uses representative points to handle
arbitrary cluster shapes, unlike K-Means.
4. Focus on Recent Papers for Long Answers:
o Prioritize your long-answer practice in this order:
1. 2025 June Paper: Every question in this paper is critical.
2. 2024 November Paper: This is your second most important source.
3. 2024 May Paper & 2023 December Paper: Use these to cover any
remaining topics and to see the repetition pattern for yourself.
5. Prepare for "Explain with Example" Questions:
o NoSQL Examples:
▪ Key-Value: Amazon DynamoDB (shopping cart).
▪ Document: MongoDB (user profile store).
▪ Column-Family: Apache Cassandra (time-series data).
▪ Graph: Neo4j (social network recommendations).
o MapReduce Examples: Be ready with examples for Matrix Multiplication and
Relational Algebra operations.
By following this targeted approach, you will be efficiently preparing for over 80-90% of the
potential exam paper, maximizing your score with minimal, focused effort. Good luck!