0% found this document useful (0 votes)
3 views13 pages

Big Data Analysis: Hadoop & NoSQL Insights

The document outlines a comprehensive analysis of exam questions related to Big Data, Hadoop, MapReduce, NoSQL, and data stream mining across multiple modules. It highlights high-priority questions that are frequently repeated in exams and suggests a strategic preparation plan focusing on core concepts, particularly those in Modules 1, 2, and 3. Additionally, it emphasizes the importance of mastering diagrams and definitions for effective exam performance.

Uploaded by

amour3255
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views13 pages

Big Data Analysis: Hadoop & NoSQL Insights

The document outlines a comprehensive analysis of exam questions related to Big Data, Hadoop, MapReduce, NoSQL, and data stream mining across multiple modules. It highlights high-priority questions that are frequently repeated in exams and suggests a strategic preparation plan focusing on core concepts, particularly those in Modules 1, 2, and 3. Additionally, it emphasizes the importance of mastering diagrams and definitions for effective exam performance.

Uploaded by

amour3255
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BIG DATA ANALYSIS PYQ AIDS

Module 1: Introduction to Big Data & Hadoop


• 1.1 Introduction to Big Data, 1.2 Big Data characteristics

o Q1 a) What is Hadoop and Why it Matters. - Dec 2023

o Q1 b) Compare traditional database and big data. - Dec 2023

o Q1 (a) Explain 5 V’s of big data. - Nov 2024

o Q1 a) Big Data and its characteristics - May 2024

o Q1 (a) Compare traditional data and big data. - Jun 2025

o Q2 (b) Explain big data enabling technologies. - Nov 2024 (Covers Hadoop
Ecosystem from 1.6)

• 1.5 Concept of Hadoop, 1.6 Core Hadoop Components; Hadoop Ecosystem

o Q2 a) Draw Hadoop Ecosystem and briefly explain its components. - Dec 2023

o Q1 (c) Write the limitations of Hadoop. - Nov 2024 (Covers 2.4 Hadoop
Limitations)

o Q4 (a) Explain Hadoop Architectural Model with both components in


detail - Nov 2024

o Q2 a) Explain HDFS architecture. - May 2024 (Covers 2.1)

o Q1 (b) What are the advantages and limitations of Hadoop - Jun 2025

o Q2 a) Draw Hadoop Ecosystem and briefly explain its components - Jun 2025

REPEATED QUESTIONS LIST

High Priority (Repeated 3-4 times)

• Hadoop Ecosystem & Components - Explain/Draw and explain the Hadoop


Ecosystem and its components. (Jun 2025, Dec 2023, Nov 2024)
• Big Data Characteristics & Vs. Traditional Data - Explain 5 V's, compare Big Data with
traditional data/database. (May 2024, Jun 2025, Dec 2023, Nov 2024)

Medium Priority (Repeated 2 times)

• Hadoop Limitations - List/Explain the advantages and limitations of Hadoop. (Jun


2025, Nov 2024)

Low Priority (Asked Once)

• Introduction to Hadoop - What is Hadoop and why it matters? (Dec 2023)

• Big Data Enabling Technologies - Explain big data enabling technologies. (Nov 2024)

• HDFS Architecture - Explain HDFS architecture. (May 2024)


Module 2: Hadoop HDFS and Map Reduce

• 2.1 Distributed File Systems (HDFS)

o Q2 a) Explain HDFS architecture. - May 2024

• 2.2 MapReduce: The Map and Reduce Tasks, Execution, Failures

o Q1 d) Compare DBMS VS DSMS. - Dec 2023 (Related to stream processing


context)

o Q1 (d) Explain how failures are handled in Map Reduce job - Nov 2024

o Q1 c) The Map and Reduce Tasks - May 2024

o Q4 a) List the main components of Mapreduce execution pipeline. - Dec 2023

o Q4 (b) Write the functions of the components and execution steps in Map
Reduce - Nov 2024

o Q4 a) List the main components of Map reduce execution pipeline. - May


2024

o Q2 (b) Write the functions of the components and execution steps in Map
Reduce - Jun 2025

• 2.3 Algorithms Using MapReduce: Relational-Algebra Operations, Matrix


Multiplication

o Q2 (a) Illustrate relational algebra operations with example. - Nov 2024

o Q3 a) Write a Map reduce pseudo code to multiply two matrices. - May 2024

o Q3 (a) Explain Selection and Projection algebraic operation using


MapReduce. - Jun 2025

• 2.4 Hadoop Limitations

o Q1 (c) Write the limitations of Hadoop. - Nov 2024

o Q1 (b) What are the advantages and limitations of Hadoop - Jun 2025
REPEATED QUESTIONS LIST:

High Priority (Repeated 3-4 times)

• MapReduce Execution & Steps - List components, explain functions, and detail the
execution steps of MapReduce. (May 2024, Dec 2023, Jun 2025, Nov 2024)

Medium Priority (Repeated 2 times)

• MapReduce Algorithms (Relational Algebra) - Explain/Illustrate Selection, Projection,


and other relational algebra operations using MapReduce.

o (Nov 2024, Jun 2025)

• Handling Failures in MapReduce - Explain how failures are handled in a MapReduce


job.(Nov 2024) [Implied in execution steps of other papers]

Low Priority (Asked Once)

• Matrix Multiplication with MapReduce - Write a MapReduce pseudo code to


multiply two matrices.(May 2024)

• DBMS vs. DSMS - Compare DBMS and DSMS (Data Stream Management System).
(Dec 2023)
Module 3: NoSQL

• 3.1 Introduction to NoSQL, SQL vs NoSQL

o Q1 c) Explain CAP theorem. State how it is different from ACID


properties. - Dec 2023

o Q1 (b) Differentiate between SQL vs NoSQL - Nov 2024

o Q1 (c) Differentiate between SQL vs NoSQL - Jun 2025

• 3.2 NoSQL Data Architecture Patterns

o Q2 b) Explain the four types of NoSQL database. - Dec 2023

o Q3 (b) Compare different types of NoSQL architectural pattern - Nov 2024

o Q2 b) Explain Column family store and Graph Store NoSQL architectural


pattern - May 2024

o Q3 (b) Explain Key-value store and Document Store NoSQL architectural


pattern with example. - Jun 2025

o Q6 (d) Four ways that NoSQL systems handle big data problems. - Jun
2025 (Covers 3.3)

• 3.3 NoSQL solution for big data

o Q6 (d) Four ways that NoSQL systems handle big data problems. - Jun 2025

REPEATED QUESTIONS LIST:


High Priority (Repeated 4 times)

• SQL vs. NoSQL - Differentiate between SQL and NoSQL databases.

o (Jun 2025, Nov 2024, Dec 2023 via CAP Theorem, May 2024 via architectural
patterns)

Medium Priority (Repeated 2-3 times)


• NoSQL Architectural Patterns / Types - Explain and compare the four
types/architectural patterns of NoSQL databases (Key-Value, Document, Column-
Family, Graph).

o (Dec 2023, Jun 2025, May 2024)

Low Priority (Asked Once)

• CAP Theorem - Explain the CAP theorem and how it differs from ACID properties.

o (Dec 2023)

• NoSQL for Big Data - Explain the four ways that NoSQL systems handle big data
problems.

o (Jun 2025)

Module 4: Mining Data Streams

• 4.1 The Stream Data Model: Issues, DSMS

o Q1 d) Compare DBMS VS DSMS. - Dec 2023

o Q5 (a) Write issues in data stream queries. Explain the issues in data
streaming - Nov 2024

o Q3 b) Explain Issues in Data stream query processing - May 2024

o Q4 (a) Draw a neat sketch, explain the architecture of the data-stream


management system - Jun 2025

• 4.3 Filtering Streams: Bloom Filter

o Q1 d) Bloom filter for stream data mining - May 2024

o Q6 (a) Bloom Filter with analysis - Jun 2025


• 4.5/4.6 Counting Frequent Items & Counting Ones (DGIM Algorithm)

o Q3 a) Explain PCY algorithm and its types with neat labeled diagram - Nov
2024

o Q3 b) Explain DGIM algorithm. - Dec 2023

o Q4 b) Explain DGIM algorithm. - May 2024

o Q4 (b) Explain DGIM algorithm for counting ones in a stream with


example - Jun 2025

REPEATED QUESTIONS LIST:


High Priority (Repeated 3 times)

• DGIM Algorithm - Explain the DGIM algorithm for counting ones in a stream.

o (May 2024, Dec 2023, Jun 2025)

Medium Priority (Repeated 2 times)

• Stream Data Model & Issues - Explain the issues in data stream query processing and
the architecture of a Data-Stream Management System (DSMS).

o (May 2024, Jun 2025, Nov 2024)

Low Priority (Asked Once)

• Bloom Filter - Explain the Bloom filter with analysis for stream data mining.

o (May 2024, Jun 2025)

• PCY Algorithm - Explain the PCY algorithm and its types.

o (Nov 2024)
Module 5: Finding Similar Items and Clustering
• 5.1 Distance Measures

o Q6 a) Explain with example two major classes of distance measures. - Dec


2023

o Q1 b) Distance measures for Big Data - May 2024

o Q1 (d) List and explain Distance measures for Big Data - Jun 2025

• 5.2 CURE Algorithm, Stream-Clustering

o Q4 b) Explain cure algorithm. - Dec 2023

o Q6 a) Explain CURE algorithm with its advantages over traditional clustering


algorithm - Nov 2024

o Q6 b) Explain CURE algorithm. - May 2024

o Q6 (b) Cure Algorithm - Jun 2025

REPEATED QUESTIONS LIST:


High Priority (Repeated 3 times)

• CURE Algorithm - Explain the CURE algorithm, its advantages, and how it works.

o (Dec 2023, May 2024, Nov 2024, Jun 2025)

Medium Priority (Repeated 2 times)

• Distance Measures - List, explain, and provide examples of major classes of distance
measures (e.g., Jaccard, Cosine, Euclidean). (May 2024, Dec 2023, Jun 2025)
Module 6: Real-Time Big Data Models
• 6.1 PageRank

o Q6 b) Explain the structure of web with suitable diagram. - Dec 2023 (Leads
to PageRank)

o Q5 (b) Explain Page rank using Map reduce, also explain spider traps and dead
ends - Nov 2024

o Q6 a) Explain PageRank algorithm. - May 2024

o Q5 (a) Explain Page rank using Map reduce, also explain spider traps and dead
ends - Jun 2025

• 6.2 A Model for Recommendation Systems

o Q5 a) What is Recommender System? Explain Types of recommender


system. - Dec 2023

o Q5 a) Explain Collaborative filtering system. How is it different from content


based system. - May 2024

o Q6 (b) Explain Movie recommendation using Collaborative -based


filtering. - Nov 2024

o Q5 (b) Explain Movie recommendation using Content -based filtering. - Jun


2025

• 6.3 Social Networks as Graphs, Clustering

o Q5 b) What is a Social Network? Give Varieties of Social Networks and the


need for social network graph. - Dec 2023

o Q5 b) What is clique percolation method Write an algorithm on (CPM). - May


2024

o Q6 (c) Clustering of Social-Network Graphs. - Jun 2025


REPEATED QUESTIONS LIST:
High Priority (Repeated 3 times)

• PageRank - Explain the PageRank algorithm, its computation using MapReduce, and
challenges like spider traps and dead ends.

o (May 2024, Nov 2024, Jun 2025)

• Recommendation Systems - Explain types of recommender systems, specifically


Content-based and Collaborative Filtering, and how they differ.

o (Dec 2023, May 2024, Nov 2024, Jun 2025)

Low Priority (Asked Once)

• Social Network Graphs - Clustering of social networks, the Clique Percolation


Method (CPM), and the structure of the web.

o (Dec 2023, May 2024, Jun 2025)


Based on the analysis of these 4 exam papers, here is a breakdown of high-probability
questions and a strategic preparation plan.

Overall Analysis & Strategic Suggestions

1. Heavy Weightage on Core Concepts: Modules 1 (Intro & Hadoop), 2 (MapReduce),


and 3 (NoSQL) form the foundation and are consistently tested in every paper, often
in the compulsory Q1.

2. Pattern of Repetition: The university question paper setters have a clear pattern of
repeating questions, sometimes verbatim, from the last 2-3 years. Focus intensely on
the 2024 and 2025 papers.

3. Question 1 is Crucial: Since Q1 is compulsory, mastering the frequently repeated


topics from this section is non-negotiable for a high score.

4. "Explain with Example" is Key: Many questions demand examples. Prepare concrete,
simple examples for algorithms, architectural patterns, and systems.

High-Probability Questions (Must Prepare)

Here are the questions with the highest chance of appearing, based on repetition and
consistent patterns.

1. Module 1: Introduction to Big Data & Hadoop

• Very High Chance: Compare Traditional Data vs. Big Data (and its 5 V's
characteristics).

• Very High Chance: Explain the Hadoop Ecosystem with a diagram and its core
components.

• High Chance: What are the limitations of Hadoop?

2. Module 2: Hadoop HDFS and MapReduce

• Very High Chance: Explain the MapReduce execution pipeline/job execution steps
with all its components and functions.

• High Chance: Explain how Selection and Projection operations are performed using
MapReduce.

3. Module 3: NoSQL

• Very High Chance: Differentiate between SQL and NoSQL databases.

• High Chance: Explain the four types of NoSQL databases / NoSQL architectural
patterns (Key-Value, Document, Column-Family, Graph) with examples.
4. Module 4: Mining Data Streams

• Very High Chance: Explain the DGIM algorithm for counting ones in a window with
an example.

• High Chance: Explain the Bloom Filter and its analysis.

5. Module 5: Finding Similar Items and Clustering

• Very High Chance: Explain the CURE algorithm and its advantages over traditional
clustering algorithms.

• High Chance: List and explain different Distance Measures (e.g., Jaccard, Cosine,
Euclidean).

6. Module 6: Real-Time Big Data Models

• Very High Chance: Explain the PageRank algorithm. (Be prepared to explain its
computation using MapReduce and concepts like Spider Traps).

• Very High Chance: Explain Recommendation Systems, specifically differentiating


between Content-based and Collaborative Filtering.

Final Preparation Strategy & To-Do List

Based on this analysis, here is a concrete action plan for you:

1. Create a "Super List" of Definitions & Differences:

o Prepare a single document with crisp, clear answers for all Q1-type short
notes, especially:

▪ Big Data vs. Traditional Data

▪ SQL vs. NoSQL

▪ Hadoop Limitations

▪ Distance Measures

▪ Map & Reduce Tasks

▪ CAP Theorem

▪ DBMS vs. DSMS

2. Master the Diagrams:

o Hadoop Ecosystem: Be able to draw this from memory and label all
components.
o MapReduce Execution Steps: Draw a flow chart of the pipeline (Input -> Split
-> Map -> Shuffle & Sort -> Reduce -> Output).

o HDFS Architecture: Know the NameNode and DataNode interaction.

o Data-Stream Management System (DSMS) Architecture: As it appeared in


the 2025 paper.

3. Practice Algorithm Explanations with Examples:

o For each algorithm below, write down a step-by-step explanation and a


simple numerical example.

▪ DGIM Algorithm: Create a sample bit stream (e.g., 101011001110)


and show how DGIM would count the 1s in a window.

▪ PageRank: Use a simple web graph with 3-4 pages and demonstrate
one iteration of the rank calculation.

▪ CURE Algorithm: Explain how it uses representative points to handle


arbitrary cluster shapes, unlike K-Means.

4. Focus on Recent Papers for Long Answers:

o Prioritize your long-answer practice in this order:

1. 2025 June Paper: Every question in this paper is critical.

2. 2024 November Paper: This is your second most important source.

3. 2024 May Paper & 2023 December Paper: Use these to cover any
remaining topics and to see the repetition pattern for yourself.

5. Prepare for "Explain with Example" Questions:

o NoSQL Examples:

▪ Key-Value: Amazon DynamoDB (shopping cart).

▪ Document: MongoDB (user profile store).

▪ Column-Family: Apache Cassandra (time-series data).

▪ Graph: Neo4j (social network recommendations).

o MapReduce Examples: Be ready with examples for Matrix Multiplication and


Relational Algebra operations.

By following this targeted approach, you will be efficiently preparing for over 80-90% of the
potential exam paper, maximizing your score with minimal, focused effort. Good luck!

You might also like