BDA Module 1
BDA Module 1
Eg: 05111988
DOB : 05/11/1988- Virat Kohli
3
Traditional data
It is the kind of information that is easy to organize and store in simple
databases, like spreadsheets or small computer systems. This could be things
like customer names, phone numbers, or sales records.
Advantages of Traditional Data
Stored, managed easier with regular database systems.
It is relatively inexpensive in storing and processing of data or datasets that are
relatively small in size.
Disadvantages of Traditional Data
This tends to be only applicable in structured formats and can therefore limit
flexibility.
This does not work well for unstructured data types such as text, images or
videos.
4
What is Big Data?
No single standard definition…
Big data refers to massive, complex data sets (either structured,
semi-structured or unstructured) that are rapidly generated and transmitted from a
wide variety of sources
“Big Data ” is data whose scale, diversity, and complexity require new
architecture, techniques, algorithms, and analytics to manage it and extract value
and hidden knowledge from it…
Massive amount of data which can not be stored, processed and analyzed using
traditional tools is called as Big Data.
5
What is Big Data?
6
THE DIFFERENCE BETWEEN TRADITIONAL DATA AND
BIG DATA
The Difference Between Traditional Data and Big Data
10
How is big data used for improving the companies?
1. Predict exactly what customers want before they ask for it
11
CHARACTERISTICS OF BIG DATA:
Volume
● The name Big Data itself is related to an enormous size.
● Big Data is a vast 'volumes' of data generated from
many sources daily, such as business processes, machines, social media
platforms, networks, human interactions, Twitter data feeds, clickstreams
on a web page or a mobile app, or sensor-enabled equipment and
many more.
13
Value
Value refers to the usefulness or benefits you can extract from Big Data.
It answers the question:
“What meaningful insights or business advantages can we get from the data?”
Eg: How Value is Extracted from Big Data in E-commerce--Recommending products based on user behavior
(Amazon, Flipkart)
14
Variety
● Data had been collected from databases and sheets in the past,
● But these days the data will comes in any forms, that are PDFs,
Emails, audios, SM posts, photos, videos, etc.
15
Variety
16
Velocity
Data is begin generated fast and need to be processed fast.
• Velocity plays an important role compared to others. Velocity creates the speed
by which the data is created in real-time. It contains the linking of incoming
data sets speeds, rate of change, and activity bursts. The primary
aspect of Big Data is to provide demanding data rapidly.
• Big data velocity deals with the speed at the data flows from sources like
application logs, business processes, networks,
17
Veracity
Veracity is a big data characteristic related to the trustworthiness, accuracy,
and quality of the data. It addresses the uncertainty, inconsistency, and
reliability of data being collected.
18
Big Data: 3V’s
19
Some Make it 4V’s
20
5 V's of Big Data
•Volume
•Value
•Variety
•Velocity
•Veracity
21
Comparison: Traditional (Small) Data vs Big Data using 5 V’s
Volume Data stored in MB, GB, TB Data stored in PB, EB (huge volume)
Data increases exponentially, may stream in
Velocity Data increases gradually, updated periodically
real-time
Mostly unstructured (images, videos, text, logs,
Variety Mostly structured (tables, rows, columns)
etc.)
Stored locally, used for specific, smaller Globally present, used for broad insights, analytics
Value
applications & AI
Tools (Extra) SQL Server, Oracle (Single Node) Hadoop, Spark (Multinode Cluster)
22
Types of Bigdata
Following are the types of Big Data:
1. Structured
2. Unstructured
[Link]-structured
23
Types of Big data
24
Structured data
Structured data has certain predefined organizational properties and is present in
structured or tabular schema, making it easier to analyze and sort.
The business data of an e-commerce website can be considered to be structured data.
Un-Structured Data
Unstructured data accounts for the majority of big data and comprises information such as
dates, numbers, and facts.
Photos we upload on Facebook or Instagram and videos that we watch on YouTube
or any other platform contribute to the growing pile of unstructured data
25
Semi Structured Data
Semi-structured data is a hybrid of structured and unstructured data. This means that it
inherits a few characteristics of structured data but nonetheless contains information that
fails to have a definite structure and does not conform with relational databases or formal
structures of data models.
For instance, JSON and XML are typical examples of semi-structured data.
26
ORIGINS AND GROWTH OF BIG DATA:
Who’s Generating Big Data( Origins)
The progress and innovation is no longer hindered by the ability to collect data
But, by the ability to manage, analyze, summarize, visualize, and discover knowledge from the collected
data in a timely manner and in a scalable fashion
28
What’s driving Big Data( Growth)
What factors are causing Big Data to grow and become important?
Main Drivers:
Explosion of digital data (from apps, social media, IoT, etc.)
Need for real-time insights (instant decisions)
Demand for advanced analytics (predictive models, AI)
Cheaper storage & faster processing (cloud, big data tools)
Growth of AI & ML (which need lots of data to learn)
Big Data is being generated by everything connected to the internet — people, machines, apps
— and it’s growing fast because of the digital lifestyle we live in. Businesses need to keep up
with this data to stay competitive
29
HARNESSING BIG DATA WITH MODERN
TECHNOLOGIES AND PROCESSING
SYSTEMS??
:
Harnessing Big Data
Using advanced technologies and architectures to collect, store, process, and analyze vast
amounts of diverse and rapidly changing data to extract meaningful insights and make
data-driven decisions.
31
Technologies Available for Big Data
2. Data Ingestion Tools: Bring data from various sources into Big Data systems
Apache Kafka – Handles real-time streaming data
Apache Flume – Gathers log data (web, mobile, IoT)
Apache NiFi – Manages data flows with automation
3. Data Storage Technologies:Store huge volumes of structured and unstructured data
HDFS (Hadoop Distributed File System) – Distributed storage
NoSQL Databases – MongoDB, Cassandra for flexible storage
Cloud Storage – Amazon S3, Google Cloud Storage
4. Data Processing Engines:Process and compute data in batch or real time
Hadoop MapReduce – Batch processing (slow but reliable)
Apache Spark – Fast, in-memory processing
Apache Storm / Flink – Real-time event stream processing
32
Technologies Available for Big Data
5. Query & Analysis Tools
Analyze large datasets using queries or scripts
Hive, Pig – Run queries on Hadoop
Presto / Trino – Fast querying across big datasets
6. ETL & Data Pipeline Tools
Clean, transform, and move data between systems
Talend – ETL (Extract, Transform, Load) tool
Apache Airflow – Automates and schedules workflows
[Link] Visualization & BI Tools
Present data insights in charts, dashboards, and reports
Tableau, Power BI – Visual dashboards
Grafana – Real-time monitoring dashboards
33
Technologies Available for Big Data
34
Advantage of Big Data Analytics
Customer Acquisition and Retention
Big data helps companies understand what customers like by analyzing their online behavior
and shopping history, so they can attract and keep them with personalized offers.
Focused and Targeted Promotions
It lets businesses send the right ads to the right people at the right time, increasing sales and
saving money on marketing.
Potential Risks Identification
By studying past data, companies can spot possible problems early, like fraud or delays, and
take action before they grow.
Innovation
Analyzing data helps businesses find new ideas for products or services based on what
customers want and market trends.
Complex Supplier Networks
Big data gives clear info on suppliers, helping companies manage inventory, cut costs, and
avoid delays in delivery.
35
Advantage of Big Data Analytics
Cost Optimization
It shows where money is being wasted, so businesses can fix inefficiencies and save on
operations.
Improve Efficiency
Data helps identify slow or repetitive tasks in a business, so they can be automated or
improved for smoother work.
Better Decision Making
With accurate and timely data, businesses can make smarter, confident decisions instead of
guessing.
Enhanced Customer Experience
By knowing customer needs better, businesses can offer personalized services, making
customers happier and more loyal.
Competitive Advantage
Using big data smartly helps businesses stay ahead of rivals by adapting faster and making
better choices.
36
Case study :Big data at Wal-Mart
Walmart, one of the largest retailers in the world, continues to leverage Big Data to
understand consumer behavior, optimize inventory, and drive sales.
Walmart's advanced analytics platform processes millions of transactions daily across its
global stores and online platforms. By analyzing this vast data, they predict customer
purchasing patterns, optimize supply chain logistics, and provide personalized offers.
Recent Innovation:
Walmart has been integrating AI and Big Data to optimize product recommendations, tailor
advertisements, and enhance customer engagement. For instance, during the 2023 holiday
season, Walmart used Big Data to forecast the types of products that would be in high
demand (such as video games, electronics, and holiday-specific items) based on historical
data and current trends.
Impact:
The use of data-driven insights improved inventory management, leading to reduced
stockouts, more personalized shopping experiences, and increased sales during peak periods
like Black Friday and Christmas.
37
Case study :Big data at Netflix
Netflix is a prime example of a company using Big Data for personalization and improving user
experience.
What Happened:
Netflix collects data on what content users are watching, how they interact with the
platform, and what time of day they watch. It also considers factors such as rating patterns,
genres, and search history.
Recent Innovation:
In 2023, Netflix introduced an AI-driven personalization engine that provides
hyper-targeted content recommendations. The system processes billions of data points
from over 200 million subscribers and recommends content that a user is most likely to
watch based on similar patterns of behavior from other viewers.
Impact:
The recommendations system drives over 80% of all content watched on Netflix, improving
user retention and helping the company save on advertising costs. This data-driven strategy
helps retain customers and ensures higher engagement rates.
38
Challenges in Handling Big Data
39
Hadoop
Hadoop is an open-source software framework that is used for
storing and processing large amounts of data in a distributed
computing environment.
● HDFS(Hadoop Distributed
File System)
● Map-Reduce
● Hadoop-Common or
Common Utilities Core Hadoop Component
41
Core Hadoop Component
1. MapReduce
MapReduce is a data processing model in Hadoop built on top of the YARN framework. Its
main feature is distributed and parallel processing across a Hadoop cluster, which significantly
boosts performance when handling Big Data. Since serial processing is inefficient at large scale,
MapReduce enables fast and efficient data processing by splitting the workload into two
phases: Map and Reduce.
42
Core Hadoop Component
MapReduce Workflow: The workflow begins when input data is passed to the Map() function.
This data, often stored in blocks, is broken down into tuples (key-value pairs). These pairs are
then passed to the Reduce() function, which aggregates them based on the keys and performs
required operations such as sorting, counting or summing to generate the final output. The
result is then written to HDFS.
2. HDFS
HDFS(Hadoop Distributed File System) is utilized for storage permission. It is mainly
designed for working on commodity Hardware devices(inexpensive devices), working on a
distributed file system design. HDFS is designed in such a way that it believes more in storing
the data in a large chunk of blocks rather than storing small data blocks.
HDFS in Hadoop provides Fault-tolerance and High availability to the storage layer and the
other devices present in that Hadoop cluster.
43
Core Hadoop Component
HDFS Architecture Components
1. NameNode (Master Node):
Manages the HDFS cluster and stores metadata (file names, sizes, permissions, block
locations).
Does not store actual data, only metadata (via edit logs and fsimage).
Controls file operations like create, delete, and replicate, directing DataNodes accordingly.
Helps clients locate the nearest DataNode for fast access.
2. DataNode (Slave Node):
Stores the actual data blocks on local disks.
Handles read/write requests from clients.
Sends heartbeats and block reports to the NameNode regularly.
Data is replicated (default 3 copies) for reliability.
44
Core Hadoop Component
4. Hadoop Common (Common Utilities)
Hadoop Common, also known as Common Utilities, includes the core Java libraries and
scripts required by all the components in a Hadoop ecosystem such as HDFS, YARN and
MapReduce. These libraries provide essential functionalities like:
File system and I/O operations
Configuration and logging
Security and authentication
Network communication
Hadoop Common ensures that the entire cluster works cohesively and that hardware failures,
which are common in Hadoop's commodity hardware setup, are handled automatically by the
software
45
Core Hadoop Component
3. YARN (Yet Another Resource Negotiator)
YARN is the resource management layer in the Hadoop ecosystem. It allows multiple data
processing engines like MapReduce, Spark, and others to run and share cluster resources
efficiently. It handles two core responsibilities:
Job Scheduling: Breaks a large task into smaller jobs and assigns them to various nodes in
the Hadoop cluster. It also tracks job priority, dependencies, and execution time.
Resource Management: Allocates and manages the cluster resources (CPU, memory, etc.)
required for running these jobs.
46
Advantages and Disadvantages of Hadoop
Advantages:
Scalability: Easily scale to thousands of machines.
Cost-effective: Uses low-cost hardware to process big data.
Fault Tolerance: Automatic recovery from node failures.
High Availability: Data replication ensures no loss even if nodes fail.
Flexibility: Can handle structured, semi-structured and unstructured data.
Open-source and Community-driven: Constant updates and wide support.
Disadvantages:
Not ideal for real-time processing (better suited for batch processing).
Complexity in programming with MapReduce.
High latency for certain types of queries.
Requires skilled professionals to manage and develop.
47
Applications Of Hadoop
Hadoop is used across a variety of industries:
Banking: Fraud detection, risk modeling.
Retail: Customer behavior analysis, inventory management.
Healthcare: Disease prediction, patient record analysis.
Telecom: Network performance monitoring.
Social Media: Trend analysis, user recommendation engines.
48
How Does Hadoop Work?
Here’s a overview of how Hadoop operates:
A large data file is uploaded to HDFS.
HDFS splits the file into fixed-size blocks (e.g., 128MB).
These blocks are distributed and stored across multiple machines (DataNodes).
Each block is replicated for fault tolerance.
The NameNode keeps track of where each block is stored.
49
High Level Architecture Of Hadoop
Hadoop works internally, using a Master-Slave Architecture to store
and process big data. It involves:
Master Node
Slave Nodes
Think of the Master as the Manager – it tells other workers (slaves) what to do and where
everything is.
51
Architecture Of Hadoop
Slave Nodes
Each slave node contains four components:
DataNode : Stores actual data blocks (pieces of big files).
Reports to the NameNode about what data it has.
Node Manager :Works under the Resource Manager.
Monitors and manages jobs (Map or Reduce tasks) on this machine.
Map Task:This is the first phase of data processing.
It processes data blocks and produces key-value pairs.
Reduce Task:This is the second phase of processing.
It aggregates or summarizes the data output from the Map phase.
Think of each slave as a worker:
They store data (DataNode)
And process tasks (Map/Reduce) when assigned by the master.
52
limitations of Hadoop
Handling Small Files:
Hadoop struggles with many small files because it's designed for large files. This can overwhelm the system and
slow things down.
Slow Processing Speed:
Hadoop's MapReduce is slow because it relies on reading and writing data to disk, which creates latency.
Batch Processing Only:
Hadoop only supports batch processing, which means it processes data in large chunks, not in real-time.
No Real-Time Data Processing:
Hadoop is not suited for real-time data, making it unsuitable for applications that need fast, continuous data
processing.
No Support for Iterative Processing:
Hadoop is inefficient for tasks that need repeated data processing, such as machine learning, because it doesn't
handle iterative steps well.
High Latency:
The time it takes for Hadoop to process and return results (latency) can be high due to the multiple steps involved
in the MapReduce process.
Difficult to Use:
Writing MapReduce jobs requires complex programming, making Hadoop harder to use, especially for beginners.
53
limitations of Hadoop
Security Issues:
Hadoop lacks strong out-of-the-box security features and requires complex configuration for
secure operations.
No Abstraction:
Hadoop doesn't offer high-level abstractions, meaning developers have to deal with
low-level details like handling data manually.
Vulnerability to Attacks:
Since Hadoop is primarily written in Java, it inherits Java's security risks, making it a target
for cyber-attacks.
Unpredictable Job Completion Time:
Hadoop doesn't provide a guarantee on when jobs will finish, making it unreliable for
time-sensitive tasks.
54
Hadoop Ecosystem
The Hadoop Ecosystem is a collection of tools that work with
Hadoop to handle different types of big data tasks
Ingesting data (getting data into Hadoop) such as:
These tools help Hadoop handle real-world big data problems more
efficiently.
55
Hadoop Ecosystem
56
Hadoop Ecosystem
57
Sqoop
A lot of applications still store data in relational databases, thus making them a very important
source of data. Therefore, Sqoop plays an important part in bringing data from Relational
Databases into HDFS.
The commands written in Sqoop internally converts into MapReduce tasks that are executed
over HDFS. It works with almost all relational databases like MySQL, Postgres, SQLite, etc. It
can also be used to export data from HDFS to RDBMS.
Apache Sqoop, a command-line interface tool, moves data between relational databases
and Hadoop.
● It is used to export data from the Hadoop file system to relational
databases andto import data from relational databases such as MySQL and
Oracle into the Hadoop file system.
58
Flume
● Flume is an open-source, reliable, and available service used to efficiently collect, aggregate,
and move large amounts of data from multiple data sources into HDFS. It can collect data in
real-time as well as in batch mode. It has a flexible architecture and is fault-tolerant with
multiple recovery mechanisms.
● Apache Flume is a distributed system for collecting, aggregating, and transferring data
from multiple sources to a centralized data store.
● Apache Flume can be used to transport large amounts of social-media
generated data, logs, network traffic data, email messages, and many more to a centralized data
store.
● In simple words, Apache Flume is a tool in the Hadoop ecosystem for
transferring data from one location to anotherefficiently andreliably. The
main goal of Apache Flume is to transfer data from applications to Hadoop HDFS.
● It is highly robust, reliable, and fault-tolerant.
Kafka
59
Pig
Pig was developed for analyzing large datasets and overcomes the difficulty to write
map and reduce functions. It consists of two components: Pig Latin and Pig Engine.
Pig Latin is the Scripting Language that is similar to SQL. Pig Engine is the execution
engine on which Pig Latin runs. Internally, the code written in Pig is converted to
MapReduce functions and makes it very easy for programmers who aren’t proficient
in Java.
Apache Pig is a high-level language platform for analyzing and querying huge
dataset that are stored in HDFS. Pig as a component of Hadoop Ecosystem uses
PigLatin language. It is very similar to SQL. It loads the data, applies the required
filters and dumps the data in the required format. For Programs execution, pig
requires Java runtime environment.
60
Hive
It allows for easy reading, writing, and managing files on HDFS. It has its own querying
language for the purpose known as Hive Querying Language (HQL) which is very similar to
SQL.
The Hadoop ecosystem component, Apache Hive, is an open source data warehouse system
for querying and analyzing large datasets stored in Hadoop files. Hive do three main
functions: data summarization, query, and analysis.
Hive use language called HiveQL (HQL), which is similar to SQL. HiveQL automatically
translates SQL-like queries into MapReduce jobs which will execute on Hadoop.
61
HBase
Apache HBase is a Hadoop ecosystem component which is a distributed database.
HBase is a Column-based NoSQL database. It runs on top of HDFS and can handle any type of
data. It allows for real-time processing and random read/write operations to be performed in
the data
HBase is scalable, distributed, and NoSQL database that is built on top of HDFS. HBase,
individual columns, and indexed by a unique row key. This architecture allows for rapid
retrieval of individual rows and columns and efficient scans over individual columns within a
table.
62
Mahout
● It is a powerful, scalable machine-learning library that runs on top of Hadoop
MapReduce.
● Mahout is supported by its 3 pillars:
○ Recommender engines: Recommenders can be classified as being user based or item
based and can be used to attract users and suggest products by mining user behaviour.
○ Clustering: Clustering groups objects of a similar nature in one place.
○ Classification: Classification techniques decide whether a thing deserves to be a part
of some type or not.
63
Hadoop Ecosystem
Oozie
:Oozie is a workflow scheduler system that allows users to link jobs written on various
platforms like MapReduce, Hive, Pig, etc. Using Oozie you can schedule a job in advance and
can create a pipeline of individual jobs to be executed sequentially or in parallel to achieve a
bigger task
Zookeeper
In a Hadoop cluster, coordinating and synchronizing nodes can be a challenging task.
Therefore, Zookeeper is the perfect tool for the problem.
It is an open-source, distributed, and centralized service for maintaining configuration
information, naming, providing distributed synchronization, and providing group services
across the cluster.
64
Spark
Spark is an alternative framework to Hadoop built on Scala but supports varied applications
written in Java, Python, etc. Compared to MapReduce it provides in-memory processing
which accounts for faster processing. In addition to batch processing offered by Hadoop, it
can also handle real-time processing.
65
Thank
You