0% found this document useful (0 votes)
18 views66 pages

BDA Module 1

Data refers to a collection of facts and information that can be structured or unstructured, while big data encompasses massive, complex datasets that require advanced tools for management and analysis. The key differences between traditional data and big data include volume, variety, velocity, and complexity, necessitating specialized technologies for effective handling. Big data analytics provides insights that can enhance customer experiences, optimize operations, and drive innovation across various industries.

Uploaded by

abhisheknaikedu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views66 pages

BDA Module 1

Data refers to a collection of facts and information that can be structured or unstructured, while big data encompasses massive, complex datasets that require advanced tools for management and analysis. The key differences between traditional data and big data include volume, variety, velocity, and complexity, necessitating specialized technologies for effective handling. Big data analytics provides insights that can enhance customer experiences, optimize operations, and drive innovation across various industries.

Uploaded by

abhisheknaikedu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

What is DATA

In general, data is a collection of facts, information, and


statistics and this can be in various forms such as numbers,
text, sound, images, or any other format.
Information:
Information is data that has been processed, organized, or structured
in a way that makes it meaningful, valuable, and useful.

Eg: 05111988
DOB : 05/11/1988- Virat Kohli

3
Traditional data
It is the kind of information that is easy to organize and store in simple
databases, like spreadsheets or small computer systems. This could be things
like customer names, phone numbers, or sales records.
Advantages of Traditional Data
Stored, managed easier with regular database systems.
It is relatively inexpensive in storing and processing of data or datasets that are
relatively small in size.
Disadvantages of Traditional Data
This tends to be only applicable in structured formats and can therefore limit
flexibility.
This does not work well for unstructured data types such as text, images or
videos.
4
What is Big Data?
No single standard definition…
Big data refers to massive, complex data sets (either structured,
semi-structured or unstructured) that are rapidly generated and transmitted from a
wide variety of sources

“Big Data ” is data whose scale, diversity, and complexity require new
architecture, techniques, algorithms, and analytics to manage it and extract value
and hidden knowledge from it…

Massive amount of data which can not be stored, processed and analyzed using
traditional tools is called as Big Data.

5
What is Big Data?

6
THE DIFFERENCE BETWEEN TRADITIONAL DATA AND
BIG DATA
The Difference Between Traditional Data and Big Data

Aspect Traditional Data Big Data


Structured,
Mostly structured (tables, semi-structured, and
Data Type
rows, columns) unstructured (text,
audio, video, logs)
Petabytes to
Volume Gigabytes to Terabytes
Zettabytes
Generated very fast
Speed of Generated slowly
(seconds/milliseconds
Generation (hours/days)
)
Distributed systems
Storage Centralized servers
(e.g., cloud, clusters)
Needs 8
Difference Between Traditional Data and Big Data
Aspect Traditional Data Big Data
Difficult to manage due to size
Management Easier to manage and process
and complexity
Relationships are often unknown
Relationships Stable relationships between data
or hidden
Hadoop, Spark, Hive, Kafka,
Tools Used SQL, MS Excel, Oracle, MS Access
NoSQL tools
MongoDB, Cassandra, HBase,
Databases MySQL, Oracle DB, SQL Server
Amazon Redshift
Bank transaction logs, employee Tweets, IoT sensor data, CCTV
Examples
records videos, YouTube logs
Real-time fraud detection,
Use Case Monthly sales reports
predictive analytics
The key differences between traditional data and big data are related to the volume, variety, velocity,
complexity, and potential value of the data. Traditional data is typically small in size, structured, and
static, while big data is large, complex, and constantly changing. As a result, big data requires
specialized tools and techniques to manage and analyze effectively.
9
What is Big Data Analytics?

Big data analytics is the often


complex process of examining
big data to uncover information
-- such as hidden patterns,
correlations, market trends and
customer preferences that can help
organizations make informed
business decisions.

10
How is big data used for improving the companies?
1. Predict exactly what customers want before they ask for it

2. Get customers excited about their own data

3. Improve customer service interactions

4. Identify customer pain points and solve them

5. Reduce health care costs and improve treatment

11
CHARACTERISTICS OF BIG DATA:
Volume
● The name Big Data itself is related to an enormous size.
● Big Data is a vast 'volumes' of data generated from
many sources daily, such as business processes, machines, social media
platforms, networks, human interactions, Twitter data feeds, clickstreams
on a web page or a mobile app, or sensor-enabled equipment and
many more.

13
Value
Value refers to the usefulness or benefits you can extract from Big Data.
It answers the question:

“What meaningful insights or business advantages can we get from the data?”

Just having large data (Volume) is not enough.


If data doesn't give useful insights, it's just noise.

Value is an essential characteristic of big data


Value is the ultimate goal — turning raw data into meaningful, actionable
outcomes
● It is valuable and reliable data that we store, process, and also analyze.

Eg: How Value is Extracted from Big Data in E-commerce--Recommending products based on user behavior
(Amazon, Flipkart)

14
Variety

● Big Data can be structured, unstructured, and semi- structured


that are being collected from different sources.

● Data had been collected from databases and sheets in the past,

● But these days the data will comes in any forms, that are PDFs,
Emails, audios, SM posts, photos, videos, etc.

15
Variety

16
Velocity
Data is begin generated fast and need to be processed fast.
• Velocity plays an important role compared to others. Velocity creates the speed
by which the data is created in real-time. It contains the linking of incoming
data sets speeds, rate of change, and activity bursts. The primary
aspect of Big Data is to provide demanding data rapidly.

• Big data velocity deals with the speed at the data flows from sources like
application logs, business processes, networks,

• and social media sites, sensors, mobile devices, etc

17
Veracity
Veracity is a big data characteristic related to the trustworthiness, accuracy,
and quality of the data. It addresses the uncertainty, inconsistency, and
reliability of data being collected.

It also refers to the presence of:


errors
incomplete
missing values

Issues in the Data (Veracity Examples):


•Missing values (? in Min, Max, Mean columns)
•Illogical combinations (e.g., Min = 15000, Max = 7.9)
•Inconsistent standard deviation (e.g., 50 million is not realistic)
•Possibly faulty or noisy sensor/recorded data

18
Big Data: 3V’s

19
Some Make it 4V’s

20
5 V's of Big Data

•Volume
•Value
•Variety
•Velocity
•Veracity

21
Comparison: Traditional (Small) Data vs Big Data using 5 V’s

5 V's Traditional Data (Small Data / RDBMS) Big Data

Volume Data stored in MB, GB, TB Data stored in PB, EB (huge volume)
Data increases exponentially, may stream in
Velocity Data increases gradually, updated periodically
real-time
Mostly unstructured (images, videos, text, logs,
Variety Mostly structured (tables, rows, columns)
etc.)

Distributed, data from multiple sources – may


Veracity Centralized, easy to control data quality have inconsistencies(The trustworthiness,
accuracy, and quality of data.)

Stored locally, used for specific, smaller Globally present, used for broad insights, analytics
Value
applications & AI
Tools (Extra) SQL Server, Oracle (Single Node) Hadoop, Spark (Multinode Cluster)

22
Types of Bigdata
Following are the types of Big Data:
1. Structured
2. Unstructured
[Link]-structured

23
Types of Big data

24
Structured data
Structured data has certain predefined organizational properties and is present in
structured or tabular schema, making it easier to analyze and sort.
The business data of an e-commerce website can be considered to be structured data.

Un-Structured Data
Unstructured data accounts for the majority of big data and comprises information such as
dates, numbers, and facts.
Photos we upload on Facebook or Instagram and videos that we watch on YouTube
or any other platform contribute to the growing pile of unstructured data

25
Semi Structured Data
Semi-structured data is a hybrid of structured and unstructured data. This means that it
inherits a few characteristics of structured data but nonetheless contains information that
fails to have a definite structure and does not conform with relational databases or formal
structures of data models.

For instance, JSON and XML are typical examples of semi-structured data.

26
ORIGINS AND GROWTH OF BIG DATA:
Who’s Generating Big Data( Origins)

Sensor technology and network


(measuring all kinds of data)
Mobile devices
(tracking all objects all the time)

Scientific instruments Healthcare


(collecting all sorts of data)
Social media and networks Media &
(all of us are generating data) Entertainment

Big Data is generated by almost every sector of society today.

The progress and innovation is no longer hindered by the ability to collect data
But, by the ability to manage, analyze, summarize, visualize, and discover knowledge from the collected
data in a timely manner and in a scalable fashion
28
What’s driving Big Data( Growth)
What factors are causing Big Data to grow and become important?

Main Drivers:
Explosion of digital data (from apps, social media, IoT, etc.)
Need for real-time insights (instant decisions)
Demand for advanced analytics (predictive models, AI)
Cheaper storage & faster processing (cloud, big data tools)
Growth of AI & ML (which need lots of data to learn)

Big Data is being generated by everything connected to the internet — people, machines, apps
— and it’s growing fast because of the digital lifestyle we live in. Businesses need to keep up
with this data to stay competitive

29
HARNESSING BIG DATA WITH MODERN
TECHNOLOGIES AND PROCESSING
SYSTEMS??
:
Harnessing Big Data
Using advanced technologies and architectures to collect, store, process, and analyze vast
amounts of diverse and rapidly changing data to extract meaningful insights and make
data-driven decisions.

Key Technologies and Systems Involved in Harnessing Big Data:


1. Data Generation & Transaction Systems
These systems are where data starts — they create the data we work with later.

System Full Form Purpose Example


Online Transaction Captures daily
OLTP Train bookings, ATM withdrawals
Processing transactions
Online Analytical sales performance across regions for
OLAP Analyzes historical data
Processing the last 5 years
Real-Time Analytics Processes live data
RTAP Fraud detection, live tracking
Processing instantly

31
Technologies Available for Big Data
2. Data Ingestion Tools: Bring data from various sources into Big Data systems
Apache Kafka – Handles real-time streaming data
Apache Flume – Gathers log data (web, mobile, IoT)
Apache NiFi – Manages data flows with automation
3. Data Storage Technologies:Store huge volumes of structured and unstructured data
HDFS (Hadoop Distributed File System) – Distributed storage
NoSQL Databases – MongoDB, Cassandra for flexible storage
Cloud Storage – Amazon S3, Google Cloud Storage
4. Data Processing Engines:Process and compute data in batch or real time
Hadoop MapReduce – Batch processing (slow but reliable)
Apache Spark – Fast, in-memory processing
Apache Storm / Flink – Real-time event stream processing

32
Technologies Available for Big Data
5. Query & Analysis Tools
Analyze large datasets using queries or scripts
Hive, Pig – Run queries on Hadoop
Presto / Trino – Fast querying across big datasets
6. ETL & Data Pipeline Tools
Clean, transform, and move data between systems
Talend – ETL (Extract, Transform, Load) tool
Apache Airflow – Automates and schedules workflows
[Link] Visualization & BI Tools
Present data insights in charts, dashboards, and reports
Tableau, Power BI – Visual dashboards
Grafana – Real-time monitoring dashboards

33
Technologies Available for Big Data

34
Advantage of Big Data Analytics
Customer Acquisition and Retention
Big data helps companies understand what customers like by analyzing their online behavior
and shopping history, so they can attract and keep them with personalized offers.
Focused and Targeted Promotions
It lets businesses send the right ads to the right people at the right time, increasing sales and
saving money on marketing.
Potential Risks Identification
By studying past data, companies can spot possible problems early, like fraud or delays, and
take action before they grow.
Innovation
Analyzing data helps businesses find new ideas for products or services based on what
customers want and market trends.
Complex Supplier Networks
Big data gives clear info on suppliers, helping companies manage inventory, cut costs, and
avoid delays in delivery.
35
Advantage of Big Data Analytics
Cost Optimization
It shows where money is being wasted, so businesses can fix inefficiencies and save on
operations.
Improve Efficiency
Data helps identify slow or repetitive tasks in a business, so they can be automated or
improved for smoother work.
Better Decision Making
With accurate and timely data, businesses can make smarter, confident decisions instead of
guessing.
Enhanced Customer Experience
By knowing customer needs better, businesses can offer personalized services, making
customers happier and more loyal.
Competitive Advantage
Using big data smartly helps businesses stay ahead of rivals by adapting faster and making
better choices.
36
Case study :Big data at Wal-Mart
Walmart, one of the largest retailers in the world, continues to leverage Big Data to
understand consumer behavior, optimize inventory, and drive sales.
Walmart's advanced analytics platform processes millions of transactions daily across its
global stores and online platforms. By analyzing this vast data, they predict customer
purchasing patterns, optimize supply chain logistics, and provide personalized offers.
Recent Innovation:
Walmart has been integrating AI and Big Data to optimize product recommendations, tailor
advertisements, and enhance customer engagement. For instance, during the 2023 holiday
season, Walmart used Big Data to forecast the types of products that would be in high
demand (such as video games, electronics, and holiday-specific items) based on historical
data and current trends.
Impact:
The use of data-driven insights improved inventory management, leading to reduced
stockouts, more personalized shopping experiences, and increased sales during peak periods
like Black Friday and Christmas.
37
Case study :Big data at Netflix
Netflix is a prime example of a company using Big Data for personalization and improving user
experience.
What Happened:
Netflix collects data on what content users are watching, how they interact with the
platform, and what time of day they watch. It also considers factors such as rating patterns,
genres, and search history.
Recent Innovation:
In 2023, Netflix introduced an AI-driven personalization engine that provides
hyper-targeted content recommendations. The system processes billions of data points
from over 200 million subscribers and recommends content that a user is most likely to
watch based on similar patterns of behavior from other viewers.
Impact:
The recommendations system drives over 80% of all content watched on Netflix, improving
user retention and helping the company save on advertising costs. This data-driven strategy
helps retain customers and ensures higher engagement rates.
38
Challenges in Handling Big Data

1. Storing exponentially growing huge


datasets: Focuses on the volume of data that
needs to be saved and managed

2. Processing data having complex


structure: Deals with the variety and format of
data (text, images, videos, logs, etc.).

3. Processing data faster.


Highlights the velocity aspect — the need for fast data
handling

4. The Bottleneck is in technology


Refers to limitations in current tools, infrastructure, and
systems that can't keep up.

5. Also in technical skills


Points to a human resource challenge, where expertise is
lacking to manage Big Data systems effectively

39
Hadoop
Hadoop is an open-source software framework that is used for
storing and processing large amounts of data in a distributed
computing environment.

Hadoop is designed to process large volumes of data (Big Data)


across many machines without relying on a single machine.

It is built to be scalable, fault-tolerant and cost-effective. Instead of


relying on expensive high-end hardware, Hadoop works by
connecting many inexpensive computers (called nodes) in a cluster.
40
Core Hadoop Component

● HDFS(Hadoop Distributed
File System)

● Map-Reduce

● Hadoop YARN(Yet another


resource negotiator)

● Hadoop-Common or
Common Utilities Core Hadoop Component

41
Core Hadoop Component
1. MapReduce
MapReduce is a data processing model in Hadoop built on top of the YARN framework. Its
main feature is distributed and parallel processing across a Hadoop cluster, which significantly
boosts performance when handling Big Data. Since serial processing is inefficient at large scale,
MapReduce enables fast and efficient data processing by splitting the workload into two
phases: Map and Reduce.

42
Core Hadoop Component
MapReduce Workflow: The workflow begins when input data is passed to the Map() function.
This data, often stored in blocks, is broken down into tuples (key-value pairs). These pairs are
then passed to the Reduce() function, which aggregates them based on the keys and performs
required operations such as sorting, counting or summing to generate the final output. The
result is then written to HDFS.

2. HDFS
HDFS(Hadoop Distributed File System) is utilized for storage permission. It is mainly
designed for working on commodity Hardware devices(inexpensive devices), working on a
distributed file system design. HDFS is designed in such a way that it believes more in storing
the data in a large chunk of blocks rather than storing small data blocks.

HDFS in Hadoop provides Fault-tolerance and High availability to the storage layer and the
other devices present in that Hadoop cluster.
43
Core Hadoop Component
HDFS Architecture Components
1. NameNode (Master Node):
Manages the HDFS cluster and stores metadata (file names, sizes, permissions, block
locations).
Does not store actual data, only metadata (via edit logs and fsimage).
Controls file operations like create, delete, and replicate, directing DataNodes accordingly.
Helps clients locate the nearest DataNode for fast access.
2. DataNode (Slave Node):
Stores the actual data blocks on local disks.
Handles read/write requests from clients.
Sends heartbeats and block reports to the NameNode regularly.
Data is replicated (default 3 copies) for reliability.

44
Core Hadoop Component
4. Hadoop Common (Common Utilities)
Hadoop Common, also known as Common Utilities, includes the core Java libraries and
scripts required by all the components in a Hadoop ecosystem such as HDFS, YARN and
MapReduce. These libraries provide essential functionalities like:
File system and I/O operations
Configuration and logging
Security and authentication
Network communication

Hadoop Common ensures that the entire cluster works cohesively and that hardware failures,
which are common in Hadoop's commodity hardware setup, are handled automatically by the
software

45
Core Hadoop Component
3. YARN (Yet Another Resource Negotiator)

YARN is the resource management layer in the Hadoop ecosystem. It allows multiple data
processing engines like MapReduce, Spark, and others to run and share cluster resources
efficiently. It handles two core responsibilities:

Job Scheduling: Breaks a large task into smaller jobs and assigns them to various nodes in
the Hadoop cluster. It also tracks job priority, dependencies, and execution time.
Resource Management: Allocates and manages the cluster resources (CPU, memory, etc.)
required for running these jobs.

46
Advantages and Disadvantages of Hadoop

Advantages:
Scalability: Easily scale to thousands of machines.
Cost-effective: Uses low-cost hardware to process big data.
Fault Tolerance: Automatic recovery from node failures.
High Availability: Data replication ensures no loss even if nodes fail.
Flexibility: Can handle structured, semi-structured and unstructured data.
Open-source and Community-driven: Constant updates and wide support.
Disadvantages:
Not ideal for real-time processing (better suited for batch processing).
Complexity in programming with MapReduce.
High latency for certain types of queries.
Requires skilled professionals to manage and develop.

47
Applications Of Hadoop
Hadoop is used across a variety of industries:
Banking: Fraud detection, risk modeling.
Retail: Customer behavior analysis, inventory management.
Healthcare: Disease prediction, patient record analysis.
Telecom: Network performance monitoring.
Social Media: Trend analysis, user recommendation engines.

48
How Does Hadoop Work?
Here’s a overview of how Hadoop operates:
A large data file is uploaded to HDFS.
HDFS splits the file into fixed-size blocks (e.g., 128MB).
These blocks are distributed and stored across multiple machines (DataNodes).
Each block is replicated for fault tolerance.
The NameNode keeps track of where each block is stored.

A MapReduce job is submitted to process the data.


The Map tasks run on the DataNodes where the data blocks are located.
Map tasks produce intermediate key-value pairs.
The Reduce tasks aggregate and summarize the data.
The final result is stored back into HDFS or exported elsewhere.

49
High Level Architecture Of Hadoop
Hadoop works internally, using a Master-Slave Architecture to store
and process big data. It involves:
Master Node
Slave Nodes

Hadoop 1 Architecture Hadoop 2 architecture (with YARN-resource manager)


50
Architecture Of Hadoop
Master Node
The master node contains two main components:
NameNode
Manages the HDFS (Hadoop Distributed File System).
Keeps track of where files are stored (metadata).
Does not store actual data, just the directory and block info.
Resource Manager
Part of YARN (Yet Another Resource Negotiator).
Allocates resources (CPU, RAM) to different tasks.
Sends jobs to slave nodes for processing.

Think of the Master as the Manager – it tells other workers (slaves) what to do and where
everything is.
51
Architecture Of Hadoop
Slave Nodes
Each slave node contains four components:
DataNode : Stores actual data blocks (pieces of big files).
Reports to the NameNode about what data it has.
Node Manager :Works under the Resource Manager.
Monitors and manages jobs (Map or Reduce tasks) on this machine.
Map Task:This is the first phase of data processing.
It processes data blocks and produces key-value pairs.
Reduce Task:This is the second phase of processing.
It aggregates or summarizes the data output from the Map phase.
Think of each slave as a worker:
They store data (DataNode)
And process tasks (Map/Reduce) when assigned by the master.

52
limitations of Hadoop
Handling Small Files:
Hadoop struggles with many small files because it's designed for large files. This can overwhelm the system and
slow things down.
Slow Processing Speed:
Hadoop's MapReduce is slow because it relies on reading and writing data to disk, which creates latency.
Batch Processing Only:
Hadoop only supports batch processing, which means it processes data in large chunks, not in real-time.
No Real-Time Data Processing:
Hadoop is not suited for real-time data, making it unsuitable for applications that need fast, continuous data
processing.
No Support for Iterative Processing:
Hadoop is inefficient for tasks that need repeated data processing, such as machine learning, because it doesn't
handle iterative steps well.
High Latency:
The time it takes for Hadoop to process and return results (latency) can be high due to the multiple steps involved
in the MapReduce process.
Difficult to Use:
Writing MapReduce jobs requires complex programming, making Hadoop harder to use, especially for beginners.
53
limitations of Hadoop
Security Issues:
Hadoop lacks strong out-of-the-box security features and requires complex configuration for
secure operations.
No Abstraction:
Hadoop doesn't offer high-level abstractions, meaning developers have to deal with
low-level details like handling data manually.
Vulnerability to Attacks:
Since Hadoop is primarily written in Java, it inherits Java's security risks, making it a target
for cyber-attacks.
Unpredictable Job Completion Time:
Hadoop doesn't provide a guarantee on when jobs will finish, making it unreliable for
time-sensitive tasks.

54
Hadoop Ecosystem
The Hadoop Ecosystem is a collection of tools that work with
Hadoop to handle different types of big data tasks
Ingesting data (getting data into Hadoop) such as:

•Ingesting data (getting data into Hadoop)


•Storing data
•Processing data
•Analyzing/querying data
•Visualizing results

These tools help Hadoop handle real-world big data problems more
efficiently.

55
Hadoop Ecosystem

56
Hadoop Ecosystem

57
Sqoop
A lot of applications still store data in relational databases, thus making them a very important
source of data. Therefore, Sqoop plays an important part in bringing data from Relational
Databases into HDFS.
The commands written in Sqoop internally converts into MapReduce tasks that are executed
over HDFS. It works with almost all relational databases like MySQL, Postgres, SQLite, etc. It
can also be used to export data from HDFS to RDBMS.
Apache Sqoop, a command-line interface tool, moves data between relational databases
and Hadoop.
● It is used to export data from the Hadoop file system to relational
databases andto import data from relational databases such as MySQL and
Oracle into the Hadoop file system.

58
Flume
● Flume is an open-source, reliable, and available service used to efficiently collect, aggregate,
and move large amounts of data from multiple data sources into HDFS. It can collect data in
real-time as well as in batch mode. It has a flexible architecture and is fault-tolerant with
multiple recovery mechanisms.
● Apache Flume is a distributed system for collecting, aggregating, and transferring data
from multiple sources to a centralized data store.
● Apache Flume can be used to transport large amounts of social-media
generated data, logs, network traffic data, email messages, and many more to a centralized data
store.
● In simple words, Apache Flume is a tool in the Hadoop ecosystem for
transferring data from one location to anotherefficiently andreliably. The
main goal of Apache Flume is to transfer data from applications to Hadoop HDFS.
● It is highly robust, reliable, and fault-tolerant.
Kafka
59
Pig
Pig was developed for analyzing large datasets and overcomes the difficulty to write
map and reduce functions. It consists of two components: Pig Latin and Pig Engine.

Pig Latin is the Scripting Language that is similar to SQL. Pig Engine is the execution
engine on which Pig Latin runs. Internally, the code written in Pig is converted to
MapReduce functions and makes it very easy for programmers who aren’t proficient
in Java.

Apache Pig is a high-level language platform for analyzing and querying huge
dataset that are stored in HDFS. Pig as a component of Hadoop Ecosystem uses
PigLatin language. It is very similar to SQL. It loads the data, applies the required
filters and dumps the data in the required format. For Programs execution, pig
requires Java runtime environment.

60
Hive
It allows for easy reading, writing, and managing files on HDFS. It has its own querying
language for the purpose known as Hive Querying Language (HQL) which is very similar to
SQL.
The Hadoop ecosystem component, Apache Hive, is an open source data warehouse system
for querying and analyzing large datasets stored in Hadoop files. Hive do three main
functions: data summarization, query, and analysis.
Hive use language called HiveQL (HQL), which is similar to SQL. HiveQL automatically
translates SQL-like queries into MapReduce jobs which will execute on Hadoop.

61
HBase
Apache HBase is a Hadoop ecosystem component which is a distributed database.
HBase is a Column-based NoSQL database. It runs on top of HDFS and can handle any type of
data. It allows for real-time processing and random read/write operations to be performed in
the data
HBase is scalable, distributed, and NoSQL database that is built on top of HDFS. HBase,

provide real-time access to read or write data in HDFS.

HBase is a column-oriented, non-relational database. This means that data is stored in

individual columns, and indexed by a unique row key. This architecture allows for rapid

retrieval of individual rows and columns and efficient scans over individual columns within a

table.
62
Mahout
● It is a powerful, scalable machine-learning library that runs on top of Hadoop
MapReduce.
● Mahout is supported by its 3 pillars:
○ Recommender engines: Recommenders can be classified as being user based or item
based and can be used to attract users and suggest products by mining user behaviour.
○ Clustering: Clustering groups objects of a similar nature in one place.
○ Classification: Classification techniques decide whether a thing deserves to be a part
of some type or not.

63
Hadoop Ecosystem
Oozie
:Oozie is a workflow scheduler system that allows users to link jobs written on various
platforms like MapReduce, Hive, Pig, etc. Using Oozie you can schedule a job in advance and
can create a pipeline of individual jobs to be executed sequentially or in parallel to achieve a
bigger task

Zookeeper
In a Hadoop cluster, coordinating and synchronizing nodes can be a challenging task.
Therefore, Zookeeper is the perfect tool for the problem.
It is an open-source, distributed, and centralized service for maintaining configuration
information, naming, providing distributed synchronization, and providing group services
across the cluster.

64
Spark
Spark is an alternative framework to Hadoop built on Scala but supports varied applications
written in Java, Python, etc. Compared to MapReduce it provides in-memory processing
which accounts for faster processing. In addition to batch processing offered by Hadoop, it
can also handle real-time processing.

65
Thank
You

You might also like