0% found this document useful (0 votes)
14 views50 pages

Understanding Big Data and Analytics

The document provides an overview of Big Data and its characteristics, defining it as large and complex datasets that traditional tools struggle to process, characterized by the 5 V's: Volume, Velocity, Variety, Veracity, and Value. It discusses the evolution of data storage and processing from the 1970s to the present, highlighting the rise of unstructured data and the importance of Big Data Analytics for informed business decisions. Additionally, it outlines the challenges associated with Big Data, including exponential growth, cloud computing complexities, and data retention issues.

Uploaded by

poojithag122
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views50 pages

Understanding Big Data and Analytics

The document provides an overview of Big Data and its characteristics, defining it as large and complex datasets that traditional tools struggle to process, characterized by the 5 V's: Volume, Velocity, Variety, Veracity, and Value. It discusses the evolution of data storage and processing from the 1970s to the present, highlighting the rise of unstructured data and the importance of Big Data Analytics for informed business decisions. Additionally, it outlines the challenges associated with Big Data, including exponential growth, cloud computing complexities, and data retention issues.

Uploaded by

poojithag122
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Department of Computer Science and Engineering

Bigdata and Analytics (BCS714D)

Module-01
Data: It refers to raw, unorganized facts or figures.

• It can be numbers, words, measurements, or observations.

Example: "John, 25, Male" — this is data.

.IN

Information is processed or organized data that is meaningful and useful.

• It is data given context.


C
• Example: "John is a 25-year-old male customer." — This is information derived from
N
data.
SY

What is Big Data?

• Big Data refers to extremely large and complex datasets that traditional tools can't
process efficiently.
U

• It is characterized by 5 V’s:
VT

o Volume: Huge amount of data.

o Velocity: Speed at which data is generated.

o Variety: Different formats (text, video, images, etc.).

o Veracity: Uncertainty of data accuracy.

o Value: Extracting useful insights.

Big Data Analytics: Big Data Analytics is the process of examining big data to uncover
patterns, trends, and insights.

• It helps businesses make informed decisions.

• Uses tools like Hadoop, Spark, Hive, Pig, etc.

• Example: Analyzing customer purchase history to recommend products.

Prof. Deepika G , Dept. of CSE, SVIT Page 1


Studied smart, not hard — thanks to [Link]
� 2.1 Characteristics of Data
Three key characteristics of data:

1. Composition

o Structure of data: sources, types, granularity.

o Static vs. real-time data.

2. Condition

.IN
o State of the data (clean, raw, needs enhancement).

o Questions: Can it be used as-is? Does it need processing?


C
3. Context
N
o Background of the data: where, when, why was it generated?
SY
U
VT

῿ 2.2 Evolution of Big Data


Overview:

• Before 1970s: Data was primitive and structured; stored using mainframes.

• 1980s–1990s: Rise of relational databases; data became relational and used in data-
intensive applications.

Prof. Deepika G , Dept. of CSE, SVIT Page 2


Studied smart, not hard — thanks to [Link]
• 2000s & beyond: Explosion of complex, unstructured data due to Web, IoT, and
multimedia technologies.

Table 2.1 – Evolution of Big Data

Data Generation &


Time Period Data Utilization Data Driven Output
Storage

1970s & Primitive and Mainframes for basic data

.IN

before Structured storage

1980s & Complex and Relational databases; data-



1990s Relational
C
intensive apps
N
2000s and Complex and Structured, unstructured,

beyond Unstructured multimedia data
SY

Key Contributors to Modern Big Data


U

• World Wide Web (WWW) – Opened floodgates for global digital content.
VT

• Internet of Things (IoT) – Devices constantly generate vast, diverse data streams.

῿ 2.3 Definition of Big Data


Common Understandings of Big Data

Several responses have been widely accepted to define Big Data:

1. Beyond Infrastructure:
Big Data refers to anything that exceeds traditional storage, processing, and
analysis capabilities of existing human and technical infrastructure.

2. Relative to Time:
What is considered "big" today may become normal or standard tomorrow, indicating
that Big Data is a moving threshold.

3. Large Volume:
It refers to enormous data sizes – Terabytes, Petabytes, Zettabytes, and even
Yottabytes.

Prof. Deepika G , Dept. of CSE, SVIT Page 3


Studied smart, not hard — thanks to [Link]
4. The 3 Vs – A standard definition framework.

Standard 3Vs Definition (Gartner’s Perspective)

According to Gartner IT Glossary and Doug Laney (2001), Big Data is defined by:

Big Data is high-volume, high-velocity, and high-variety information assets that demand
cost-effective, innovative forms of information processing for enhanced insight and
decision making.

• Volume: Huge amounts of data.

.IN
• Velocity: Speed of data generation and processing.

• Variety: Different types and formats (structured, unstructured, semi-structured).


C
N
Key Conceptual Flow
SY

Big Data ➝ Information ➝ Actionable Intelligence ➝ Better Decisions ➝ Enhanced


Business Value
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 4


Studied smart, not hard — thanks to [Link]
1. Volume-Based Definition

• Focuses on the massive size of data.

• Big Data is often defined by its scale, measured in terabytes, petabytes, zettabytes, or
even yottabytes.

2. Infrastructure-Based Definition

• Refers to data that exceeds the capacity of current tools or infrastructure.

.IN
• Suggests Big Data begins where traditional systems struggle to store, process, or
analyze data efficiently. C
N
3. Relativity Definition
SY

• Emphasizes the relative and evolving nature of Big Data.

• What we call “big” today may become standard in the near future due to technological
advances.
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 5


Studied smart, not hard — thanks to [Link]
Definition of Big Data (Gartner’s Perspective)

Three-Part Definition (3Vs Model):

"Big Data is high-volume, high-velocity, and high-variety information assets that


demand cost-effective, innovative forms of information processing for enhanced insight
and decision making."

⿿ Part I: High-Volume, High-Velocity, High-Variety

• Refers to large and fast-growing data.

.IN
• Includes:

o Structured, semi-structured, and unstructured data.


C
• Requires:
N
o Fast storage, processing, and analysis capabilities.
SY

Part II: Cost-Effective & Innovative Processing

• Focuses on:
U

o Using new technologies and methods to: Ingest, Store, Process, Persist,
Integrate, Visualize.
VT

• Goal:

o Efficient handling of large, fast, and diverse data sets.

Part III: Enhanced Insight & Decision-Making

• Involves:

o Gaining deep, rich, and actionable insights.

o Using those insights for better and faster decisions.

• Leads to:

o Improved business value and competitive advantage.

Prof. Deepika G , Dept. of CSE, SVIT Page 6


Studied smart, not hard — thanks to [Link]
Flow Summary:

Data → Information → Actionable Intelligence → Better Decisions → Enhanced Business


Value

Challenges with Big Data

1. Exponential Data Growth

o Data volume is increasing rapidly, mostly generated in the last 2-3 years.

.IN
o Raises key questions:

▪ Will all this data be useful for analysis?

Should we work with all data or a subset?



C
▪ How to separate valuable knowledge from noise?
N
2. Cloud Computing and Virtualization
SY

o Cloud is essential for cost-efficient, elastic, and scalable infrastructure.

o However, deciding whether to host big data solutions on-premises or in the cloud
U

is complex.

3. Data Retention Period


VT

o Determining how long to keep big data is tricky.

o Some data is valuable for long-term decisions.

o Other data becomes irrelevant or obsolete quickly (sometimes within hours).

Prof. Deepika G , Dept. of CSE, SVIT Page 7


Studied smart, not hard — thanks to [Link]
.IN
C
N
SY

2.5 What is Big Data?


U

• Big data is defined by three main characteristics:


VT

o Volume (amount of data)

o Velocity (speed of data generation and processing)

o Variety (different types and formats of data)

2.5.1 Volume

• Data size has grown exponentially, evolving through units:

o Bits → Bytes → Kilobytes → Megabytes → Gigabytes → Terabytes → Petabytes →


Exabytes → Zettabytes → Yottabytes

• The sources of big data are diverse and can include:

o XLS, DOC, PDF (unstructured data)

o Videos (e.g., YouTube)

o Chat conversations (e.g., Internet Messenger)

o Customer feedback forms on retail websites


Prof. Deepika G , Dept. of CSE, SVIT Page 8
Studied smart, not hard — thanks to [Link]
[Link] Where does Data get generated?

Data Characteristics Illustrated:

• Data Volume (X-axis, moving outward):

o Data size grows from small (MB, GB) to very large (TB, PB).

o The more data you have, the further out on the volume axis it is.

• Data Velocity (Y-axis, moving upward):

.IN
o Speed of data generation and processing:

▪ From Batch processing (slowest) to Periodic, Near real-time, and Real-time


(fastest).
C
N
• Data Variety (Z-axis, moving diagonally):
SY

o Types of data include:

▪ Table, Database (structured data)

Photo, Social, Web, Video, Audio, Mobile (mostly unstructured or semi-


U


structured data)
VT

Data Growth and Units (Table 2.2)

• Data Size Units (Binary scale):

o Bit = 0 or 1

o Byte = 8 bits

o Kilobyte (KB) = 1024 bytes

o Megabyte (MB) = 1024² bytes

o Gigabyte (GB) = 1024³ bytes

o Terabyte (TB) = 1024⁴ bytes

o Petabyte (PB) = 1024⁵ bytes

o Exabyte (EB) = 1024⁶ bytes

Prof. Deepika G , Dept. of CSE, SVIT Page 9


Studied smart, not hard — thanks to [Link]
o Zettabyte (ZB) = 1024⁷ bytes

o Yottabyte (YB) = 1024⁸ bytes

• Data Size Units (Decimal scale as per Figure 2.6):

o 1 KB = 1,000 bytes

o 1 MB = 1,000,000 bytes

o 1 GB = 1,000,000,000 bytes

.IN
o 1 TB = 1,000,000,000,000 bytes

o 1 PB = 1,000,000,000,000,000 bytes C
o 1 EB = 1,000,000,000,000,000,000 bytes
N
o 1 ZB = 1,000,000,000,000,000,000,000 bytes

o 1 YB = 1,000,000,000,000,000,000,000,000 bytes
SY

Note: Binary units are based on powers of 2 (1024), while decimal units are powers of 10 (1000).
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 10


Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 11


Studied smart, not hard — thanks to [Link]
Sources of Big Data (Figure 2.7)
.IN
C
• Internal Data Sources:
N
o Data storage: File systems, SQL databases (Oracle, MySQL, PostgreSQL,
SY

MongoDB, Cassandra, etc.)

o Archives: Scanned documents, paper archives, customer correspondence, health


records, admission records, assessment records, etc.
U

• External Data Sources:


VT

o Public Web: Wikipedia, weather, regulatory data, compliance, census, etc.

• Both Internal + External:

o Sensor data: Car sensors, smart meters, HVAC systems, refrigerators, etc.

o Machine log data: Event logs, app logs, business process logs, audit logs,
clickstream data.

o Social media: Twitter, Facebook, LinkedIn, YouTube, Instagram.

o Business apps: ERP, CRM, HR systems, Google Docs.

o Media: Audio, video, images, podcasts.

o Docs: CSV files, Word documents, PDFs, XLS, PPT.

Prof. Deepika G , Dept. of CSE, SVIT Page 12


Studied smart, not hard — thanks to [Link]
.IN
C
N
External Data Sources
SY

• Data that resides outside an organization’s firewall.

• Examples include:
U

o Public Web: Wikipedia, weather data, regulatory info, compliance, census data,
etc.
VT

Both (Internal + External Data Sources)

• Data that can come from both inside and outside the organization.

• Examples include:

➢ Sensor data: Car sensors, smart electric meters, office building sensors, air
conditioning units, refrigerators, etc.

➢ Machine log data: Event logs, application logs, business process logs, audit logs,
clickstream data, etc.

➢ Social media: Twitter, blogs, Facebook, LinkedIn, YouTube, Instagram, etc.

➢ Business apps: ERP (Enterprise Resource Planning), CRM (Customer Relationship


Management), HR systems, Google Docs, etc.

➢ Media: Audio, video, images, podcasts, etc.

Prof. Deepika G , Dept. of CSE, SVIT Page 13


Studied smart, not hard — thanks to [Link]
➢ Docs: CSV (Comma Separated Values), Word documents, PDF, XLS, PPT, and
similar files.

2.5.2 Velocity

• Velocity refers to the speed at which data is processed.

• There's been a shift from batch processing (e.g., traditional payroll applications) to real-
time processing.

• The progression in data processing speed is:

.IN
Batch → Periodic → Near real-time → Real-time processing

2.5.3 Variety


C
Variety deals with the different types and formats of data.

• It classifies data into three categories:


N
1. Structured data: Comes from traditional systems like transaction processing
SY

systems and RDBMS (Relational Database Management Systems).

2. Semi-structured data: Examples include HTML, XML (Hyper Text Markup


Language and eXtensible Markup Language).
U

3. Unstructured data: Includes unstructured text documents, audio, video, emails,


VT

photos, PDFs, social media content, etc.

3.2 WHAT IS BIG DATA ANALYTICS?

Big Data Analytics is…

1. Technology-enabled analytics:
Uses various data analytics and visualization tools from vendors like IBM, Tableau, SAS, R Analytics, etc.,
to process and analyze big data.
2. Gaining meaningful insights:
Helps businesses gain deeper, richer insights to better understand customers (demographics, preferences) for
cross-selling, up-selling, and better vendor/supplier management.
Example: Online retailers recommend products based on stored purchase and preference data.
3. Competitive edge:
Enables quicker and better decision-making, giving an advantage over competitors.

Prof. Deepika G , Dept. of CSE, SVIT Page 14


Studied smart, not hard — thanks to [Link]
4. Collaboration:
Involves close cooperation between IT teams, business users, and data scientists.
5. Handling large datasets:
Deals with datasets whose volume and variety surpass current storage and processing capabilities of
enterprises.
6. Moving code to data:
Instead of moving data around, programs (small in size) are moved close to where data resides (especially
large datasets in terabytes or petabytes), improving efficiency. This will become more important as data

.IN
grows to exabytes or zettabytes.

3.5 CLASSIFICATION OF ANALYTICS:


C
1. Classification by stages of analytics:
N
o Basic
SY

o Operationalized

o Advanced
U

o Monetized
VT

2. Classification by versions of analytics:

o Analytics 1.0

o Analytics 2.0

o Analytics 3.0

Figure 3.5 illustrates what big data entails, showing a cycle:

• More data produced →

• More data stored →

• More data analyzed →

• Better predictions →

• Steady growth of analysis → (loops back to more data produced)

This cycle highlights how increasing data leads to improved insights and predictions, fueling
further data generation and analysis.
Prof. Deepika G , Dept. of CSE, SVIT Page 15
Studied smart, not hard — thanks to [Link]
.IN
C
N
3.5.1 First School of Thought
SY

1. Basic analytics: This primarily is slicing and dicing of data to help with basic business
insights. This is about reporting on historical data, basic visualization, etc.
U

2. Operationalized analytics: It is operationalized analytics if it gets woven into the


enterprise’s business processes.
VT

3. Advanced analytics: This largely is about forecasting for the future by way of predictive
and prescriptive modeling.

4. Monetized analytics: This is analytics in use to derive direct business revenue.

3.5.2 Second School of Thought

Table 3.1 Analytics 1.0, 2.0, and 3.0


Analytics 1.0 Analytics 2.0 Analytics 3.0

Era: Mid 1950s to 2009 2005 to 2012 2012 to present

Descriptive + predictive + prescriptive


statistics (use data from the past to
Descriptive statistics (report Descriptive statistics + predictive
make prophecies for the future and
Stats used: on events, occurrences, etc. statistics (use data from the past to
at the same time make
of the past) make predictions for the future)
recommendations to leverage the
situation to one’s advantage)

Prof. Deepika G , Dept. of CSE, SVIT Page 16


Studied smart, not hard — thanks to [Link]
Analytics 1.0 Analytics 2.0 Analytics 3.0

What will happen? When will it


Key
What happened? Why did it happen? Why will it happen? What
questions What will happen? Why will it happen?
happen? should be the action taken to take
asked:
advantage of what will happen?

Data from legacy systems, Big data is being taken up seriously.


A blend of big data and data from
ERP, CRM, and 3rd party Data is mainly unstructured, arriving at
legacy systems, ERP, CRM, and 3rd
applications. Small and a much higher pace. This fast flow of
Data party applications. A blend of big data

.IN
structured data sources. data entailed that the influx of big
sources: and traditional analytics to yield
Data stored in enterprise volume data had to be stored and
insights and offerings with speed and
data warehouses or data processed rapidly, often on massive
impact.
marts. parallel servers running Hadoop.
C
Data is both being internally and
N
Sourcing: Data was internally sourced. Data was often externally sourced.
externally sourced.
SY

In memory analytics, in database


Database appliances, Hadoop clusters,
Technology: Relational databases processing, agile analytical methods,
SQL to Hadoop environments, etc.
machine learning techniques, etc.
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 17


Studied smart, not hard — thanks to [Link]
3.8 WHY IS BIG DATA ANALYTICS IMPORTANT?

1. Reactive – Business Intelligence:


What does Business Intelligence (BI) help us with? It allows the businesses to make
faster and better decisions by providing the right information to the right person at the
right time in the right format. It is about analysis of the past or historical data and then
displaying the findings of the analysis or reports in the form of enterprise dashboards,
alerts, notifications, etc. It has support for both pre-specified reports as well as ad hoc
querying.

2. Reactive – Big Data Analytics:

.IN
Here the analysis is done on huge datasets but the approach is still reactive as it is still
based on static data.

3. Proactive – Analytics:
C
This is to support futuristic decision making by the use of data mining, predictive
N
modeling, text mining, and statistical analysis. This analysis is not on big data as it still
uses the traditional database management practices on big data and therefore has
SY

severe limitations on the storage capacity and the processing capability.

4. Proactive – Big Data Analytics:


This is sieving through terabytes, petabytes, exabytes of information to filter out the
U

relevant data to analyze. This also includes high performance analytics to gain rapid
insights from big data and the ability to solve complex problems using more data.
VT

3.12 TERMINOLOGIES USED IN BIG DATA ENVIRONMENTS

1. In-Memory Analytics

• Data is fetched from RAM (Random Access Memory) instead of slow hard disk storage.

• Pre-processed and frequently-used data (e.g., cubes, aggregates) is stored in-memory.

• Enables faster access, quicker deployment, better insights, and minimal IT involvement.

• Eliminates repeated disk reads for analysis.

2. In-Database Processing

• Also known as in-database analytics.

• Combines data warehouses with analytical systems.

• Analytical computations are done within the database itself.

Prof. Deepika G , Dept. of CSE, SVIT Page 18


Studied smart, not hard — thanks to [Link]
• Avoids time-consuming export of data.

• Saves time and supports complex/extensive computations.

3. Symmetric Multiprocessor System (SMP)

• Uses a single main memory shared by two or more identical processors.

• All processors:

o Have access to I/O devices.

o Are controlled by a single OS instance.

.IN
• Processors are tightly coupled.

• Each processor has its own high-speed cache memory.


C
• Communication happens via a system bus.
N
4. Massively Parallel Processing (MPP)
SY

• Involves multiple processors working in parallel on different parts of the same program.

• Each processor:

o Has its own OS and dedicated memory.


U

o Communicates through messaging interfaces.


VT

• Suitable for large-scale data analytics.

• More complex to program than SMP.

• Referred to as loosely coupled systems (unlike tightly-coupled SMP).

5. Difference Between Parallel and Distributed Systems

• Parallel systems:

o Are tightly coupled.

o Use shared memory.

o User is unaware of how tasks are parallelized.

o Common in parallel databases.

• Distributed systems (explained later in text):

o Are loosely coupled.

o Use multiple systems possibly at different locations.

Prof. Deepika G , Dept. of CSE, SVIT Page 19


Studied smart, not hard — thanks to [Link]
o Require coordination over a network.

.IN
C
N
SY
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 20


Studied smart, not hard — thanks to [Link]
3.12.6 Shared Nothing Architecture

There are three common types of multiprocessor architecture used in high transaction rate
systems:

1. Shared Memory (SM)

• All processors share a common central memory.

• Suitable for systems needing tight coordination.

• Memory access is faster but can cause contention (conflicts).

.IN
2. Shared Disk (SD)

• Each processor has its own private memory.


C
• Disks are shared among multiple processors.
N
• Coordination required to manage access to shared disks.
SY

3. Shared Nothing (SN)

• No sharing of memory or disk among processors.


U

• Each processor has its own private memory and disk.


VT

• Ideal for scalability – processors operate independently.

• Most distributed systems (like Hadoop) follow this model.

Diagram – Figure 3.12: Distributed System

• Shows multiple processors (P1, P2, P3), each serving its own set of users.

• Each processor has its own storage.

• Processors communicate over a network.

• Represents a Shared Nothing Architecture, where processors work independently and


manage their own data.

Prof. Deepika G , Dept. of CSE, SVIT Page 21


Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT

[Link] Advantages of a “Shared Nothing Architecture”

Fault Isolation

• A "Shared Nothing Architecture" helps isolate faults effectively.

• If a fault occurs in a node, it is contained within that specific node.

• Other nodes remain unaffected since communication is done only through messages (or
the absence of them).

Prof. Deepika G , Dept. of CSE, SVIT Page 22


Studied smart, not hard — thanks to [Link]
Scalability

• In shared systems, disk bandwidth and controller become bottlenecks due to resource
sharing.

• In a shared nothing setup, each node has its own resources, eliminating contention.

• No synchronization is needed for shared data access, so nodes can scale easily without
performance degradation.

3.12.7 CAP Theorem

.IN
The CAP Theorem (also known as Brewer's Theorem) states:

In a distributed computing environment, it is impossible for a system to simultaneously


C
guarantee all three of the following:

1. Consistency
N
2. Availability
SY

3. Partition Tolerance

At best, only two of the three can be fully achieved.


U

[Link] CAP Theorem – Term Explanations


VT

1. Consistency

o Every read receives the most recent (last) write.

o No stale or outdated data is returned.

2. Availability

o Every request (read or write) receives a successful response (not necessarily the
most recent).

o The system is always available to serve requests.

3. Partition Tolerance

o The system continues to function even if communication between nodes is lost


(i.e., network failure).

o No single point of failure crashes the entire system.

Prof. Deepika G , Dept. of CSE, SVIT Page 23


Studied smart, not hard — thanks to [Link]
Diagram Summary (Figure 3.14)

A triangle with each corner representing one aspect of the CAP Theorem:

• Consistency

• Availability

• Partition Tolerance

.IN
C
N
SY

Training Institute Example


U

You visit Amey (Office Admin) to check your schedule:

1. You ask Amey if you're scheduled to train at 3:00 PM.


VT

2. Amey checks his records and finds no such training scheduled.

3. You recall that the training coordinator informed yesterday about this session.

4. Amey suspects it could have been Joey (the second admin) who got the update.

5. Joey confirms: You are indeed scheduled for training at 3:00 PM.

Problem Identified:

濿 "A clear case of inconsistent system!!!"

• The update was shared only with Joey, but you were checking with Amey.

• Both records were not synchronized, leading to data inconsistency.

Resolution Plan by Training Coordinator:

To ensure data consistency, the coordinator proposes:

Prof. Deepika G , Dept. of CSE, SVIT Page 24


Studied smart, not hard — thanks to [Link]
1. Dual Update Rule:
o If either Amey or Joey gets an update, they must update their own file and inform
the other to update theirs too.

2. Email Backup Rule:

o If the other admin is unavailable, the first one must send all updates via email.

3. Update on Resumption:

o When the unavailable admin returns to duty, he must update his file using the
email updates.

.IN
Key Learnings:

• The inconsistency happened due to a lack of synchronization between two systems


(admins).
C
N
• The solution ensures that availability may be slightly compromised (delays), but
consistency is preserved.
SY

• This directly reflects the CAP Theorem:

o You cannot guarantee Consistency, Availability, and Partition Tolerance


simultaneously.
U

o This solution chooses to sacrifice some availability to preserve consistency.


VT

忿 Summary: CAP Theorem Trade-Offs

You can only choose two out of three:

1. Consistency (C):

o All users always get the latest, updated information.

2. Availability (A):

o The system always responds (e.g., admins always provide schedule info).

3. Partition Tolerance (P):

o System functions even if parts (admins/servers) can’t communicate.

俿 When to Choose What?

• Choose Availability over Consistency when:

o Business can tolerate some delay in data sync.


Prof. Deepika G , Dept. of CSE, SVIT Page 25
Studied smart, not hard — thanks to [Link]
• Choose Consistency over Availability when:

o Business requires accurate, up-to-date reads/writes.

㿿 Examples of Database Systems Based on CAP:

Combination Description Example Databases

System is always responsive and


AP (Availability + Riak, Cassandra,
continues despite partitions, but may
Partition Tolerance) CouchDB, DynamoDB
have inconsistent data.

.IN
CP (Consistency + Strongly consistent but may become HBase, MongoDB, Redis,
Partition Tolerance) unavailable during network issues.
C MemcacheDB, BigTable

CA (Consistency + Consistent and always available, but can't Traditional RDBMS,


N
Availability) tolerate partitions. PostgreSQL, MySQL, etc.
SY
U
VT

Data Growth and Units (Table 2.2)

• Data Size Units (Binary scale):

o Bit = 0 or 1

o Byte = 8 bits

o Kilobyte (KB) = 1024 bytes

Prof. Deepika G , Dept. of CSE, SVIT Page 26


Studied smart, not hard — thanks to [Link]
o Megabyte (MB) = 1024² bytes

o Gigabyte (GB) = 1024³ bytes

o Terabyte (TB) = 1024⁴ bytes

o Petabyte (PB) = 1024⁵ bytes

o Exabyte (EB) = 1024⁶ bytes

o Zettabyte (ZB) = 1024⁷ bytes

.IN
o Yottabyte (YB) = 1024⁸ bytes

• Data Size Units (Decimal scale as per Figure 2.6):


C
o 1 KB = 1,000 bytes
N
o 1 MB = 1,000,000 bytes

o 1 GB = 1,000,000,000 bytes
SY

o 1 TB = 1,000,000,000,000 bytes

o 1 PB = 1,000,000,000,000,000 bytes
U

o 1 EB = 1,000,000,000,000,000,000 bytes
VT

o 1 ZB = 1,000,000,000,000,000,000,000 bytes

o 1 YB = 1,000,000,000,000,000,000,000,000 bytes

Note: Binary units are based on powers of 2 (1024), while decimal units are powers of 10 (1000).

Figure 3.15 – CAP Triangle

• Illustrates that you can pick only two of the three: C, A, or P.

• Each DB system aligns with one of the triangle’s edges:

o CA: Traditional RDBMS (e.g., PostgreSQL)

o CP: Systems like HBase, Redis

o AP: Systems like Cassandra, DynamoDB

Prof. Deepika G , Dept. of CSE, SVIT Page 27


Studied smart, not hard — thanks to [Link]
4.1 NoSQL (NOT ONLY SQL)

• The term NoSQL was first coined by Carlo Strozzi in 1998 for his lightweight, open-source
relational database that did not expose the standard SQL interface.

• In 2009, Johan Oskarsson reintroduced the term at an event focused on open-source


distributed networks.

• The hashtag #NoSQL was coined by Eric Evans and others to describe non-relational
databases.

.IN
Key Features of NoSQL Databases:

1. They are open source.

2. They are non-relational.


C
3. They are distributed.
N
4. They are schema-less.
SY

5. They are cluster-friendly.

6. They emerged from 21st-century web applications.


U

4.1.1 Where is it Used?


VT

• Widely used in big data and real-time web applications.

• Used to stock large volumes of data for analysis.


• Ideal for storing social media data and other data types that are difficult to store and
analyze in traditional RDBMS.

4.1.2 What is it?

• NoSQL stands for Not Only SQL.

• These databases are:

o Non-relational

o Open source

o Distributed

• Popular due to:

o Ability to scale out or scale horizontally.

o Aptitude for handling structured, semi-structured, and unstructured data.


Prof. Deepika G , Dept. of CSE, SVIT Page 28
Studied smart, not hard — thanks to [Link]
Additional Feature of NoSQL Databases:

• Non-relational: They do not follow a relational data model.

• Types of NoSQL databases include:

o Key–value pairs

o Document-oriented

o Column-oriented

o Graph-based databases.

Characteristics of NoSQL Databases:

.IN
1. Are non-relational
C
N
o NoSQL databases do not adhere to the relational data model.

They can be key-value pairs, document-oriented, column-oriented, or graph-


SY

o
based databases.

Figure 4.1: Where to use NoSQL?


U

• Log analysis

• Social networking feeds


VT

• Time-based data (not easily analyzed in traditional RDBMS)

Figure 4.2: What is NoSQL?

• Non-relational data storage systems


• No fixed table schema
• No joins

Prof. Deepika G , Dept. of CSE, SVIT Page 29


Studied smart, not hard — thanks to [Link]
• No multi-document transactions
• Relaxes one or more ACID properties

.IN
C
N
Additional Features:
SY

2. Are distributed:

o Data is distributed across several nodes in a cluster, which is often made up of low-
U

cost commodity hardware.


VT

3. Offer no support for ACID properties (Atomicity, Consistency, Isolation, Durability):

o They generally do not support traditional ACID transaction properties.

o Instead, they adhere to Brewer's CAP theorem (Consistency, Availability, and


Partition tolerance), often compromising on consistency in favor of availability
and partition tolerance.

4. Provide no fixed schema:

o NoSQL databases allow flexible schema design.

o They do not require data to strictly follow any schema, making them suitable for
evolving or semi-structured data.

4.1.3 Types of NoSQL Databases

NoSQL databases are non-relational and broadly classified into two types:

1. Key-Value (the big hash table)

o Maintains a large hash table of keys and values.

Prof. Deepika G , Dept. of CSE, SVIT Page 30


Studied smart, not hard — thanks to [Link]
o Examples: Dynamo, Redis, Riak, Amazon S3, Scalaris.

o Sample Key-Value Pair:

Key Value

First Name Simmonds

Last Name David

.IN
2. Schema-less (includes Document, Column, and Graph-based)

o These databases do not have a fixed schema.


C
o Examples:
N
Document: MongoDB, Apache CouchDB, Couchbase, MarkLogic.

Store data as collections of documents.


SY

Sample Document:

{
U

"Book Name": "Fundamentals of Business Analytics",


VT

"Publisher": "Wiley India",

"Year of Publication": "2011"

Column: Cassandra, HBase.

Each storage block contains data from only one column.

Graph-based: Neo4j.

Prof. Deepika G , Dept. of CSE, SVIT Page 31


Studied smart, not hard — thanks to [Link]
Graph Databases (also called network databases)
• Store data in nodes.
• Examples: Neo4j, HyperGraphDB.
• Sample graph structure example:
o Nodes:
▪ ID: 1001, Name: John, Age: 28
▪ ID: 1002, Name: Joe, Age: 32
▪ ID: 1003, Name: Group, Age: AAA
o Relationships:
▪ John knows Joe since 2002

.IN
▪ John is a member of Group since 2003
▪ Joe is a member of Group since 2002

4.1.4 Why NoSQL?


C
N
1. Scale-out architecture: Instead of monolithic relational databases.

2. Can house large volumes of structured, semi-structured, and unstructured data.


SY

3. Dynamic schema: Allows insertion of data without a predefined schema. Supports


faster development and easier code integration.
U

4. Auto-sharding: Automatically spreads data across multiple servers, balancing load and
providing quick recovery if a server goes down.
VT

5. Replication: Supports replication for high availability, fault tolerance, and disaster
recovery.

4.1.5 Advantages of NoSQL

• Can easily scale up and down: Supports rapid, elastic scaling including scaling to the
cloud.

Table 4.1: Popular Schema-less Databases

Key-Value Data Column-Oriented Document Data Graph Data


Store Data Store Store Store

Riak Cassandra MongoDB Infinite Graph

Redis HBase CouchDB Neo4j

Membase Hyper Table Raven DB Allegro Graph

Prof. Deepika G , Dept. of CSE, SVIT Page 32


Studied smart, not hard — thanks to [Link]
.IN
More Detailed Advantages of NoSQL
C
N
1. Cluster scale:
o Supports distribution across 100+ nodes (often in multiple data centers).
SY

2. Performance scale:
o Handles over 100,000+ database reads/writes per second.
3. Data scale:
o Can store over 1 billion+ documents.
U

2. Doesn’t require a pre-defined schema


VT

• NoSQL (e.g., MongoDB) allows records with different sets of key-value pairs.
• Example (from MongoDB):

{
"_id": 101,
"BookName": "Fundamentals of Business Analytics",
"AuthorName": "Seema Acharya",
"Publisher": "Wiley India"
},
{
"_id": 102,
"BookName": "Big Data and Analytics"
}

Prof. Deepika G , Dept. of CSE, SVIT Page 33


Studied smart, not hard — thanks to [Link]
3. Cheap, easy to implement
• Allows benefits of scale, fault tolerance, high availability at low cost.

4. Relaxes the data consistency requirement


• Adheres to the CAP theorem (Compromises consistency, but ensures availability and
partition tolerance).
• Supports eventual consistency.

5. Data can be replicated and partitioned

.IN
(a) Sharding
• Distributes data automatically across multiple servers.
• Handles server addition/removal without application downtime.
• Balances data and query load.
C
(b) Replication
N
• Stores multiple copies of data across clusters/data centers.
• Ensures high availability and fault tolerance.
SY

4.1.6 What We Miss With NoSQL?


U

• NoSQL solves scalability and schema flexibility issues.


VT

• But some RDBMS features are still superior – details in next figure (Figure 4.5).

1. Joins – NoSQL databases generally do not support JOIN operations like SQL databases.

2. Group By – Advanced aggregations like GROUP BY can be complex or unsupported.

3. ACID properties – No built-in support for Atomicity, Consistency, Isolation,


Durability.

Prof. Deepika G , Dept. of CSE, SVIT Page 34


Studied smart, not hard — thanks to [Link]
4. SQL – No standard query language (although MongoDB uses its own query syntax, and
Cassandra uses CQL).

5. Easy integration with other SQL-based applications – Since it lacks standard SQL, it
is harder to integrate with tools designed for SQL-based databases.

However, MongoDB and Cassandra mitigate this to some extent with:

• MongoDB Query Language

• CQL (Cassandra Query Language)

.IN
4.1.7 Use of NoSQL in Industry

NoSQL is used across various industries, particularly for big data and real-time applications.
C
Breakdown of how different NoSQL models are used:
N
1. Key–Value Pairs
o Use: Shopping carts, user data analysis
SY

o Examples: Amazon, LinkedIn


2. Column-Oriented
o Use: Analyze large user actions, sensor feeds
o Examples: Facebook, Twitter, eBay, Netflix
U

3. Document Based
o Use: Real-time analytics, logging, document management
VT

o Examples: MongoDB, CouchDB use cases


4. Graph-Based
o Use: Network modeling, recommendations, upsell and cross-sell analytics
o Examples: Walmart, social networks

Prof. Deepika G , Dept. of CSE, SVIT Page 35


Studied smart, not hard — thanks to [Link]
4.1.8 NoSQL Vendors

A few popular NoSQL vendors and their products:

Company Product Most Widely Used by

Amazon DynamoDB LinkedIn, Mozilla

Facebook Cassandra Netflix, Twitter, eBay

Google BigTable Adobe Photoshop

4.1.9 SQL vs NoSQL


.IN
C
Table 4.3 summarizes the key differences:
N
SQL NoSQL
SY

Relational database Non-relational, distributed database

Relational model Model-less approach


U

Pre-defined schema Dynamic schema for unstructured data


VT

Table-based databases Document, graph, wide-column, or key-value store

Vertically scalable Horizontally scalable (scale-out using clusters)

Uses SQL Uses UnQL (Unstructured Query Language)

Not ideal for large datasets Preferred for large datasets

Not a good fit for hierarchical data Ideal for hierarchical/JSON-like data

Emphasis on ACID properties Follows CAP theorem (sacrifices ACID)

Excellent vendor support Heavily community supported

Supports complex queries Weak at complex querying

Prof. Deepika G , Dept. of CSE, SVIT Page 36


Studied smart, not hard — thanks to [Link]
SQL NoSQL

Can be configured for strong


Some support eventual consistency (e.g., Cassandra)
consistency

Examples: Oracle, MySQL, Examples: MongoDB, Cassandra, HBase, Redis, Neo4j,


PostgreSQL, etc. CouchDB

.IN
4.1.10 NewSQL

• NewSQL is a modern RDBMS that combines features of both SQL and NoSQL.

It offers the scalability of NoSQL systems used for Online Transaction Processing

C
(OLTP).
N
• At the same time, it maintains the ACID guarantees of traditional databases.
SY

• NewSQL supports the relational data model and uses SQL as the primary interface.

[Link] Characteristics of NewSQL


U

• Based on the shared-nothing architecture.


• Uses a SQL interface for application interaction.
VT

4.1.11 Comparison of SQL, NoSQL, and NewSQL

Feature SQL NoSQL NewSQL

Adherence to ACID properties Yes No Yes

OLTP/OLAP Yes No Yes


Schema rigidity Yes No Maybe

Adherence to
Adherence to data model No Yes
relational model

Data Format Flexibility No Yes Maybe


Scale out
Scale up (Vertical
Scalability (Horizontal Scale out
Scaling)
Scaling)
Distributed Computing Yes Yes Yes

Community Support Huge Growing Slowly growing

Prof. Deepika G , Dept. of CSE, SVIT Page 37


Studied smart, not hard — thanks to [Link]
4.2 HADOOP

• Hadoop is an open-source project by the Apache Foundation.


• It is a framework written in Java, originally developed by Doug Cutting in 2005.
• Named after Doug Cutting’s son’s toy elephant.
• Initially developed to support distribution for Nutch (a text search engine).
• Inspired by Google MapReduce and Google File System.
• Now a core part of computing infrastructure for companies like Yahoo, Facebook,
LinkedIn, Twitter, etc.

.IN
4.2.1 Features of Hadoop

• Optimized to handle massive amounts of structured, semi-structured, and unstructured


data using inexpensive, commodity hardware.
C
• Based on a shared-nothing architecture.
N
• Data replication across multiple computers ensures fault tolerance and availability.
SY

• Designed for high throughput (not low latency). It's batch-oriented.


• Supports OLTP and OLAP, but not a replacement for traditional RDBMS.
• Not ideal when work cannot be parallelized or when data dependencies exist.
U

• Not suitable for small files—performs best with huge datasets.


VT

4.2.2 Key Advantages of Hadoop

Stores data in its native format:


• HDFS allows storing data without enforcing structure at input.
• Structure is applied only during processing.
Scalable:
• Can store and distribute data across thousands of low-cost servers.
• Proven scalability (e.g., used by Facebook & Yahoo).
Cost-effective:
• Reduced cost per terabyte due to scale-out architecture.
Resilient to failure:
• Ensures data replication across nodes for fault tolerance.
• Automatically recovers from node failures.

Prof. Deepika G , Dept. of CSE, SVIT Page 38


Studied smart, not hard — thanks to [Link]
Flexible:
• Handles all types of data (structured, semi-structured, unstructured).
• Supports various applications: log analysis, data mining, recommendation systems,
market campaigns, etc.
Fast:
• High-speed data processing due to “move code to data” paradigm.

.IN
C
N
SY
U
VT

4.2.3 Versions of Hadoop

There are two versions of Hadoop available:


1. Hadoop 1.0
2. Hadoop 2.0

Prof. Deepika G , Dept. of CSE, SVIT Page 39


Studied smart, not hard — thanks to [Link]
.IN
• [Link] Hadoop 1.0 C
• It has two main parts:
N
1. Data storage framework:
▪ A general-purpose file system called Hadoop Distributed File System
SY

(HDFS).
▪ HDFS is schema-less.
▪ It simply stores data files.
U

▪ These data files can be in just about any format.


VT

2. Data processing framework:


Based on a simple functional programming model called MapReduce (popularized by Google).

o Uses two key functions:


▪ Map: Takes in a set of key–value pairs and generates intermediate data
(another list of key–value pairs).
▪ Reduce: Acts on the intermediate data to produce the output data.
o Map and Reduce work in isolation from each other.
o Processing is highly distributed, fault-tolerant, and scalable.

Limitations of Hadoop 1.0

1. Requirement for MapReduce expertise:

o Developers needed to be proficient in MapReduce and programming languages like


Java.

2. Batch processing only:

Prof. Deepika G , Dept. of CSE, SVIT Page 40


Studied smart, not hard — thanks to [Link]
o Hadoop1.0 supported only batch processing.
o Suitable for:
▪ Log analysis
▪ Large-scale data mining
o Not suitable for:
▪ Other kinds of projects requiring real-time or interactive processing.
3. Tight coupling with MapReduce:
o Hadoop 1.0 was tightly coupled with MapReduce.
o Vendors had two poor options:
▪ Rewrite their tools to work with MapReduce.

.IN
▪ Extract data from HDFS and process it outside Hadoop.
o Both options led to inefficiencies due to data movement in and out of the Hadoop
cluster.
o
C
[Link] Hadoop 2.0
N
• HDFS remains the data storage framework in Hadoop 2.0.
SY

• Introduced a new resource management framework called:


o YARN (Yet Another Resource Negotiator):
▪ Supports dividing applications into parallel tasks.
▪ Enhances flexibility, scalability, and efficiency.
U

▪ Replaces the old JobTracker with ApplicationMaster.


▪ Replaces TaskTracker with NodeManager.
VT

▪ Allows any application (not just MapReduce) to run on Hadoop.

• Key Benefits:
o MapReduce expertise is no longer required.
o Supports both:
▪ Batch processing
▪ Real-time processing
o MapReduce is no longer the only option:
▪ Alternative data processing tools are now supported.
▪ Native features like data standardization and master data management can
be performed in HDFS.

Prof. Deepika G , Dept. of CSE, SVIT Page 41


Studied smart, not hard — thanks to [Link]
.IN
HDFS (Hadoop Distributed File System)

• It is the distributed storage unit of Hadoop.


C
• Provides streaming access to file system data.
N
• Supports file permissions and authentication.
SY

• Based on GFS (Google File System).

• Scales a single cluster node to hundreds or thousands of nodes.


U

• Handles large datasets on commodity hardware.


VT

• HDFS is highly fault-tolerant:

o Stores files across multiple machines.

o Files are stored in a redundant fashion to allow data recovery in case of failure.

(Example Scenario)

• An e-commerce website stores millions of customers' data in a distributed manner.

• Data has been collected over 4–5 years.

• Batch analytics is run on archived data to analyze:

o Customer behavior

o Buying patterns

o Preferences

Prof. Deepika G , Dept. of CSE, SVIT Page 42


Studied smart, not hard — thanks to [Link]
o Requirements
• Helps identify:
o Which products are purchased
o In which months
o By which types of customers

HBase

• Stores data in HDFS.


• It is the first non-batch component of the Hadoop ecosystem.

.IN
• Works as a database on top of HDFS.
• Offers quick random access to stored data.
• Has very low latency compared to HDFS.
• It is a:
C
o NoSQL database
N
o Non-relational
o Column-oriented database
SY

• Data Structure:
o A table can have thousands of columns.
o A row can have several column families.
U

o Each column family can contain multiple columns.


o Each column can have several key-value pairs.
VT

• Based on: Google BigTable


• Used by: Facebook, Twitter, Yahoo, etc.

(Example Scenario)
• The same e-commerce website also stores millions of product data.
• To search among millions of products and get real-time results, optimization is needed.

• HBase supports real-time analytics.

• Due to high data velocity, they chose HBase over HDFS, as HDFS does not support
real-time writes.

• Impact:

o Query time reduced from 3 days to 3 minutes.

Prof. Deepika G , Dept. of CSE, SVIT Page 43


Studied smart, not hard — thanks to [Link]
Difference Between HBase and Hadoop/HDFS

HDFS (Hadoop Distributed File


Aspect HBase (Hadoop Database)
System)

Type File System NoSQL Database

Analogy Like NTFS Like MySQL

Real-time random read and


Write/Read Model WORM (Write Once Read Many)

.IN
write

Underlying System Based on Google File System (GFS)


C Based on Google BigTable

Random access to small


Data Access Pattern Full table scan or partition scan
ranges
N
Performance with
SY

Very good 4–5 times slower than HDFS


Hive

Via Java APIs, REST, Avro,


Access Methods Via MapReduce jobs only
Thrift
U

Storage Type Static and rigid Dynamic and flexible


VT

Latency High latency Low latency

Best Suited For Batch processing and analytics Real-time analytics

Hadoop Ecosystem Components for Data Processing

1. MapReduce

• A programming model for processing large datasets in parallel.

• Developed by Google in 2004.

• Works in two phases:

o Map: Converts input data into key-value pairs.

o Reduce: Combines and processes these pairs to generate output.

Prof. Deepika G , Dept. of CSE, SVIT Page 44


Studied smart, not hard — thanks to [Link]
• Data flows from and back to HDFS.

2. Spark

• A programming and computing model developed at UC Berkeley in 2009.

• Open-source and written in Scala.

• Performs in-memory processing (much faster than disk-based).

• Can use HDFS as a data source but does not require MapReduce.

.IN
• Can run with or without Hadoop.

Spark Libraries: C
• Spark SQL – Supports querying using SQL.
N
• Spark Streaming – For real-time data analysis.

• MLlib – For machine learning tasks.


SY

• GraphX – For graph computations.


U

Key Notes
VT

• Hadoop is used for batch processing of unstructured data.

• Spark is widely used for fast, real-time, in-memory processing.

Hadoop Ecosystem Components for Data Analysis

1. Pig

• A high-level scripting language used with Hadoop.

• Acts as an alternative to MapReduce.

• Developed by Yahoo.

Pig has two main parts:

• Pig Latin:

oA scripting language that gets translated into MapReduce jobs.

o Used for ETL (Extract, Transform, Load) and data analysis.

Prof. Deepika G , Dept. of CSE, SVIT Page 45


Studied smart, not hard — thanks to [Link]
o Supports operations like grouping, filtering, sorting, joining, etc.

o Loads data from HDFS, processes it, and sends it back or displays it.

• Pig runtime:

o The execution environment where Pig scripts are run.

2. Hive

• A data warehouse software on top of Hadoop.

.IN
• Used for summarization, querying, and analysis.

• Uses HQL (Hive Query Language) or HiveQL, similar to SQL.


C
• Converts SQL-like queries into MapReduce jobs for execution on Hadoop.
N
Difference Between Hive and RDBMS
SY

Aspect Hive RDBMS (MySQL, SQL Server, etc.)

Schema Enforced on read (schema-on-


U

Enforced on write (schema-on-write)


Enforcement read)
VT

Fast, as schema is not enforced Slower due to schema enforcement


Data Load
during insert during load

Query Better read performance after Good performance after schema


Performance faster load validation

Data Access
Write once, read many times Read and write many times
Model

Batch-oriented, not ideal for Designed for OLTP and day-to-day


Processing Type
OLTP transactions

Analytical workloads (OLAP-like


Best Suited For Transactional workloads (OLTP)
use cases)

Queries are converted into Queries are executed directly in the


Query Execution
MapReduce jobs database engine

Prof. Deepika G , Dept. of CSE, SVIT Page 46


Studied smart, not hard — thanks to [Link]
4.2.5 Hadoop Distributions

Hadoop is an open-source Apache project.

Core components of Hadoop:

1. Hadoop Common

2. HDFS (Hadoop Distributed File System)

3. YARN (Yet Another Resource Negotiator)

.IN
4. MapReduce

Several companies have built custom distributions of Hadoop for easier use:
C
• IBM
N
• Amazon Web Services

• Microsoft
SY

• Teradata

• Hortonworks
U

• Cloudera
VT

Despite different strategies, all distributions aim to:

• Distribute data and workloads across many servers.

• Make big data manageable and scalable.

4.2.5 Hadoop Distributions

• Hadoop is an open-source project from Apache, freely available for download.

• Core components of Hadoop include:

1. Hadoop Common

2. Hadoop Distributed File System (HDFS)

3. Hadoop YARN (Yet Another Resource Negotiator)

4. Hadoop MapReduce

Prof. Deepika G , Dept. of CSE, SVIT Page 47


Studied smart, not hard — thanks to [Link]
• Various companies have created their own distributions (packaged versions) of Hadoop
to make it more user-friendly and consumable. These include: IBM, Amazon, Web
Services, Microsoft, Teradata, Hortonworks, Cloudera.

• Though their strategies differ slightly, the main goal is to distribute data and workloads
across many servers for scalable and manageable big data processing.

.IN
C
N
SY
U
VT

4.2.8 Hadoop versus SQL

Hadoop SQL

Scale out Scale up

Key-Value pairs Relational table

Functional Programming Declarative Queries

Offline batch processing Online transaction processing

Prof. Deepika G , Dept. of CSE, SVIT Page 48


Studied smart, not hard — thanks to [Link]
• Scale out (Hadoop) means adding more machines to handle increased workload, while
scale up (SQL) means increasing the power of a single machine.
• Hadoop uses a key-value pair data model, whereas SQL uses relational tables.
• Hadoop programming is functional in nature, while SQL uses declarative queries.
• Hadoop is optimized for offline batch processing of large datasets, while SQL is designed
for online transaction processing (OLTP) with quick response times.

4.2.7 Integrated Hadoop Systems Offered by Leading Market Vendors

These vendors package and offer Hadoop-based big data platforms that are ready for

.IN
enterprise use.

The major vendors mentioned are:

• EMC Greenplum
C
• Oracle Big Data Appliance
N
• Microsoft Big Data Solution
• IBM InfoSphere
SY

• HP Big Data Solutions



These integrated systems typically combine hardware, software, and support services optimized
for Hadoop workloads, helping organizations deploy big data solutions more easily and
U

efficiently.
VT

4.2.8 Cloud-Based Hadoop Solutions

Amazon Web Services (AWS): Provides a comprehensive portfolio of cloud services aimed at
managing big data efficiently. AWS emphasizes reducing costs, scaling on demand, and
accelerating innovation speed.

Google Cloud Storage connector for Hadoop: Allows direct execution of MapReduce jobs on
data stored in Google Cloud Storage without copying it locally. This simplifies Hadoop
deployment, reduces costs, and offers performance comparable to Hadoop's HDFS. It also
enhances reliability by removing the single point of failure (name node).

Prof. Deepika G , Dept. of CSE, SVIT Page 49


Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT

Prof. Deepika G , Dept. of CSE, SVIT Page 50


Studied smart, not hard — thanks to [Link]

You might also like