Department of Computer Science and Engineering
Bigdata and Analytics (BCS714D)
Module-01
Data: It refers to raw, unorganized facts or figures.
• It can be numbers, words, measurements, or observations.
Example: "John, 25, Male" — this is data.
.IN
•
Information is processed or organized data that is meaningful and useful.
• It is data given context.
C
• Example: "John is a 25-year-old male customer." — This is information derived from
N
data.
SY
What is Big Data?
• Big Data refers to extremely large and complex datasets that traditional tools can't
process efficiently.
U
• It is characterized by 5 V’s:
VT
o Volume: Huge amount of data.
o Velocity: Speed at which data is generated.
o Variety: Different formats (text, video, images, etc.).
o Veracity: Uncertainty of data accuracy.
o Value: Extracting useful insights.
Big Data Analytics: Big Data Analytics is the process of examining big data to uncover
patterns, trends, and insights.
• It helps businesses make informed decisions.
• Uses tools like Hadoop, Spark, Hive, Pig, etc.
• Example: Analyzing customer purchase history to recommend products.
Prof. Deepika G , Dept. of CSE, SVIT Page 1
Studied smart, not hard — thanks to [Link]
� 2.1 Characteristics of Data
Three key characteristics of data:
1. Composition
o Structure of data: sources, types, granularity.
o Static vs. real-time data.
2. Condition
.IN
o State of the data (clean, raw, needs enhancement).
o Questions: Can it be used as-is? Does it need processing?
C
3. Context
N
o Background of the data: where, when, why was it generated?
SY
U
VT
2.2 Evolution of Big Data
Overview:
• Before 1970s: Data was primitive and structured; stored using mainframes.
• 1980s–1990s: Rise of relational databases; data became relational and used in data-
intensive applications.
Prof. Deepika G , Dept. of CSE, SVIT Page 2
Studied smart, not hard — thanks to [Link]
• 2000s & beyond: Explosion of complex, unstructured data due to Web, IoT, and
multimedia technologies.
Table 2.1 – Evolution of Big Data
Data Generation &
Time Period Data Utilization Data Driven Output
Storage
1970s & Primitive and Mainframes for basic data
.IN
—
before Structured storage
1980s & Complex and Relational databases; data-
—
1990s Relational
C
intensive apps
N
2000s and Complex and Structured, unstructured,
—
beyond Unstructured multimedia data
SY
Key Contributors to Modern Big Data
U
• World Wide Web (WWW) – Opened floodgates for global digital content.
VT
• Internet of Things (IoT) – Devices constantly generate vast, diverse data streams.
2.3 Definition of Big Data
Common Understandings of Big Data
Several responses have been widely accepted to define Big Data:
1. Beyond Infrastructure:
Big Data refers to anything that exceeds traditional storage, processing, and
analysis capabilities of existing human and technical infrastructure.
2. Relative to Time:
What is considered "big" today may become normal or standard tomorrow, indicating
that Big Data is a moving threshold.
3. Large Volume:
It refers to enormous data sizes – Terabytes, Petabytes, Zettabytes, and even
Yottabytes.
Prof. Deepika G , Dept. of CSE, SVIT Page 3
Studied smart, not hard — thanks to [Link]
4. The 3 Vs – A standard definition framework.
Standard 3Vs Definition (Gartner’s Perspective)
According to Gartner IT Glossary and Doug Laney (2001), Big Data is defined by:
Big Data is high-volume, high-velocity, and high-variety information assets that demand
cost-effective, innovative forms of information processing for enhanced insight and
decision making.
• Volume: Huge amounts of data.
.IN
• Velocity: Speed of data generation and processing.
• Variety: Different types and formats (structured, unstructured, semi-structured).
C
N
Key Conceptual Flow
SY
Big Data ➝ Information ➝ Actionable Intelligence ➝ Better Decisions ➝ Enhanced
Business Value
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 4
Studied smart, not hard — thanks to [Link]
1. Volume-Based Definition
• Focuses on the massive size of data.
• Big Data is often defined by its scale, measured in terabytes, petabytes, zettabytes, or
even yottabytes.
2. Infrastructure-Based Definition
• Refers to data that exceeds the capacity of current tools or infrastructure.
.IN
• Suggests Big Data begins where traditional systems struggle to store, process, or
analyze data efficiently. C
N
3. Relativity Definition
SY
• Emphasizes the relative and evolving nature of Big Data.
• What we call “big” today may become standard in the near future due to technological
advances.
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 5
Studied smart, not hard — thanks to [Link]
Definition of Big Data (Gartner’s Perspective)
Three-Part Definition (3Vs Model):
"Big Data is high-volume, high-velocity, and high-variety information assets that
demand cost-effective, innovative forms of information processing for enhanced insight
and decision making."
Part I: High-Volume, High-Velocity, High-Variety
• Refers to large and fast-growing data.
.IN
• Includes:
o Structured, semi-structured, and unstructured data.
C
• Requires:
N
o Fast storage, processing, and analysis capabilities.
SY
Part II: Cost-Effective & Innovative Processing
• Focuses on:
U
o Using new technologies and methods to: Ingest, Store, Process, Persist,
Integrate, Visualize.
VT
• Goal:
o Efficient handling of large, fast, and diverse data sets.
Part III: Enhanced Insight & Decision-Making
• Involves:
o Gaining deep, rich, and actionable insights.
o Using those insights for better and faster decisions.
• Leads to:
o Improved business value and competitive advantage.
Prof. Deepika G , Dept. of CSE, SVIT Page 6
Studied smart, not hard — thanks to [Link]
Flow Summary:
Data → Information → Actionable Intelligence → Better Decisions → Enhanced Business
Value
Challenges with Big Data
1. Exponential Data Growth
o Data volume is increasing rapidly, mostly generated in the last 2-3 years.
.IN
o Raises key questions:
▪ Will all this data be useful for analysis?
Should we work with all data or a subset?
▪
C
▪ How to separate valuable knowledge from noise?
N
2. Cloud Computing and Virtualization
SY
o Cloud is essential for cost-efficient, elastic, and scalable infrastructure.
o However, deciding whether to host big data solutions on-premises or in the cloud
U
is complex.
3. Data Retention Period
VT
o Determining how long to keep big data is tricky.
o Some data is valuable for long-term decisions.
o Other data becomes irrelevant or obsolete quickly (sometimes within hours).
Prof. Deepika G , Dept. of CSE, SVIT Page 7
Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
2.5 What is Big Data?
U
• Big data is defined by three main characteristics:
VT
o Volume (amount of data)
o Velocity (speed of data generation and processing)
o Variety (different types and formats of data)
2.5.1 Volume
• Data size has grown exponentially, evolving through units:
o Bits → Bytes → Kilobytes → Megabytes → Gigabytes → Terabytes → Petabytes →
Exabytes → Zettabytes → Yottabytes
• The sources of big data are diverse and can include:
o XLS, DOC, PDF (unstructured data)
o Videos (e.g., YouTube)
o Chat conversations (e.g., Internet Messenger)
o Customer feedback forms on retail websites
Prof. Deepika G , Dept. of CSE, SVIT Page 8
Studied smart, not hard — thanks to [Link]
[Link] Where does Data get generated?
Data Characteristics Illustrated:
• Data Volume (X-axis, moving outward):
o Data size grows from small (MB, GB) to very large (TB, PB).
o The more data you have, the further out on the volume axis it is.
• Data Velocity (Y-axis, moving upward):
.IN
o Speed of data generation and processing:
▪ From Batch processing (slowest) to Periodic, Near real-time, and Real-time
(fastest).
C
N
• Data Variety (Z-axis, moving diagonally):
SY
o Types of data include:
▪ Table, Database (structured data)
Photo, Social, Web, Video, Audio, Mobile (mostly unstructured or semi-
U
▪
structured data)
VT
Data Growth and Units (Table 2.2)
• Data Size Units (Binary scale):
o Bit = 0 or 1
o Byte = 8 bits
o Kilobyte (KB) = 1024 bytes
o Megabyte (MB) = 1024² bytes
o Gigabyte (GB) = 1024³ bytes
o Terabyte (TB) = 1024⁴ bytes
o Petabyte (PB) = 1024⁵ bytes
o Exabyte (EB) = 1024⁶ bytes
Prof. Deepika G , Dept. of CSE, SVIT Page 9
Studied smart, not hard — thanks to [Link]
o Zettabyte (ZB) = 1024⁷ bytes
o Yottabyte (YB) = 1024⁸ bytes
• Data Size Units (Decimal scale as per Figure 2.6):
o 1 KB = 1,000 bytes
o 1 MB = 1,000,000 bytes
o 1 GB = 1,000,000,000 bytes
.IN
o 1 TB = 1,000,000,000,000 bytes
o 1 PB = 1,000,000,000,000,000 bytes C
o 1 EB = 1,000,000,000,000,000,000 bytes
N
o 1 ZB = 1,000,000,000,000,000,000,000 bytes
o 1 YB = 1,000,000,000,000,000,000,000,000 bytes
SY
Note: Binary units are based on powers of 2 (1024), while decimal units are powers of 10 (1000).
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 10
Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 11
Studied smart, not hard — thanks to [Link]
Sources of Big Data (Figure 2.7)
.IN
C
• Internal Data Sources:
N
o Data storage: File systems, SQL databases (Oracle, MySQL, PostgreSQL,
SY
MongoDB, Cassandra, etc.)
o Archives: Scanned documents, paper archives, customer correspondence, health
records, admission records, assessment records, etc.
U
• External Data Sources:
VT
o Public Web: Wikipedia, weather, regulatory data, compliance, census, etc.
• Both Internal + External:
o Sensor data: Car sensors, smart meters, HVAC systems, refrigerators, etc.
o Machine log data: Event logs, app logs, business process logs, audit logs,
clickstream data.
o Social media: Twitter, Facebook, LinkedIn, YouTube, Instagram.
o Business apps: ERP, CRM, HR systems, Google Docs.
o Media: Audio, video, images, podcasts.
o Docs: CSV files, Word documents, PDFs, XLS, PPT.
Prof. Deepika G , Dept. of CSE, SVIT Page 12
Studied smart, not hard — thanks to [Link]
.IN
C
N
External Data Sources
SY
• Data that resides outside an organization’s firewall.
• Examples include:
U
o Public Web: Wikipedia, weather data, regulatory info, compliance, census data,
etc.
VT
Both (Internal + External Data Sources)
• Data that can come from both inside and outside the organization.
• Examples include:
➢ Sensor data: Car sensors, smart electric meters, office building sensors, air
conditioning units, refrigerators, etc.
➢ Machine log data: Event logs, application logs, business process logs, audit logs,
clickstream data, etc.
➢ Social media: Twitter, blogs, Facebook, LinkedIn, YouTube, Instagram, etc.
➢ Business apps: ERP (Enterprise Resource Planning), CRM (Customer Relationship
Management), HR systems, Google Docs, etc.
➢ Media: Audio, video, images, podcasts, etc.
Prof. Deepika G , Dept. of CSE, SVIT Page 13
Studied smart, not hard — thanks to [Link]
➢ Docs: CSV (Comma Separated Values), Word documents, PDF, XLS, PPT, and
similar files.
2.5.2 Velocity
• Velocity refers to the speed at which data is processed.
• There's been a shift from batch processing (e.g., traditional payroll applications) to real-
time processing.
• The progression in data processing speed is:
.IN
Batch → Periodic → Near real-time → Real-time processing
2.5.3 Variety
•
C
Variety deals with the different types and formats of data.
• It classifies data into three categories:
N
1. Structured data: Comes from traditional systems like transaction processing
SY
systems and RDBMS (Relational Database Management Systems).
2. Semi-structured data: Examples include HTML, XML (Hyper Text Markup
Language and eXtensible Markup Language).
U
3. Unstructured data: Includes unstructured text documents, audio, video, emails,
VT
photos, PDFs, social media content, etc.
3.2 WHAT IS BIG DATA ANALYTICS?
Big Data Analytics is…
1. Technology-enabled analytics:
Uses various data analytics and visualization tools from vendors like IBM, Tableau, SAS, R Analytics, etc.,
to process and analyze big data.
2. Gaining meaningful insights:
Helps businesses gain deeper, richer insights to better understand customers (demographics, preferences) for
cross-selling, up-selling, and better vendor/supplier management.
Example: Online retailers recommend products based on stored purchase and preference data.
3. Competitive edge:
Enables quicker and better decision-making, giving an advantage over competitors.
Prof. Deepika G , Dept. of CSE, SVIT Page 14
Studied smart, not hard — thanks to [Link]
4. Collaboration:
Involves close cooperation between IT teams, business users, and data scientists.
5. Handling large datasets:
Deals with datasets whose volume and variety surpass current storage and processing capabilities of
enterprises.
6. Moving code to data:
Instead of moving data around, programs (small in size) are moved close to where data resides (especially
large datasets in terabytes or petabytes), improving efficiency. This will become more important as data
.IN
grows to exabytes or zettabytes.
3.5 CLASSIFICATION OF ANALYTICS:
C
1. Classification by stages of analytics:
N
o Basic
SY
o Operationalized
o Advanced
U
o Monetized
VT
2. Classification by versions of analytics:
o Analytics 1.0
o Analytics 2.0
o Analytics 3.0
Figure 3.5 illustrates what big data entails, showing a cycle:
• More data produced →
• More data stored →
• More data analyzed →
• Better predictions →
• Steady growth of analysis → (loops back to more data produced)
This cycle highlights how increasing data leads to improved insights and predictions, fueling
further data generation and analysis.
Prof. Deepika G , Dept. of CSE, SVIT Page 15
Studied smart, not hard — thanks to [Link]
.IN
C
N
3.5.1 First School of Thought
SY
1. Basic analytics: This primarily is slicing and dicing of data to help with basic business
insights. This is about reporting on historical data, basic visualization, etc.
U
2. Operationalized analytics: It is operationalized analytics if it gets woven into the
enterprise’s business processes.
VT
3. Advanced analytics: This largely is about forecasting for the future by way of predictive
and prescriptive modeling.
4. Monetized analytics: This is analytics in use to derive direct business revenue.
3.5.2 Second School of Thought
Table 3.1 Analytics 1.0, 2.0, and 3.0
Analytics 1.0 Analytics 2.0 Analytics 3.0
Era: Mid 1950s to 2009 2005 to 2012 2012 to present
Descriptive + predictive + prescriptive
statistics (use data from the past to
Descriptive statistics (report Descriptive statistics + predictive
make prophecies for the future and
Stats used: on events, occurrences, etc. statistics (use data from the past to
at the same time make
of the past) make predictions for the future)
recommendations to leverage the
situation to one’s advantage)
Prof. Deepika G , Dept. of CSE, SVIT Page 16
Studied smart, not hard — thanks to [Link]
Analytics 1.0 Analytics 2.0 Analytics 3.0
What will happen? When will it
Key
What happened? Why did it happen? Why will it happen? What
questions What will happen? Why will it happen?
happen? should be the action taken to take
asked:
advantage of what will happen?
Data from legacy systems, Big data is being taken up seriously.
A blend of big data and data from
ERP, CRM, and 3rd party Data is mainly unstructured, arriving at
legacy systems, ERP, CRM, and 3rd
applications. Small and a much higher pace. This fast flow of
Data party applications. A blend of big data
.IN
structured data sources. data entailed that the influx of big
sources: and traditional analytics to yield
Data stored in enterprise volume data had to be stored and
insights and offerings with speed and
data warehouses or data processed rapidly, often on massive
impact.
marts. parallel servers running Hadoop.
C
Data is both being internally and
N
Sourcing: Data was internally sourced. Data was often externally sourced.
externally sourced.
SY
In memory analytics, in database
Database appliances, Hadoop clusters,
Technology: Relational databases processing, agile analytical methods,
SQL to Hadoop environments, etc.
machine learning techniques, etc.
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 17
Studied smart, not hard — thanks to [Link]
3.8 WHY IS BIG DATA ANALYTICS IMPORTANT?
1. Reactive – Business Intelligence:
What does Business Intelligence (BI) help us with? It allows the businesses to make
faster and better decisions by providing the right information to the right person at the
right time in the right format. It is about analysis of the past or historical data and then
displaying the findings of the analysis or reports in the form of enterprise dashboards,
alerts, notifications, etc. It has support for both pre-specified reports as well as ad hoc
querying.
2. Reactive – Big Data Analytics:
.IN
Here the analysis is done on huge datasets but the approach is still reactive as it is still
based on static data.
3. Proactive – Analytics:
C
This is to support futuristic decision making by the use of data mining, predictive
N
modeling, text mining, and statistical analysis. This analysis is not on big data as it still
uses the traditional database management practices on big data and therefore has
SY
severe limitations on the storage capacity and the processing capability.
4. Proactive – Big Data Analytics:
This is sieving through terabytes, petabytes, exabytes of information to filter out the
U
relevant data to analyze. This also includes high performance analytics to gain rapid
insights from big data and the ability to solve complex problems using more data.
VT
3.12 TERMINOLOGIES USED IN BIG DATA ENVIRONMENTS
1. In-Memory Analytics
• Data is fetched from RAM (Random Access Memory) instead of slow hard disk storage.
• Pre-processed and frequently-used data (e.g., cubes, aggregates) is stored in-memory.
• Enables faster access, quicker deployment, better insights, and minimal IT involvement.
• Eliminates repeated disk reads for analysis.
2. In-Database Processing
• Also known as in-database analytics.
• Combines data warehouses with analytical systems.
• Analytical computations are done within the database itself.
Prof. Deepika G , Dept. of CSE, SVIT Page 18
Studied smart, not hard — thanks to [Link]
• Avoids time-consuming export of data.
• Saves time and supports complex/extensive computations.
3. Symmetric Multiprocessor System (SMP)
• Uses a single main memory shared by two or more identical processors.
• All processors:
o Have access to I/O devices.
o Are controlled by a single OS instance.
.IN
• Processors are tightly coupled.
• Each processor has its own high-speed cache memory.
C
• Communication happens via a system bus.
N
4. Massively Parallel Processing (MPP)
SY
• Involves multiple processors working in parallel on different parts of the same program.
• Each processor:
o Has its own OS and dedicated memory.
U
o Communicates through messaging interfaces.
VT
• Suitable for large-scale data analytics.
• More complex to program than SMP.
• Referred to as loosely coupled systems (unlike tightly-coupled SMP).
5. Difference Between Parallel and Distributed Systems
• Parallel systems:
o Are tightly coupled.
o Use shared memory.
o User is unaware of how tasks are parallelized.
o Common in parallel databases.
• Distributed systems (explained later in text):
o Are loosely coupled.
o Use multiple systems possibly at different locations.
Prof. Deepika G , Dept. of CSE, SVIT Page 19
Studied smart, not hard — thanks to [Link]
o Require coordination over a network.
.IN
C
N
SY
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 20
Studied smart, not hard — thanks to [Link]
3.12.6 Shared Nothing Architecture
There are three common types of multiprocessor architecture used in high transaction rate
systems:
1. Shared Memory (SM)
• All processors share a common central memory.
• Suitable for systems needing tight coordination.
• Memory access is faster but can cause contention (conflicts).
.IN
2. Shared Disk (SD)
• Each processor has its own private memory.
C
• Disks are shared among multiple processors.
N
• Coordination required to manage access to shared disks.
SY
3. Shared Nothing (SN)
• No sharing of memory or disk among processors.
U
• Each processor has its own private memory and disk.
VT
• Ideal for scalability – processors operate independently.
• Most distributed systems (like Hadoop) follow this model.
Diagram – Figure 3.12: Distributed System
• Shows multiple processors (P1, P2, P3), each serving its own set of users.
• Each processor has its own storage.
• Processors communicate over a network.
• Represents a Shared Nothing Architecture, where processors work independently and
manage their own data.
Prof. Deepika G , Dept. of CSE, SVIT Page 21
Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT
[Link] Advantages of a “Shared Nothing Architecture”
Fault Isolation
• A "Shared Nothing Architecture" helps isolate faults effectively.
• If a fault occurs in a node, it is contained within that specific node.
• Other nodes remain unaffected since communication is done only through messages (or
the absence of them).
Prof. Deepika G , Dept. of CSE, SVIT Page 22
Studied smart, not hard — thanks to [Link]
Scalability
• In shared systems, disk bandwidth and controller become bottlenecks due to resource
sharing.
• In a shared nothing setup, each node has its own resources, eliminating contention.
• No synchronization is needed for shared data access, so nodes can scale easily without
performance degradation.
3.12.7 CAP Theorem
.IN
The CAP Theorem (also known as Brewer's Theorem) states:
In a distributed computing environment, it is impossible for a system to simultaneously
C
guarantee all three of the following:
1. Consistency
N
2. Availability
SY
3. Partition Tolerance
At best, only two of the three can be fully achieved.
U
[Link] CAP Theorem – Term Explanations
VT
1. Consistency
o Every read receives the most recent (last) write.
o No stale or outdated data is returned.
2. Availability
o Every request (read or write) receives a successful response (not necessarily the
most recent).
o The system is always available to serve requests.
3. Partition Tolerance
o The system continues to function even if communication between nodes is lost
(i.e., network failure).
o No single point of failure crashes the entire system.
Prof. Deepika G , Dept. of CSE, SVIT Page 23
Studied smart, not hard — thanks to [Link]
Diagram Summary (Figure 3.14)
A triangle with each corner representing one aspect of the CAP Theorem:
• Consistency
• Availability
• Partition Tolerance
.IN
C
N
SY
Training Institute Example
U
You visit Amey (Office Admin) to check your schedule:
1. You ask Amey if you're scheduled to train at 3:00 PM.
VT
2. Amey checks his records and finds no such training scheduled.
3. You recall that the training coordinator informed yesterday about this session.
4. Amey suspects it could have been Joey (the second admin) who got the update.
5. Joey confirms: You are indeed scheduled for training at 3:00 PM.
Problem Identified:
濿 "A clear case of inconsistent system!!!"
• The update was shared only with Joey, but you were checking with Amey.
• Both records were not synchronized, leading to data inconsistency.
Resolution Plan by Training Coordinator:
To ensure data consistency, the coordinator proposes:
Prof. Deepika G , Dept. of CSE, SVIT Page 24
Studied smart, not hard — thanks to [Link]
1. Dual Update Rule:
o If either Amey or Joey gets an update, they must update their own file and inform
the other to update theirs too.
2. Email Backup Rule:
o If the other admin is unavailable, the first one must send all updates via email.
3. Update on Resumption:
o When the unavailable admin returns to duty, he must update his file using the
email updates.
.IN
Key Learnings:
• The inconsistency happened due to a lack of synchronization between two systems
(admins).
C
N
• The solution ensures that availability may be slightly compromised (delays), but
consistency is preserved.
SY
• This directly reflects the CAP Theorem:
o You cannot guarantee Consistency, Availability, and Partition Tolerance
simultaneously.
U
o This solution chooses to sacrifice some availability to preserve consistency.
VT
忿 Summary: CAP Theorem Trade-Offs
You can only choose two out of three:
1. Consistency (C):
o All users always get the latest, updated information.
2. Availability (A):
o The system always responds (e.g., admins always provide schedule info).
3. Partition Tolerance (P):
o System functions even if parts (admins/servers) can’t communicate.
俿 When to Choose What?
• Choose Availability over Consistency when:
o Business can tolerate some delay in data sync.
Prof. Deepika G , Dept. of CSE, SVIT Page 25
Studied smart, not hard — thanks to [Link]
• Choose Consistency over Availability when:
o Business requires accurate, up-to-date reads/writes.
㿿 Examples of Database Systems Based on CAP:
Combination Description Example Databases
System is always responsive and
AP (Availability + Riak, Cassandra,
continues despite partitions, but may
Partition Tolerance) CouchDB, DynamoDB
have inconsistent data.
.IN
CP (Consistency + Strongly consistent but may become HBase, MongoDB, Redis,
Partition Tolerance) unavailable during network issues.
C MemcacheDB, BigTable
CA (Consistency + Consistent and always available, but can't Traditional RDBMS,
N
Availability) tolerate partitions. PostgreSQL, MySQL, etc.
SY
U
VT
Data Growth and Units (Table 2.2)
• Data Size Units (Binary scale):
o Bit = 0 or 1
o Byte = 8 bits
o Kilobyte (KB) = 1024 bytes
Prof. Deepika G , Dept. of CSE, SVIT Page 26
Studied smart, not hard — thanks to [Link]
o Megabyte (MB) = 1024² bytes
o Gigabyte (GB) = 1024³ bytes
o Terabyte (TB) = 1024⁴ bytes
o Petabyte (PB) = 1024⁵ bytes
o Exabyte (EB) = 1024⁶ bytes
o Zettabyte (ZB) = 1024⁷ bytes
.IN
o Yottabyte (YB) = 1024⁸ bytes
• Data Size Units (Decimal scale as per Figure 2.6):
C
o 1 KB = 1,000 bytes
N
o 1 MB = 1,000,000 bytes
o 1 GB = 1,000,000,000 bytes
SY
o 1 TB = 1,000,000,000,000 bytes
o 1 PB = 1,000,000,000,000,000 bytes
U
o 1 EB = 1,000,000,000,000,000,000 bytes
VT
o 1 ZB = 1,000,000,000,000,000,000,000 bytes
o 1 YB = 1,000,000,000,000,000,000,000,000 bytes
Note: Binary units are based on powers of 2 (1024), while decimal units are powers of 10 (1000).
Figure 3.15 – CAP Triangle
• Illustrates that you can pick only two of the three: C, A, or P.
• Each DB system aligns with one of the triangle’s edges:
o CA: Traditional RDBMS (e.g., PostgreSQL)
o CP: Systems like HBase, Redis
o AP: Systems like Cassandra, DynamoDB
Prof. Deepika G , Dept. of CSE, SVIT Page 27
Studied smart, not hard — thanks to [Link]
4.1 NoSQL (NOT ONLY SQL)
• The term NoSQL was first coined by Carlo Strozzi in 1998 for his lightweight, open-source
relational database that did not expose the standard SQL interface.
• In 2009, Johan Oskarsson reintroduced the term at an event focused on open-source
distributed networks.
• The hashtag #NoSQL was coined by Eric Evans and others to describe non-relational
databases.
.IN
Key Features of NoSQL Databases:
1. They are open source.
2. They are non-relational.
C
3. They are distributed.
N
4. They are schema-less.
SY
5. They are cluster-friendly.
6. They emerged from 21st-century web applications.
U
4.1.1 Where is it Used?
VT
• Widely used in big data and real-time web applications.
• Used to stock large volumes of data for analysis.
• Ideal for storing social media data and other data types that are difficult to store and
analyze in traditional RDBMS.
4.1.2 What is it?
• NoSQL stands for Not Only SQL.
• These databases are:
o Non-relational
o Open source
o Distributed
• Popular due to:
o Ability to scale out or scale horizontally.
o Aptitude for handling structured, semi-structured, and unstructured data.
Prof. Deepika G , Dept. of CSE, SVIT Page 28
Studied smart, not hard — thanks to [Link]
Additional Feature of NoSQL Databases:
• Non-relational: They do not follow a relational data model.
• Types of NoSQL databases include:
o Key–value pairs
o Document-oriented
o Column-oriented
o Graph-based databases.
Characteristics of NoSQL Databases:
.IN
1. Are non-relational
C
N
o NoSQL databases do not adhere to the relational data model.
They can be key-value pairs, document-oriented, column-oriented, or graph-
SY
o
based databases.
Figure 4.1: Where to use NoSQL?
U
• Log analysis
• Social networking feeds
VT
• Time-based data (not easily analyzed in traditional RDBMS)
Figure 4.2: What is NoSQL?
• Non-relational data storage systems
• No fixed table schema
• No joins
Prof. Deepika G , Dept. of CSE, SVIT Page 29
Studied smart, not hard — thanks to [Link]
• No multi-document transactions
• Relaxes one or more ACID properties
.IN
C
N
Additional Features:
SY
2. Are distributed:
o Data is distributed across several nodes in a cluster, which is often made up of low-
U
cost commodity hardware.
VT
3. Offer no support for ACID properties (Atomicity, Consistency, Isolation, Durability):
o They generally do not support traditional ACID transaction properties.
o Instead, they adhere to Brewer's CAP theorem (Consistency, Availability, and
Partition tolerance), often compromising on consistency in favor of availability
and partition tolerance.
4. Provide no fixed schema:
o NoSQL databases allow flexible schema design.
o They do not require data to strictly follow any schema, making them suitable for
evolving or semi-structured data.
4.1.3 Types of NoSQL Databases
NoSQL databases are non-relational and broadly classified into two types:
1. Key-Value (the big hash table)
o Maintains a large hash table of keys and values.
Prof. Deepika G , Dept. of CSE, SVIT Page 30
Studied smart, not hard — thanks to [Link]
o Examples: Dynamo, Redis, Riak, Amazon S3, Scalaris.
o Sample Key-Value Pair:
Key Value
First Name Simmonds
Last Name David
.IN
2. Schema-less (includes Document, Column, and Graph-based)
o These databases do not have a fixed schema.
C
o Examples:
N
Document: MongoDB, Apache CouchDB, Couchbase, MarkLogic.
Store data as collections of documents.
SY
Sample Document:
{
U
"Book Name": "Fundamentals of Business Analytics",
VT
"Publisher": "Wiley India",
"Year of Publication": "2011"
Column: Cassandra, HBase.
Each storage block contains data from only one column.
Graph-based: Neo4j.
Prof. Deepika G , Dept. of CSE, SVIT Page 31
Studied smart, not hard — thanks to [Link]
Graph Databases (also called network databases)
• Store data in nodes.
• Examples: Neo4j, HyperGraphDB.
• Sample graph structure example:
o Nodes:
▪ ID: 1001, Name: John, Age: 28
▪ ID: 1002, Name: Joe, Age: 32
▪ ID: 1003, Name: Group, Age: AAA
o Relationships:
▪ John knows Joe since 2002
.IN
▪ John is a member of Group since 2003
▪ Joe is a member of Group since 2002
4.1.4 Why NoSQL?
C
N
1. Scale-out architecture: Instead of monolithic relational databases.
2. Can house large volumes of structured, semi-structured, and unstructured data.
SY
3. Dynamic schema: Allows insertion of data without a predefined schema. Supports
faster development and easier code integration.
U
4. Auto-sharding: Automatically spreads data across multiple servers, balancing load and
providing quick recovery if a server goes down.
VT
5. Replication: Supports replication for high availability, fault tolerance, and disaster
recovery.
4.1.5 Advantages of NoSQL
• Can easily scale up and down: Supports rapid, elastic scaling including scaling to the
cloud.
Table 4.1: Popular Schema-less Databases
Key-Value Data Column-Oriented Document Data Graph Data
Store Data Store Store Store
Riak Cassandra MongoDB Infinite Graph
Redis HBase CouchDB Neo4j
Membase Hyper Table Raven DB Allegro Graph
Prof. Deepika G , Dept. of CSE, SVIT Page 32
Studied smart, not hard — thanks to [Link]
.IN
More Detailed Advantages of NoSQL
C
N
1. Cluster scale:
o Supports distribution across 100+ nodes (often in multiple data centers).
SY
2. Performance scale:
o Handles over 100,000+ database reads/writes per second.
3. Data scale:
o Can store over 1 billion+ documents.
U
2. Doesn’t require a pre-defined schema
VT
• NoSQL (e.g., MongoDB) allows records with different sets of key-value pairs.
• Example (from MongoDB):
{
"_id": 101,
"BookName": "Fundamentals of Business Analytics",
"AuthorName": "Seema Acharya",
"Publisher": "Wiley India"
},
{
"_id": 102,
"BookName": "Big Data and Analytics"
}
Prof. Deepika G , Dept. of CSE, SVIT Page 33
Studied smart, not hard — thanks to [Link]
3. Cheap, easy to implement
• Allows benefits of scale, fault tolerance, high availability at low cost.
4. Relaxes the data consistency requirement
• Adheres to the CAP theorem (Compromises consistency, but ensures availability and
partition tolerance).
• Supports eventual consistency.
5. Data can be replicated and partitioned
.IN
(a) Sharding
• Distributes data automatically across multiple servers.
• Handles server addition/removal without application downtime.
• Balances data and query load.
C
(b) Replication
N
• Stores multiple copies of data across clusters/data centers.
• Ensures high availability and fault tolerance.
SY
4.1.6 What We Miss With NoSQL?
U
• NoSQL solves scalability and schema flexibility issues.
VT
• But some RDBMS features are still superior – details in next figure (Figure 4.5).
1. Joins – NoSQL databases generally do not support JOIN operations like SQL databases.
2. Group By – Advanced aggregations like GROUP BY can be complex or unsupported.
3. ACID properties – No built-in support for Atomicity, Consistency, Isolation,
Durability.
Prof. Deepika G , Dept. of CSE, SVIT Page 34
Studied smart, not hard — thanks to [Link]
4. SQL – No standard query language (although MongoDB uses its own query syntax, and
Cassandra uses CQL).
5. Easy integration with other SQL-based applications – Since it lacks standard SQL, it
is harder to integrate with tools designed for SQL-based databases.
However, MongoDB and Cassandra mitigate this to some extent with:
• MongoDB Query Language
• CQL (Cassandra Query Language)
.IN
4.1.7 Use of NoSQL in Industry
NoSQL is used across various industries, particularly for big data and real-time applications.
C
Breakdown of how different NoSQL models are used:
N
1. Key–Value Pairs
o Use: Shopping carts, user data analysis
SY
o Examples: Amazon, LinkedIn
2. Column-Oriented
o Use: Analyze large user actions, sensor feeds
o Examples: Facebook, Twitter, eBay, Netflix
U
3. Document Based
o Use: Real-time analytics, logging, document management
VT
o Examples: MongoDB, CouchDB use cases
4. Graph-Based
o Use: Network modeling, recommendations, upsell and cross-sell analytics
o Examples: Walmart, social networks
Prof. Deepika G , Dept. of CSE, SVIT Page 35
Studied smart, not hard — thanks to [Link]
4.1.8 NoSQL Vendors
A few popular NoSQL vendors and their products:
Company Product Most Widely Used by
Amazon DynamoDB LinkedIn, Mozilla
Facebook Cassandra Netflix, Twitter, eBay
Google BigTable Adobe Photoshop
4.1.9 SQL vs NoSQL
.IN
C
Table 4.3 summarizes the key differences:
N
SQL NoSQL
SY
Relational database Non-relational, distributed database
Relational model Model-less approach
U
Pre-defined schema Dynamic schema for unstructured data
VT
Table-based databases Document, graph, wide-column, or key-value store
Vertically scalable Horizontally scalable (scale-out using clusters)
Uses SQL Uses UnQL (Unstructured Query Language)
Not ideal for large datasets Preferred for large datasets
Not a good fit for hierarchical data Ideal for hierarchical/JSON-like data
Emphasis on ACID properties Follows CAP theorem (sacrifices ACID)
Excellent vendor support Heavily community supported
Supports complex queries Weak at complex querying
Prof. Deepika G , Dept. of CSE, SVIT Page 36
Studied smart, not hard — thanks to [Link]
SQL NoSQL
Can be configured for strong
Some support eventual consistency (e.g., Cassandra)
consistency
Examples: Oracle, MySQL, Examples: MongoDB, Cassandra, HBase, Redis, Neo4j,
PostgreSQL, etc. CouchDB
.IN
4.1.10 NewSQL
• NewSQL is a modern RDBMS that combines features of both SQL and NoSQL.
It offers the scalability of NoSQL systems used for Online Transaction Processing
•
C
(OLTP).
N
• At the same time, it maintains the ACID guarantees of traditional databases.
SY
• NewSQL supports the relational data model and uses SQL as the primary interface.
[Link] Characteristics of NewSQL
U
• Based on the shared-nothing architecture.
• Uses a SQL interface for application interaction.
VT
4.1.11 Comparison of SQL, NoSQL, and NewSQL
Feature SQL NoSQL NewSQL
Adherence to ACID properties Yes No Yes
OLTP/OLAP Yes No Yes
Schema rigidity Yes No Maybe
Adherence to
Adherence to data model No Yes
relational model
Data Format Flexibility No Yes Maybe
Scale out
Scale up (Vertical
Scalability (Horizontal Scale out
Scaling)
Scaling)
Distributed Computing Yes Yes Yes
Community Support Huge Growing Slowly growing
Prof. Deepika G , Dept. of CSE, SVIT Page 37
Studied smart, not hard — thanks to [Link]
4.2 HADOOP
• Hadoop is an open-source project by the Apache Foundation.
• It is a framework written in Java, originally developed by Doug Cutting in 2005.
• Named after Doug Cutting’s son’s toy elephant.
• Initially developed to support distribution for Nutch (a text search engine).
• Inspired by Google MapReduce and Google File System.
• Now a core part of computing infrastructure for companies like Yahoo, Facebook,
LinkedIn, Twitter, etc.
.IN
4.2.1 Features of Hadoop
• Optimized to handle massive amounts of structured, semi-structured, and unstructured
data using inexpensive, commodity hardware.
C
• Based on a shared-nothing architecture.
N
• Data replication across multiple computers ensures fault tolerance and availability.
SY
• Designed for high throughput (not low latency). It's batch-oriented.
• Supports OLTP and OLAP, but not a replacement for traditional RDBMS.
• Not ideal when work cannot be parallelized or when data dependencies exist.
U
• Not suitable for small files—performs best with huge datasets.
VT
4.2.2 Key Advantages of Hadoop
Stores data in its native format:
• HDFS allows storing data without enforcing structure at input.
• Structure is applied only during processing.
Scalable:
• Can store and distribute data across thousands of low-cost servers.
• Proven scalability (e.g., used by Facebook & Yahoo).
Cost-effective:
• Reduced cost per terabyte due to scale-out architecture.
Resilient to failure:
• Ensures data replication across nodes for fault tolerance.
• Automatically recovers from node failures.
Prof. Deepika G , Dept. of CSE, SVIT Page 38
Studied smart, not hard — thanks to [Link]
Flexible:
• Handles all types of data (structured, semi-structured, unstructured).
• Supports various applications: log analysis, data mining, recommendation systems,
market campaigns, etc.
Fast:
• High-speed data processing due to “move code to data” paradigm.
.IN
C
N
SY
U
VT
4.2.3 Versions of Hadoop
There are two versions of Hadoop available:
1. Hadoop 1.0
2. Hadoop 2.0
Prof. Deepika G , Dept. of CSE, SVIT Page 39
Studied smart, not hard — thanks to [Link]
.IN
• [Link] Hadoop 1.0 C
• It has two main parts:
N
1. Data storage framework:
▪ A general-purpose file system called Hadoop Distributed File System
SY
(HDFS).
▪ HDFS is schema-less.
▪ It simply stores data files.
U
▪ These data files can be in just about any format.
VT
2. Data processing framework:
Based on a simple functional programming model called MapReduce (popularized by Google).
o Uses two key functions:
▪ Map: Takes in a set of key–value pairs and generates intermediate data
(another list of key–value pairs).
▪ Reduce: Acts on the intermediate data to produce the output data.
o Map and Reduce work in isolation from each other.
o Processing is highly distributed, fault-tolerant, and scalable.
Limitations of Hadoop 1.0
1. Requirement for MapReduce expertise:
o Developers needed to be proficient in MapReduce and programming languages like
Java.
2. Batch processing only:
Prof. Deepika G , Dept. of CSE, SVIT Page 40
Studied smart, not hard — thanks to [Link]
o Hadoop1.0 supported only batch processing.
o Suitable for:
▪ Log analysis
▪ Large-scale data mining
o Not suitable for:
▪ Other kinds of projects requiring real-time or interactive processing.
3. Tight coupling with MapReduce:
o Hadoop 1.0 was tightly coupled with MapReduce.
o Vendors had two poor options:
▪ Rewrite their tools to work with MapReduce.
.IN
▪ Extract data from HDFS and process it outside Hadoop.
o Both options led to inefficiencies due to data movement in and out of the Hadoop
cluster.
o
C
[Link] Hadoop 2.0
N
• HDFS remains the data storage framework in Hadoop 2.0.
SY
• Introduced a new resource management framework called:
o YARN (Yet Another Resource Negotiator):
▪ Supports dividing applications into parallel tasks.
▪ Enhances flexibility, scalability, and efficiency.
U
▪ Replaces the old JobTracker with ApplicationMaster.
▪ Replaces TaskTracker with NodeManager.
VT
▪ Allows any application (not just MapReduce) to run on Hadoop.
• Key Benefits:
o MapReduce expertise is no longer required.
o Supports both:
▪ Batch processing
▪ Real-time processing
o MapReduce is no longer the only option:
▪ Alternative data processing tools are now supported.
▪ Native features like data standardization and master data management can
be performed in HDFS.
Prof. Deepika G , Dept. of CSE, SVIT Page 41
Studied smart, not hard — thanks to [Link]
.IN
HDFS (Hadoop Distributed File System)
• It is the distributed storage unit of Hadoop.
C
• Provides streaming access to file system data.
N
• Supports file permissions and authentication.
SY
• Based on GFS (Google File System).
• Scales a single cluster node to hundreds or thousands of nodes.
U
• Handles large datasets on commodity hardware.
VT
• HDFS is highly fault-tolerant:
o Stores files across multiple machines.
o Files are stored in a redundant fashion to allow data recovery in case of failure.
(Example Scenario)
• An e-commerce website stores millions of customers' data in a distributed manner.
• Data has been collected over 4–5 years.
• Batch analytics is run on archived data to analyze:
o Customer behavior
o Buying patterns
o Preferences
Prof. Deepika G , Dept. of CSE, SVIT Page 42
Studied smart, not hard — thanks to [Link]
o Requirements
• Helps identify:
o Which products are purchased
o In which months
o By which types of customers
HBase
• Stores data in HDFS.
• It is the first non-batch component of the Hadoop ecosystem.
.IN
• Works as a database on top of HDFS.
• Offers quick random access to stored data.
• Has very low latency compared to HDFS.
• It is a:
C
o NoSQL database
N
o Non-relational
o Column-oriented database
SY
• Data Structure:
o A table can have thousands of columns.
o A row can have several column families.
U
o Each column family can contain multiple columns.
o Each column can have several key-value pairs.
VT
• Based on: Google BigTable
• Used by: Facebook, Twitter, Yahoo, etc.
(Example Scenario)
• The same e-commerce website also stores millions of product data.
• To search among millions of products and get real-time results, optimization is needed.
• HBase supports real-time analytics.
• Due to high data velocity, they chose HBase over HDFS, as HDFS does not support
real-time writes.
• Impact:
o Query time reduced from 3 days to 3 minutes.
Prof. Deepika G , Dept. of CSE, SVIT Page 43
Studied smart, not hard — thanks to [Link]
Difference Between HBase and Hadoop/HDFS
HDFS (Hadoop Distributed File
Aspect HBase (Hadoop Database)
System)
Type File System NoSQL Database
Analogy Like NTFS Like MySQL
Real-time random read and
Write/Read Model WORM (Write Once Read Many)
.IN
write
Underlying System Based on Google File System (GFS)
C Based on Google BigTable
Random access to small
Data Access Pattern Full table scan or partition scan
ranges
N
Performance with
SY
Very good 4–5 times slower than HDFS
Hive
Via Java APIs, REST, Avro,
Access Methods Via MapReduce jobs only
Thrift
U
Storage Type Static and rigid Dynamic and flexible
VT
Latency High latency Low latency
Best Suited For Batch processing and analytics Real-time analytics
Hadoop Ecosystem Components for Data Processing
1. MapReduce
• A programming model for processing large datasets in parallel.
• Developed by Google in 2004.
• Works in two phases:
o Map: Converts input data into key-value pairs.
o Reduce: Combines and processes these pairs to generate output.
Prof. Deepika G , Dept. of CSE, SVIT Page 44
Studied smart, not hard — thanks to [Link]
• Data flows from and back to HDFS.
2. Spark
• A programming and computing model developed at UC Berkeley in 2009.
• Open-source and written in Scala.
• Performs in-memory processing (much faster than disk-based).
• Can use HDFS as a data source but does not require MapReduce.
.IN
• Can run with or without Hadoop.
Spark Libraries: C
• Spark SQL – Supports querying using SQL.
N
• Spark Streaming – For real-time data analysis.
• MLlib – For machine learning tasks.
SY
• GraphX – For graph computations.
U
Key Notes
VT
• Hadoop is used for batch processing of unstructured data.
• Spark is widely used for fast, real-time, in-memory processing.
Hadoop Ecosystem Components for Data Analysis
1. Pig
• A high-level scripting language used with Hadoop.
• Acts as an alternative to MapReduce.
• Developed by Yahoo.
Pig has two main parts:
• Pig Latin:
oA scripting language that gets translated into MapReduce jobs.
o Used for ETL (Extract, Transform, Load) and data analysis.
Prof. Deepika G , Dept. of CSE, SVIT Page 45
Studied smart, not hard — thanks to [Link]
o Supports operations like grouping, filtering, sorting, joining, etc.
o Loads data from HDFS, processes it, and sends it back or displays it.
• Pig runtime:
o The execution environment where Pig scripts are run.
2. Hive
• A data warehouse software on top of Hadoop.
.IN
• Used for summarization, querying, and analysis.
• Uses HQL (Hive Query Language) or HiveQL, similar to SQL.
C
• Converts SQL-like queries into MapReduce jobs for execution on Hadoop.
N
Difference Between Hive and RDBMS
SY
Aspect Hive RDBMS (MySQL, SQL Server, etc.)
Schema Enforced on read (schema-on-
U
Enforced on write (schema-on-write)
Enforcement read)
VT
Fast, as schema is not enforced Slower due to schema enforcement
Data Load
during insert during load
Query Better read performance after Good performance after schema
Performance faster load validation
Data Access
Write once, read many times Read and write many times
Model
Batch-oriented, not ideal for Designed for OLTP and day-to-day
Processing Type
OLTP transactions
Analytical workloads (OLAP-like
Best Suited For Transactional workloads (OLTP)
use cases)
Queries are converted into Queries are executed directly in the
Query Execution
MapReduce jobs database engine
Prof. Deepika G , Dept. of CSE, SVIT Page 46
Studied smart, not hard — thanks to [Link]
4.2.5 Hadoop Distributions
Hadoop is an open-source Apache project.
Core components of Hadoop:
1. Hadoop Common
2. HDFS (Hadoop Distributed File System)
3. YARN (Yet Another Resource Negotiator)
.IN
4. MapReduce
Several companies have built custom distributions of Hadoop for easier use:
C
• IBM
N
• Amazon Web Services
• Microsoft
SY
• Teradata
• Hortonworks
U
• Cloudera
VT
Despite different strategies, all distributions aim to:
• Distribute data and workloads across many servers.
• Make big data manageable and scalable.
4.2.5 Hadoop Distributions
• Hadoop is an open-source project from Apache, freely available for download.
• Core components of Hadoop include:
1. Hadoop Common
2. Hadoop Distributed File System (HDFS)
3. Hadoop YARN (Yet Another Resource Negotiator)
4. Hadoop MapReduce
Prof. Deepika G , Dept. of CSE, SVIT Page 47
Studied smart, not hard — thanks to [Link]
• Various companies have created their own distributions (packaged versions) of Hadoop
to make it more user-friendly and consumable. These include: IBM, Amazon, Web
Services, Microsoft, Teradata, Hortonworks, Cloudera.
• Though their strategies differ slightly, the main goal is to distribute data and workloads
across many servers for scalable and manageable big data processing.
.IN
C
N
SY
U
VT
4.2.8 Hadoop versus SQL
Hadoop SQL
Scale out Scale up
Key-Value pairs Relational table
Functional Programming Declarative Queries
Offline batch processing Online transaction processing
Prof. Deepika G , Dept. of CSE, SVIT Page 48
Studied smart, not hard — thanks to [Link]
• Scale out (Hadoop) means adding more machines to handle increased workload, while
scale up (SQL) means increasing the power of a single machine.
• Hadoop uses a key-value pair data model, whereas SQL uses relational tables.
• Hadoop programming is functional in nature, while SQL uses declarative queries.
• Hadoop is optimized for offline batch processing of large datasets, while SQL is designed
for online transaction processing (OLTP) with quick response times.
4.2.7 Integrated Hadoop Systems Offered by Leading Market Vendors
These vendors package and offer Hadoop-based big data platforms that are ready for
.IN
enterprise use.
The major vendors mentioned are:
• EMC Greenplum
C
• Oracle Big Data Appliance
N
• Microsoft Big Data Solution
• IBM InfoSphere
SY
• HP Big Data Solutions
•
These integrated systems typically combine hardware, software, and support services optimized
for Hadoop workloads, helping organizations deploy big data solutions more easily and
U
efficiently.
VT
4.2.8 Cloud-Based Hadoop Solutions
Amazon Web Services (AWS): Provides a comprehensive portfolio of cloud services aimed at
managing big data efficiently. AWS emphasizes reducing costs, scaling on demand, and
accelerating innovation speed.
Google Cloud Storage connector for Hadoop: Allows direct execution of MapReduce jobs on
data stored in Google Cloud Storage without copying it locally. This simplifies Hadoop
deployment, reduces costs, and offers performance comparable to Hadoop's HDFS. It also
enhances reliability by removing the single point of failure (name node).
Prof. Deepika G , Dept. of CSE, SVIT Page 49
Studied smart, not hard — thanks to [Link]
.IN
C
N
SY
U
VT
Prof. Deepika G , Dept. of CSE, SVIT Page 50
Studied smart, not hard — thanks to [Link]