0% found this document useful (0 votes)
18 views48 pages

Understanding Big Data Analytics

The document discusses the significance of Big Data, highlighting its generation from various sources such as social media, sensors, and transactions. It emphasizes the challenges and opportunities in managing and analyzing vast amounts of data, which requires new skills and technologies. The document also outlines the characteristics of Big Data, including volume, velocity, and variety, and presents key use cases and technologies associated with Big Data analytics.

Uploaded by

sayedheiba88
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views48 pages

Understanding Big Data Analytics

The document discusses the significance of Big Data, highlighting its generation from various sources such as social media, sensors, and transactions. It emphasizes the challenges and opportunities in managing and analyzing vast amounts of data, which requires new skills and technologies. The document also outlines the characteristics of Big Data, including volume, velocity, and variety, and presents key use cases and technologies associated with Big Data analytics.

Uploaded by

sayedheiba88
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Big Data Analytics

Data is the new Oil. Data is just like crude. It’s valuable, but if unrefined it
cannot really be used.
– Clive Humby, DunnHumby

We have for the first time an economy based on


a key resource [Information] that is not only renewable,
but self-generating. Running out of it is not a problem,
but drowning in it is.
– John Naisbitt

2
What Does ‘Big Data’ Mean
How is big data generated?
Sensors gathering information: e.g. Social media: posts, pictures and
Climate, traffic etc. videos

Digital satellite images

Purchase transaction records

Mobile phone GPS signals


High volume administrative
& transactional records
Who’s Generating Big Data

Mobile devices
(tracking all objects all the time)

Social media and networks Scientific instruments


(all of us are generating data) (collecting all sorts of data)

Sensor technology and networks


(measuring all kinds of data)

• The progress and innovation is no longer hindered by the ability to collect data
• But, by the ability to manage, analyze, summarize, visualize, and discover
knowledge from the collected data in a timely manner and in a scalable fashion

5
With Big Data, We’ve Moved into a New Era of Analytics

12+ terabytes 5+ million


of Tweets trade events
create daily. per second.

Volume Velocity

Variety Veracity
100’s Only 1 in 3
of different decision makers trust
types of data. their information.
Whom does it matter
• Research Community 
• Business Community - New tools, new
capabilities, new infrastructure, new business
models etc.,
• On sectors

Financial Services..
The number of organizations who see analytics as a
competitive advantage70is%growing.

57 %

63%
2010 BUSINESS
2011 IMPERATIVE
2012

business initiative
IQ
Studies show that organizations competing on
analytics outperform their peers
substantially outperform

IBM IBV/MIT Sloan Management Review Study 2011


Copyright Massachusetts Institute of Technology 2011

1.6x Revenue
Growth 2.5x Stock Price
Appreciation 2.0x EBITDA
Growth
9
Usage Example of Big Data 12

US 2012 Election

- predictive modeling - data mining for


- [Link] individualized ad targeting
- drive traffic to other campaign sites
Facebook page (33 million "likes") - Orca big-data app
YouTube channel (240,000 subscribers
and 246 million page views). - YouTube channel( 23,700 subscribers
- a contest to dine with Sarah Jessica Parker and 26 million page views)
- Every single night, the team ran 66,000
computer simulations, Reddit!!! - Ace of Spades HQ
- Amazon web services
The 5 Key Big Data Use Cases

Big Data Exploration Enhanced 360o View Security/Intelligence


Find, visualize, understand all of the Customer Extension
big data to improve decision Extend existing customer views Lower risk, detect fraud and
making (MDM, CRM, etc) by monitor cyber security in real-
incorporating additional internal time
and external information
sources

Operations Analysis Data Warehouse Augmentation


Analyze a variety of machine Integrate big data and data warehouse
data for improved business results capabilities to increase operational efficiency

1 © 2013 IBM Corporation


How much data?
• Google processes 20 petabytes (PB) a day (2017)
• We produce every day 2.5 quintillion bytes of data (2018)
• Hive is Facebook’s data warehouse, with 300 PB of data (2019).
• Facebook generates 4 new petabytes of data per day (2019)
• Facebook sees 100 million hours of daily video watch time (2019).
• Facebook users generate 4 million likes every minute (2019).

640K ought to be
enough for anybody.
Challenges

How to transfer Big Data?


AN ACUTE SHORTAGE OF SKILLS THREATENS OUR ABILITY TO ADDRESS
EMERGING OPPORTUNITIES AND RISKS

Among organizations worldwide today…

have major skill gaps in


mobile, business analytics,
and security

has all the skills it needs to be successful


applying advanced technology* for
business benefit

report a skills shortage in the ability to


manage information

* Includes business analytics, mobile computing, social business, and cloud computing

14 Sources: IBM Tech Trends report 2012, Enterprise Strategy Group, CompTIA
BIG DATA REQUIRES A BROAD SET OF SKILLS
"By 2015, big data demand will reach 4.4 million jobs globally, but only one-
third of those jobs will be filled."
Source: Gartner "Gartner's Top Predictions for IT Organizations and Users, 2013 and Beyond: Balancing Economics, Risk, Opportunity and
Innovation" 19 Oct 2012

Math and
Operations Research
Expertise
Data Experts Develop analytic algorithms
Decision Making
Data architecture, management,
Executive and
governance, policy Management
Apply information to solve
business issues

Tool Developers
Industry Vertical
Mask complexity and
analytics to lower skills
Domain Expertise
boundaries Develop hypothesis, identify
relevant business issues,
Visualization
ask the right questions
Expertise
Interpret data sets,
determine correlations and
15 present in meaningful ways
THE BIG DATA APPROACH TO ANALYTICS IS DIFFERENT

Traditional Analytics Big Data Analytics


Structured & Repeatable Iterative & Exploratory
Structure built to store data Data is the structure

Business IT Team
Users Delivers Data
Analyzed
Determine Information On Flexible
Questions Platform

Available Information Analyze ALL Available Information


Capacity constrained down sampling of Whole population analytics connects the
available information dots
Analyzed
IT Team
Analyzed
Information
Information Business
Builds System Users
To Answer Explore and
Known Questions Ask Any Question
Carefully cleanse a small information before Analyze information as is & cleanse as needed
any analysis & existing repeatable
THE BIG DATA APPROACH TO ANALYTICS IS DIFFERENT

Traditional Analytics Big Data Analytics


Structured & Repeatable Iterative & Exploratory
Structure built to store data Data is the structure
Hypothesis Question Data Exploration

?
All Information

Analyzed
Information

Answer Data Actionable Insight Correlation


Start with hypothesis Data leads the way
Test against selected data Explore all data, identify correlations

Analyze after landing… Analyze in motion…


The Myth About Big Data

 Big Data Is New


 Big Data Is Only About Massive Data Volume
 Big Data Means Hadoop
 Big Data Need A Data Warehouse
 Big Data Means Unstructured Data
 Big Data Is for Social Media & Sentiment Analysis
Big Data Technologies
Cloud Computing Parallel Computing

NoSQL Databases

General Programming

Data Visualization

Machine Learning
Short Explanation
OpenStack is an open source platform that uses pooled virtual resources to build and manage private
and public clouds. The tools that comprise the OpenStack platform, called "projects," handle the core
cloud-computing services of compute, networking, storage, identity, and image services.

Apache Mahout is a powerful, scalable machine-learning library that runs on top of Hadoop
MapReduce. Mahout is supported by its 3 pillars: Recommender engines: Recommenders can be
classified as being user based or item based and can be used to attract users and suggest products by
mining user behaviour.

SAS is an analytics software used by a number of sectors, including healthcare, finance and retail. It is
used for advanced analytics, data management and business intelligence. SAS has a strong market share
in the analytics software market, with a significant presence in the healthcare sector.

Apache Spark is an open-source, distributed processing system used for big data workloads. It utilizes in-
memory caching, and optimized query execution for fast analytic queries against data of any size.

NoSQL is a type of database management system (DBMS) that is designed to handle and store large
volumes of unstructured and semi-structured data. Unlike traditional relational databases that use
tables with pre-defined schemas to store data, NoSQL databases use flexible data models that can adapt
to changes in data structures and are capable of scaling horizontally to handle growing amounts of data.

The term NoSQL originally referred to “non-SQL” or “non-relational” databases, but the term has since
evolved to mean “not only SQL,” as NoSQL databases have expanded to include a wide range of different
database architectures and data models.
What is big data?

• “Every day, we create 2.5 quintillion bytes of data — so


much that 90% of the data in the world today has been
created in the last two years alone. This data comes
from everywhere: sensors used to gather climate
information, posts to social media sites, digital pictures
and videos, purchase transaction records, and cell phone
GPS signals to name a few.
This data is “big data.”
Terminology for using and analyzing data

Term Time Frame Specific Meaning

Decision Support 1970 - 85 Use of data analysis to support decision


making
Executive 1980 - 90 Focus on data analysis for decision by
Support senior executives
OLAP 1990 - 2000 Software for analyzing multi-dimensional
tables
Business 1989 - 2005 Tools to support data driven decisions with
Intelligence emphasis on reporting

Analytics 2005 - 2010 Focus on statistical and mathematical


analysis for decisions
Big Data 2010 - Present Focus on very large unstructured fast
moving data
What is Big Data?
“Big data are high volume, high velocity, and high variety information assets that require
new forms of processing to enable enhanced decision making, insight discovery and
process optimization” (Gartner 2012)

Volume Vertical scalability


- exceeds limits of traditional Requires - ability to grow storage to
column and row relational DB accommodate new ‘records’
- constantly growing

Velocity Data streaming


Requires - real time processing,
- arrives rapidly, often in real
analysis and transformation
time

Variety Horizontal scalability


- does not have a standard Requires - ability to add additional data
structure, e.g. text, images structures
Characteristics of Big Data:
1-Scale (Volume)

• Data Volume
– 44x increase from 2009 to 2020
– From 0.8 zettabytes to 35zb
• Data volume is increasing exponentially

Exponential increase in
collected/generated data

24
Characteristics of Big Data:
2-Complexity (Varity)
• Various formats, types, and structures
• Text, numerical, images, audio, video,
sequences, time series, social media
data, multi-dim arrays, etc…
• Static data vs. streaming data
• A single application can be
generating/collecting many types of
data

To extract knowledge all these types of data


need to linked together

25
Characteristics of Big Data:
3-Speed (Velocity)

• Data is begin generated fast and need to be


processed fast
• Online Data Analytics
• Late decisions  missing opportunities
• Examples
– E-Promotions: Based on your current location, your purchase history, what
you like  send promotions right now for store next to you

– Healthcare monitoring: sensors monitoring your activities and body  any


abnormal measurements require immediate reaction

26
Simple to start
• What is the maximum file size you have dealt so far?
– Movies/Files/Streaming video that you have used?
– What have you observed?
• What is the maximum download speed you get?
• Simple computation
– How much time to just transfer.
Huge amount of data
• There are huge volumes of data in the world:
+ From the beginning of recorded time until 2003,
+ We created 5 billion gigabytes (exabytes) of data.
+ In 2011, the same amount was created every two days
+ In 2013, the same amount of data is created every 10
minutes.
Big data spans three dimensions: Volume, Velocity and Variety
• Volume: Enterprises are awash with ever-growing data of all types, easily amassing
terabytes—even petabytes—of information.
– Turn 12 terabytes of Tweets created each day into improved product sentiment
analysis
– Convert 350 billion annual meter readings to better predict power consumption
• Velocity: Sometimes 2 minutes is too late. For time-sensitive processes such as catching
fraud, big data must be used as it streams into your enterprise in order to maximize its
value.
– Scrutinize 5 million trade events created each day to identify potential fraud
– Analyze 500 million daily call detail records in real-time to predict customer churn
faster
– The latest I have heard is 10 nano seconds delay is too much.
• Variety: Big data is any type of data - structured and unstructured data such as text,
sensor data, audio, video, click streams, log files and more. New insights are found when
analyzing these data types together.
– Monitor 100’s of live video feeds from surveillance cameras to target points of
interest
– Exploit the 80% data growth in images, video and documents to improve customer
satisfaction
Finally….
`Big- Data’ is similar to ‘Small-data’ but bigger

.. But having data bigger it requires different


approaches:
Techniques, tools, architecture
… with an aim to solve new problems
Or old problems in a better way
BIG DATA is not just HADOOP

Large scale data processing Spark Analytics Engine

Understand and navigate


Federated Discovery and Navigation
federated big data sources

Manage & store huge volume Hadoop File System


of any data MapReduce

Structure and control data Data Warehousing

Manage streaming data Stream Computing

Analyze unstructured data Text Analytics Engine

Integrate and govern all Integration, Data Quality, Security,


data sources Lifecycle Management, MDM
Types of tools typically used in Big
Data Scenario
• Where is the processing hosted?
– Distributed server/cloud
• Where data is stored?
– Distributed Storage (eg: Amazon s3)
• Where is the programming model?
– Distributed processing (Map Reduce)
• How data is stored and indexed?
– High performance schema free database
• What operations are performed on the data?
– Analytic/Semantic Processing (Eg. RDF/OWL)
When dealing with Big Data is hard
• When the operations on data are complex:
– Eg. Simple counting is not a complex problem.
– Modeling and reasoning with data of different kinds can
get extremely complex
• Good news with big-data:
– Often, because of the vast amount of data, modeling
techniques can get simpler (e.g., smart counting can
replace complex model-based analytics)…
– …as long as we deal with the scale.
Why Big-Data?
• Key enablers for the appearance and growth
of ‘Big-Data’ are:
+ Increase in storage capabilities
+ Increase in processing power
+ Availability of data
The Meaning of Big Data - 4 V’s
• Big Volume
– With simple (SQL) analytics
– With complex (non-SQL) analytics

• Big Velocity
– Drink from the fire hose

• Big Variety
– Large number of diverse data sources to integrate
• Veracity
- Verification of the diverse data from diverse sources
Big Volume - Little Analytics
• Well addressed by data warehouse crowd

• Who are pretty good at SQL analytics on


– Hundreds of nodes
– Petabytes of data
The Big Data Stack
Typical Datawarehouse Environment
A Big Data Technology Eco-system
Big Data and Datawarehouse existence
The Analytics Process Model
Analytical Model Requirements: Critical Success Factors

• Business relevance
• Statistical performance
• Interpretable and justifiable
• Operationally efficient
• Economical cost
• Compliance with regulation and legislation
(Local and International)
17
Implementation of Big Data
Platforms for Large-scale Data Analysis
• Parallel DBMS technologies
– Proposed in late eighties
– Matured over the last two decades
– Multi-billion dollar industry: Proprietary DBMS Engines intended as
Data Warehousing solutions for very large enterprises
• Map Reduce
– pioneered by Google
– popularized by Yahoo! (Hadoop)
18
Implementation of Big Data
MapReduce Parallel DBMS technologies
 Popularly used for more than two decades
• Overview:  Research Projects: Gamma, Grace, …
– Data-parallel programming model
 Commercial: Multi-billion dollar industry
– An associated parallel and distributed but access to only a privileged few
implementation for commodity clusters  Relational Data Model
• Pioneered by Google  Indexing
– Processes 20 PB of data per day  Familiar SQL interface
• Popularized by open-source Hadoop  Advanced query optimization
– Used by Yahoo!, Facebook,  Well understood and studied
Amazon, and the list is growing …
20
Implementation of Big Data
MapReduce Advantages
• Automatic Parallelization:
– Depending on the size of RAW INPUT DATA  instantiate
multiple MAP tasks
– Similarly, depending upon the number of intermediate <key,
value> partitions  instantiate multiple REDUCE tasks
• Run-time:
– Data partitioning
– Task scheduling
– Handling machine failures
– Managing inter-machine communication
• Completely transparent to the programmer/analyst/user
21
Implementation of Big Data
Map Reduce vs Parallel DBMS
Parallel DBMS MapReduce

Schema Support  Not out of the box

Indexing  Not out of the box


Imperative
(C/C++, Java, …)
Programming Declarative
Extensions
Model (SQL)
through
Pig and Hive
Optimizations
(Compression,
 Not out of the box
Query Optimiza-
tion)
Flexibility Not out of the box 
Coarse grained tech-
Hadoop…..

• Simple analytics
– X100 times a parallel DBMS
• Complex analytics (Mahout or roll-your-own)
– X100 times Scalapack
• Parallel programming
– Parallel grep (great)
• Hadoop lacks
– Stateful computations
– Point-to-point communication
Thank you…..

You might also like