Big Data Analytics
Data is the new Oil. Data is just like crude. It’s valuable, but if unrefined it
cannot really be used.
– Clive Humby, DunnHumby
We have for the first time an economy based on
a key resource [Information] that is not only renewable,
but self-generating. Running out of it is not a problem,
but drowning in it is.
– John Naisbitt
2
What Does ‘Big Data’ Mean
How is big data generated?
Sensors gathering information: e.g. Social media: posts, pictures and
Climate, traffic etc. videos
Digital satellite images
Purchase transaction records
Mobile phone GPS signals
High volume administrative
& transactional records
Who’s Generating Big Data
Mobile devices
(tracking all objects all the time)
Social media and networks Scientific instruments
(all of us are generating data) (collecting all sorts of data)
Sensor technology and networks
(measuring all kinds of data)
• The progress and innovation is no longer hindered by the ability to collect data
• But, by the ability to manage, analyze, summarize, visualize, and discover
knowledge from the collected data in a timely manner and in a scalable fashion
5
With Big Data, We’ve Moved into a New Era of Analytics
12+ terabytes 5+ million
of Tweets trade events
create daily. per second.
Volume Velocity
Variety Veracity
100’s Only 1 in 3
of different decision makers trust
types of data. their information.
Whom does it matter
• Research Community
• Business Community - New tools, new
capabilities, new infrastructure, new business
models etc.,
• On sectors
Financial Services..
The number of organizations who see analytics as a
competitive advantage70is%growing.
57 %
63%
2010 BUSINESS
2011 IMPERATIVE
2012
business initiative
IQ
Studies show that organizations competing on
analytics outperform their peers
substantially outperform
IBM IBV/MIT Sloan Management Review Study 2011
Copyright Massachusetts Institute of Technology 2011
1.6x Revenue
Growth 2.5x Stock Price
Appreciation 2.0x EBITDA
Growth
9
Usage Example of Big Data 12
US 2012 Election
- predictive modeling - data mining for
- [Link] individualized ad targeting
- drive traffic to other campaign sites
Facebook page (33 million "likes") - Orca big-data app
YouTube channel (240,000 subscribers
and 246 million page views). - YouTube channel( 23,700 subscribers
- a contest to dine with Sarah Jessica Parker and 26 million page views)
- Every single night, the team ran 66,000
computer simulations, Reddit!!! - Ace of Spades HQ
- Amazon web services
The 5 Key Big Data Use Cases
Big Data Exploration Enhanced 360o View Security/Intelligence
Find, visualize, understand all of the Customer Extension
big data to improve decision Extend existing customer views Lower risk, detect fraud and
making (MDM, CRM, etc) by monitor cyber security in real-
incorporating additional internal time
and external information
sources
Operations Analysis Data Warehouse Augmentation
Analyze a variety of machine Integrate big data and data warehouse
data for improved business results capabilities to increase operational efficiency
1 © 2013 IBM Corporation
How much data?
• Google processes 20 petabytes (PB) a day (2017)
• We produce every day 2.5 quintillion bytes of data (2018)
• Hive is Facebook’s data warehouse, with 300 PB of data (2019).
• Facebook generates 4 new petabytes of data per day (2019)
• Facebook sees 100 million hours of daily video watch time (2019).
• Facebook users generate 4 million likes every minute (2019).
640K ought to be
enough for anybody.
Challenges
How to transfer Big Data?
AN ACUTE SHORTAGE OF SKILLS THREATENS OUR ABILITY TO ADDRESS
EMERGING OPPORTUNITIES AND RISKS
Among organizations worldwide today…
have major skill gaps in
mobile, business analytics,
and security
has all the skills it needs to be successful
applying advanced technology* for
business benefit
report a skills shortage in the ability to
manage information
* Includes business analytics, mobile computing, social business, and cloud computing
14 Sources: IBM Tech Trends report 2012, Enterprise Strategy Group, CompTIA
BIG DATA REQUIRES A BROAD SET OF SKILLS
"By 2015, big data demand will reach 4.4 million jobs globally, but only one-
third of those jobs will be filled."
Source: Gartner "Gartner's Top Predictions for IT Organizations and Users, 2013 and Beyond: Balancing Economics, Risk, Opportunity and
Innovation" 19 Oct 2012
Math and
Operations Research
Expertise
Data Experts Develop analytic algorithms
Decision Making
Data architecture, management,
Executive and
governance, policy Management
Apply information to solve
business issues
Tool Developers
Industry Vertical
Mask complexity and
analytics to lower skills
Domain Expertise
boundaries Develop hypothesis, identify
relevant business issues,
Visualization
ask the right questions
Expertise
Interpret data sets,
determine correlations and
15 present in meaningful ways
THE BIG DATA APPROACH TO ANALYTICS IS DIFFERENT
Traditional Analytics Big Data Analytics
Structured & Repeatable Iterative & Exploratory
Structure built to store data Data is the structure
Business IT Team
Users Delivers Data
Analyzed
Determine Information On Flexible
Questions Platform
Available Information Analyze ALL Available Information
Capacity constrained down sampling of Whole population analytics connects the
available information dots
Analyzed
IT Team
Analyzed
Information
Information Business
Builds System Users
To Answer Explore and
Known Questions Ask Any Question
Carefully cleanse a small information before Analyze information as is & cleanse as needed
any analysis & existing repeatable
THE BIG DATA APPROACH TO ANALYTICS IS DIFFERENT
Traditional Analytics Big Data Analytics
Structured & Repeatable Iterative & Exploratory
Structure built to store data Data is the structure
Hypothesis Question Data Exploration
?
All Information
Analyzed
Information
Answer Data Actionable Insight Correlation
Start with hypothesis Data leads the way
Test against selected data Explore all data, identify correlations
Analyze after landing… Analyze in motion…
The Myth About Big Data
Big Data Is New
Big Data Is Only About Massive Data Volume
Big Data Means Hadoop
Big Data Need A Data Warehouse
Big Data Means Unstructured Data
Big Data Is for Social Media & Sentiment Analysis
Big Data Technologies
Cloud Computing Parallel Computing
NoSQL Databases
General Programming
Data Visualization
Machine Learning
Short Explanation
OpenStack is an open source platform that uses pooled virtual resources to build and manage private
and public clouds. The tools that comprise the OpenStack platform, called "projects," handle the core
cloud-computing services of compute, networking, storage, identity, and image services.
Apache Mahout is a powerful, scalable machine-learning library that runs on top of Hadoop
MapReduce. Mahout is supported by its 3 pillars: Recommender engines: Recommenders can be
classified as being user based or item based and can be used to attract users and suggest products by
mining user behaviour.
SAS is an analytics software used by a number of sectors, including healthcare, finance and retail. It is
used for advanced analytics, data management and business intelligence. SAS has a strong market share
in the analytics software market, with a significant presence in the healthcare sector.
Apache Spark is an open-source, distributed processing system used for big data workloads. It utilizes in-
memory caching, and optimized query execution for fast analytic queries against data of any size.
NoSQL is a type of database management system (DBMS) that is designed to handle and store large
volumes of unstructured and semi-structured data. Unlike traditional relational databases that use
tables with pre-defined schemas to store data, NoSQL databases use flexible data models that can adapt
to changes in data structures and are capable of scaling horizontally to handle growing amounts of data.
The term NoSQL originally referred to “non-SQL” or “non-relational” databases, but the term has since
evolved to mean “not only SQL,” as NoSQL databases have expanded to include a wide range of different
database architectures and data models.
What is big data?
• “Every day, we create 2.5 quintillion bytes of data — so
much that 90% of the data in the world today has been
created in the last two years alone. This data comes
from everywhere: sensors used to gather climate
information, posts to social media sites, digital pictures
and videos, purchase transaction records, and cell phone
GPS signals to name a few.
This data is “big data.”
Terminology for using and analyzing data
Term Time Frame Specific Meaning
Decision Support 1970 - 85 Use of data analysis to support decision
making
Executive 1980 - 90 Focus on data analysis for decision by
Support senior executives
OLAP 1990 - 2000 Software for analyzing multi-dimensional
tables
Business 1989 - 2005 Tools to support data driven decisions with
Intelligence emphasis on reporting
Analytics 2005 - 2010 Focus on statistical and mathematical
analysis for decisions
Big Data 2010 - Present Focus on very large unstructured fast
moving data
What is Big Data?
“Big data are high volume, high velocity, and high variety information assets that require
new forms of processing to enable enhanced decision making, insight discovery and
process optimization” (Gartner 2012)
Volume Vertical scalability
- exceeds limits of traditional Requires - ability to grow storage to
column and row relational DB accommodate new ‘records’
- constantly growing
Velocity Data streaming
Requires - real time processing,
- arrives rapidly, often in real
analysis and transformation
time
Variety Horizontal scalability
- does not have a standard Requires - ability to add additional data
structure, e.g. text, images structures
Characteristics of Big Data:
1-Scale (Volume)
• Data Volume
– 44x increase from 2009 to 2020
– From 0.8 zettabytes to 35zb
• Data volume is increasing exponentially
Exponential increase in
collected/generated data
24
Characteristics of Big Data:
2-Complexity (Varity)
• Various formats, types, and structures
• Text, numerical, images, audio, video,
sequences, time series, social media
data, multi-dim arrays, etc…
• Static data vs. streaming data
• A single application can be
generating/collecting many types of
data
To extract knowledge all these types of data
need to linked together
25
Characteristics of Big Data:
3-Speed (Velocity)
• Data is begin generated fast and need to be
processed fast
• Online Data Analytics
• Late decisions missing opportunities
• Examples
– E-Promotions: Based on your current location, your purchase history, what
you like send promotions right now for store next to you
– Healthcare monitoring: sensors monitoring your activities and body any
abnormal measurements require immediate reaction
26
Simple to start
• What is the maximum file size you have dealt so far?
– Movies/Files/Streaming video that you have used?
– What have you observed?
• What is the maximum download speed you get?
• Simple computation
– How much time to just transfer.
Huge amount of data
• There are huge volumes of data in the world:
+ From the beginning of recorded time until 2003,
+ We created 5 billion gigabytes (exabytes) of data.
+ In 2011, the same amount was created every two days
+ In 2013, the same amount of data is created every 10
minutes.
Big data spans three dimensions: Volume, Velocity and Variety
• Volume: Enterprises are awash with ever-growing data of all types, easily amassing
terabytes—even petabytes—of information.
– Turn 12 terabytes of Tweets created each day into improved product sentiment
analysis
– Convert 350 billion annual meter readings to better predict power consumption
• Velocity: Sometimes 2 minutes is too late. For time-sensitive processes such as catching
fraud, big data must be used as it streams into your enterprise in order to maximize its
value.
– Scrutinize 5 million trade events created each day to identify potential fraud
– Analyze 500 million daily call detail records in real-time to predict customer churn
faster
– The latest I have heard is 10 nano seconds delay is too much.
• Variety: Big data is any type of data - structured and unstructured data such as text,
sensor data, audio, video, click streams, log files and more. New insights are found when
analyzing these data types together.
– Monitor 100’s of live video feeds from surveillance cameras to target points of
interest
– Exploit the 80% data growth in images, video and documents to improve customer
satisfaction
Finally….
`Big- Data’ is similar to ‘Small-data’ but bigger
.. But having data bigger it requires different
approaches:
Techniques, tools, architecture
… with an aim to solve new problems
Or old problems in a better way
BIG DATA is not just HADOOP
Large scale data processing Spark Analytics Engine
Understand and navigate
Federated Discovery and Navigation
federated big data sources
Manage & store huge volume Hadoop File System
of any data MapReduce
Structure and control data Data Warehousing
Manage streaming data Stream Computing
Analyze unstructured data Text Analytics Engine
Integrate and govern all Integration, Data Quality, Security,
data sources Lifecycle Management, MDM
Types of tools typically used in Big
Data Scenario
• Where is the processing hosted?
– Distributed server/cloud
• Where data is stored?
– Distributed Storage (eg: Amazon s3)
• Where is the programming model?
– Distributed processing (Map Reduce)
• How data is stored and indexed?
– High performance schema free database
• What operations are performed on the data?
– Analytic/Semantic Processing (Eg. RDF/OWL)
When dealing with Big Data is hard
• When the operations on data are complex:
– Eg. Simple counting is not a complex problem.
– Modeling and reasoning with data of different kinds can
get extremely complex
• Good news with big-data:
– Often, because of the vast amount of data, modeling
techniques can get simpler (e.g., smart counting can
replace complex model-based analytics)…
– …as long as we deal with the scale.
Why Big-Data?
• Key enablers for the appearance and growth
of ‘Big-Data’ are:
+ Increase in storage capabilities
+ Increase in processing power
+ Availability of data
The Meaning of Big Data - 4 V’s
• Big Volume
– With simple (SQL) analytics
– With complex (non-SQL) analytics
• Big Velocity
– Drink from the fire hose
• Big Variety
– Large number of diverse data sources to integrate
• Veracity
- Verification of the diverse data from diverse sources
Big Volume - Little Analytics
• Well addressed by data warehouse crowd
• Who are pretty good at SQL analytics on
– Hundreds of nodes
– Petabytes of data
The Big Data Stack
Typical Datawarehouse Environment
A Big Data Technology Eco-system
Big Data and Datawarehouse existence
The Analytics Process Model
Analytical Model Requirements: Critical Success Factors
• Business relevance
• Statistical performance
• Interpretable and justifiable
• Operationally efficient
• Economical cost
• Compliance with regulation and legislation
(Local and International)
17
Implementation of Big Data
Platforms for Large-scale Data Analysis
• Parallel DBMS technologies
– Proposed in late eighties
– Matured over the last two decades
– Multi-billion dollar industry: Proprietary DBMS Engines intended as
Data Warehousing solutions for very large enterprises
• Map Reduce
– pioneered by Google
– popularized by Yahoo! (Hadoop)
18
Implementation of Big Data
MapReduce Parallel DBMS technologies
Popularly used for more than two decades
• Overview: Research Projects: Gamma, Grace, …
– Data-parallel programming model
Commercial: Multi-billion dollar industry
– An associated parallel and distributed but access to only a privileged few
implementation for commodity clusters Relational Data Model
• Pioneered by Google Indexing
– Processes 20 PB of data per day Familiar SQL interface
• Popularized by open-source Hadoop Advanced query optimization
– Used by Yahoo!, Facebook, Well understood and studied
Amazon, and the list is growing …
20
Implementation of Big Data
MapReduce Advantages
• Automatic Parallelization:
– Depending on the size of RAW INPUT DATA instantiate
multiple MAP tasks
– Similarly, depending upon the number of intermediate <key,
value> partitions instantiate multiple REDUCE tasks
• Run-time:
– Data partitioning
– Task scheduling
– Handling machine failures
– Managing inter-machine communication
• Completely transparent to the programmer/analyst/user
21
Implementation of Big Data
Map Reduce vs Parallel DBMS
Parallel DBMS MapReduce
Schema Support Not out of the box
Indexing Not out of the box
Imperative
(C/C++, Java, …)
Programming Declarative
Extensions
Model (SQL)
through
Pig and Hive
Optimizations
(Compression,
Not out of the box
Query Optimiza-
tion)
Flexibility Not out of the box
Coarse grained tech-
Hadoop…..
• Simple analytics
– X100 times a parallel DBMS
• Complex analytics (Mahout or roll-your-own)
– X100 times Scalapack
• Parallel programming
– Parallel grep (great)
• Hadoop lacks
– Stateful computations
– Point-to-point communication
Thank you…..