0% found this document useful (0 votes)
15 views86 pages

Big Data Analytics Overview and Insights

The document outlines a Big Data Analytics Learning Lab held at the University of Nairobi, focusing on the significance of Big Data, its applications, and the creation of Big Data-enabled organizations. It discusses the characteristics of Big Data, its value across various sectors, and the methods for leveraging Big Data analytics to derive actionable insights. Additionally, it highlights the impact of Big Data analytics on decision-making, marketing, performance improvement, and new business models.

Uploaded by

mheba11
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views86 pages

Big Data Analytics Overview and Insights

The document outlines a Big Data Analytics Learning Lab held at the University of Nairobi, focusing on the significance of Big Data, its applications, and the creation of Big Data-enabled organizations. It discusses the characteristics of Big Data, its value across various sectors, and the methods for leveraging Big Data analytics to derive actionable insights. Additionally, it highlights the impact of Big Data analytics on decision-making, marketing, performance improvement, and new business models.

Uploaded by

mheba11
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

[Link].

com

Big Data Analytics


Learning Lab 1

UN Data Innovation Lab


4
University of Nairobi
March 13-14, 2017
Agenda

I. Introduction to Big Data


What it is and why it matters

[Link] Data Analytics


Putting Big Data to work

[Link] a Big Data-Enabled Organization


Bringing Big Data Analytics home

[Link] Study
‘Nowcasting’ economic activity in Colombia

PwC
Introduction to
Big Data
What it is and Why it
Matters

01
What is Big Data?
“Big Data” exceeds the capacity of traditional analytics and information
management paradigms across what is known as the 4 V’s: Volume, Variety,
Velocity, and Veracity

Veracity Velocity Variety Volume

Uncertainty of Data Analysis of Streaming Different Forms of Scale of Data


Data Data

With exponential The speed at Represents the Reflects the size of


increases of data which data is diversity of the a data set. New
from unfiltered generated and data. Data sets will information is
and constantly used. New data is vary by type (e.g. generated daily
flowing data being created social networking, and in some cases
sources, data every second and media, text) and hourly, creating
quality often in some cases it they will vary how data sets that are
suffers and new may need to be well they are measured in
methods must find analyzed just as structured terabytes and
ways to “sift” quickly petabytes
through junk to
find meaning
PwC 4
The Promise of Big Data
Even more important than its definition is what Big Data promises to
achieve: intelligence in the moment.
Traditional Techniques &
Big Data Differentiators
Issues
Veracity

• Does not account for biases, • Data is stored, and mined meaningful to the
noise and abnormality in problem being analyzed
data
• Keeps data clean and processes to keep ‘dirty
data’ from accumulating in your systems

In real-time:
Velocity

• No real time analysis


• Dynamically analyze data
• Consistently integrate new information
• Auto deletes unwanted to ensure optimal storage

• Compatibility issues • Frameworks accommodate varying data types


Variety

• Advanced analytics struggle and data models


with non-numerical data • Insightful analysis with very few parameters

• Analysis is limited to small data • Scalable for huge amounts of multi-sourced data
Volume

sets
• Facilitation of massively parallel processing
• Analyzing large data sets =
• Low-cost data storage
High Costs & High Memory

PwC 5
Types of Big Data
Variety is the most unique aspect of Big Data. New technologies and new
types of data have driven much of the evolution around Big Data.

Twitter, Linkedin, Facebook, Tumblr,


Images, videos, audio, Flash, live
Blog, SlideShare, YouTube, Google+,
streams, podcasts, etc.
Social Instagram, Flickr, Pinterest, Vimeo,
Media WordPress, IM, RSS, Review, Chatter,
Media
Jive, Yammer, etc.
Medical devices, smart
XLS, PDF, CSV, email,
electric meters, car sensors,
Word, PPT, HTML, HTML 5,
road cameras, satellites,
plain text, XML, JSON, etc. Sensor
Docs traffic recording devices,
data processors found within
vehicles, video games, cable
Government, weather, boxes, assembly lines, office
competitive, traffic, building, cell towers, jet
regulatory, compliance, engines, air conditioning
health care services, units, refrigerators, trucks,
economic, census, public Public Machine Event logs, server data,
farm machinery, etc..
finance, stock, OSINT, the Web Log application logs, business
World Bank, SEC/Edgar, Data process logs, audit logs, call
Wikipedia, IMDb, etc. detail records (CDRs), mobile
location, mobile app usage,
Archives of scanned documents,
clickstream data, etc.
statements, insurance forms, medical Business
Archive Project management, marketing
record and customer correspondence, Apps automation, productivity, CRM, ERP
paper archives, and print stream files
content management system, HR,
that contain original systems of record
storage, talent management,
between organizations and their
procurement, expense management
customers
Google Docs, intranets, portals, etc.
PwC 6
“Single sources of data are no longer
sufficient to cope with the increasingly
complicated problems in many policy
arenas.” 1

Big data “is not notable because of its size,


but because of its relationality to other
data. Due to efforts to mine and aggregate
data, Big Data is fundamentally
networked.”2
(1) M. Milakovich, “Anticipatory Government: Integrating big data for Smaller Government”, in Oxford Internet Institute “Internet, Politics, Policy 2012”
Conference, Oxford, 2012
(2) D. Boyd and K. Crawford, “Six Provocations for big data,” in A Decade in Internet Time: Symposium on the Dynamics of the Internet and Society, 2011
PwC 7
Why is Big Data valuable?
We have identified 5 key areas where Big Data is uniquely valuable:

Accessibility to Enhanced visibility of relevant information and better transparency to


Data massive amounts of data. Improved reporting to stakeholders.

Next generation analytics can enable automated decision making


Decision (inventory management, financial risk assessment, sensor data
Making management, machinery tuning).

Marketing Segmentation of population to customize offerings and marketing


Trends campaigns (consumer goods, retail, social, clinical data, etc).

Performance Exploration for, and discovery of, new needs, can drive organizations
Improvement to fine tune for optimal performance and efficiency (employee data).

New Business Discovery of trends will lead organizations to form new business
models to adapt by creating new service offerings for their customers.
Models/Service
Intermediary companies with big data expertise will provide analytics
s
to 3rd parties.

PwC 8
$1 One study estimated the potential value of big
data in the U.S. health care, European public
sector administration, global personal location

Trillion data, U.S. retail, and global manufacturing to be


over $1 trillion U.S. dollars per year. 1

Another study estimated the value of big data in


the areas of customer intelligence, supply chain
$41
intelligence, performance improvements, fraud
detection, and quality and risk management to
be $41 billion per year in the UK alone. 2 Billion
(1) J. Manyika, M. Chui, B. Brown, J. Bughin, R. Dobbs, C. Roxburgh and A. H. Byers, “Big data: The next frontier for innovation, competition, and productivity,”
McKinsey & Company, 2011.
(2) Centre for Economics and Business Research, “Data equity: unlocking the value of big data,” SAS, 2012.
PwC 9
Not to be confused with…

Structured, semi-structured
or unstructured information
distinguished by one or
more of the four “V”s:
Veracity, Velocity, Variety,
Volume. Big Data Open
Data
Public, freely available
data

Crowdsourced
Data

Data collected through contributions


from a large number of individuals.

Graphic and definitions based on “Big Data in Action for Development,” World Bank, [Link]

PwC 10
Big Data
Analytics
Putting Big Data to Work

02
It’s not just about the data…
It is important to understand the distinction between Big Data sets (large,
unstructured, fast, and uncertain data) and ‘Big Data Analytics’.

Big + Big Data Analytics


Data
Refers to the DATA only Methods of using Big Data to generate insight

Machine Learning/Deep • Leveraging a computer’s ability to learn


1 Learning without being explicitly programmed to
solve business problems

• Understanding value drivers from the


IoT (Internet of Things)
2 & Sensor Analytics
ever-growing network of connected
physical objects and the communication
between them

3 Modeling Willingness-
• Mining product reviews to estimate
willingness-to-pay for product features
to-Pay

4 • Understanding human speech as it is


spoken through application of computer
Natural Language science, AI, and computational
Processing linguistics
5 • Using distributed computing and
machine learning tools to analyze
hundreds of gigabytes of data
Analyzing Data @ Scale
6 • Mining social data in real time to
understand when and where consumers
Creating a Streaming are making choices
Consumer Behavior Data
Lake

PwC 12
… It’s also about what, how, and why you use it
Big Data Analytics – the process of harnessing Big Data to yield actionable
insights – is a combination of five key elements:

Decisions Analytics Data Technology Mindset & Skills

Big Data Analytics


To leverage the variety Big Data Analytics is
The value of Big Data To store, manage, and requires firm
and volume of Big Data about operationalizing
Analytics is driven by the use Big Data often commitment to using
while managing its new and more data, but
unique decisions facing requires investments in analytics in decision-
volatility, advanced it is also about data
leaders, companies, and new technologies and making; a decisive
analytical approaches quality, data
countries today. In turn, data processing mentality capable of
are necessary, such as interoperability, data
the type, frequency, methods, such as employing in-the-
natural language disaggregation, and the
speed, and complexity distributed processing moment intelligence;
processing, network ability to modularize
of decisions drive how (e.g., Hadoop), NoSQL and investment in
analysis, simulative data structures to
Big Data Analytics is storage, and Cloud analytical technology,
modeling, artificial quickly absorb new data
deployed. computing. resources, and skills.
intelligence, etc. and new types of data.

PwC 13
Big Data Analytical Capabilities
Continuing increases in processing capacity have opened the door to a range
of advanced algorithms and modeling techniques that can produce valuable
insights from Big Data.
Structured Unstructured

Time Series Signal Analysis Cluster Analysis


Regression
Traditional

Analysis Distinguish Discover


Discover
Discover between noise and meaningful
relationships
relationships over meaningful groupings of data
between variables
time information points

A/B/N Testing
Experiment to find Classification Simulation Spatial Analysis
the most effective Organize data Modeling Extract geographic
variation of a points into known Experiment with a or topological
website, product, categories system virtually information
etc
Sentiment
Visualization Complex Event
Predictive Analysis
Use visual Processing
Modeling Extract consumer
representations of Combine data
Use data to forecast reactions based on
data to find and sources to
Emerging

or infer behavior social media


communicate info recognize events
behavior
Network Analysis Deep QA Natural Language
Optimization
Discover Find answers to Processing
Improve a process
meaningful nodes human questions Extract meaning
or function based
and relationships on using artificial from human speech
on criteria
networks intelligence or writing

PwC * For more information on these analytic methods, see Appendix. 14


Forward-Looking vs. Rear-View Analytics
Big Data Analytics improves the speed and efficiency with which we
understand the past, and opens up entirely new avenues for preparing for and
adapting to the future.
Rear-view Forward-looking

Continuous
Analytics
Prescriptive
Analytics How do we adapt
Predictive to change?
What should be
Analytics Monitor, decide, and
done?
Increasing Business Value

act autonomously or
Diagnostic Recommend ‘right’
What could semi-autonomously
Analytics or optimal actions
happen?
Descriptive or decisions
Why did it Predict future • Monitor results on a
Analytics outcomes based on continuous basis
happen? • Real-time product
Identify causes of the past • Dynamically adjust
What happened? and service
Describe, trends and propositions (graph strategies based on
outcomes • Forward-looking
summarize and analysis, entity changing
view of current and
analyze historical • Observed behavior resolution on data environment and
future value
data or events lakes to infer improved
• Sentiment Scoring predictions
• Non-traditional data present customer
• Observed behavior • Graph analysis and need) • Agent-based and
sources such as
or events Natural Language • Rapid evaluation of dynamic simulation
social listening and
• Non-traditional data Processing to multiple ‘what-if’ models, time-series
web crawling
sources such as identify hidden scenarios analysis
• Statistical and relationships and
social listening and • Optimization
regression analysis themes
web crawling decisions and
• Dynamic • Dual objective
actions
visualization models
• Behavioral
Increasing Sophistication of Data & Analytics
economics
PwC 15
Examples of Big Data Analytics in Action
Market Leaders are leveraging Big Data Analytics to generate value by
starting with a business need and focusing on implementing actionable
insights quickly and decisively
Business Need Big Data Analytics Impact
Company Business Need Data and Analytics Impact
Greater tailoring of credit Statistical model based on public Net revenue grew at a
card offers to fit customer credit and demographic data to CAGR of 32% from 1994
needs target customized products to to 2003; prompted
customers competitors to shift focus
to data and analytics
Data-enabled engine Analysis of sensor data from Over 70% annual
prognostics, monitoring, hundreds of sensors in 4,000 revenue from the aircraft
maintenance and repair engines to identify and solve engine division
issues weeks in advance attributable to this
service
Search-to-purchase Semantic search, which enables Increases 10-15% the
conversion by anticipating discovery using algorithms that likelihood that a customer
intent of a shopper’s search rank results via social signals will complete their
and delivering relevant results from around the web purchase – translating to
millions of dollars in
revenue
Transformation from Analysis of data from 66 million Revenue and
subscription streaming service subscribers’ viewing habits and subscriber base
to original content producer preferences increased by 15% and
9% respectively in 2013
Leverage Internet of Things Launched software to help Estimated 1% reduction
(IoT) by connecting machines airlines and railroads move their in fuel costs, projected to
PwC to facilitate data-enabled data to the cloud and predict save the airline industry 16
prognostics, increase efficiency mechanical malfunctions, $30 billion over 15
Big Data Analytics in Development
Big Data Analytics is making an equally impressive impact on Development
interventions – allowing decision-makers to reach and serve previously
neglected populations.
Business Need Big Data Analytics Impact
Company Business Need Data and Analytics Impact
More transparent, reliable, Web scraping of online price data Government statistical
and low-cost method to used to produce price indices, offices shifting to
track inflation in Argentina and econometric analysis used to accept Big Data. Central
model disaggregated impacts of banks using Big Data to
policies see day-to-day volatility.
Understand how migrants Iterative analysis of call detail Informing labor policy
act as arbitrageurs to bring records (CDRs) to track design in low-income
labor markets into equilibrium movement of migrants in countries to incentivize
response to local shocks to labor or disincentivize
demand (weather, economy, migratory behavior
conflict, etc.)
The city of Rio de Janeiro The city combines data from 30 Rio has improved
wanted to improve its city agencies – including weather, emergency response
emergency response by satellite, video, GPS, historic time by 30%, catalogued
better predicting heavy rainfall rainfall, and topographic survey 200+ flood points, and
and subsequent severe data – in a central Operations can now predict heavy
landslides and flooding Center rains 48 hours in advance
on a half-km basis
Create a better ecosystem Remote crowdsourced data M-PESA is being used to
for mobile services in the gathered via cell phones used to lower costs for farmers
agricultural sectors of connect farmers to markets, to receive loans and
Kenya, Tanzania, and assess farmers’ credit perform transactions
Mozambique worthiness, and incubate new with distributers and
PwC mobile businesses with greater buyers, as well as to 17
predictors of success provide geography-
Creating a Big
Data-Enabled
Organization
Bringing Big Data home

03
Step 1: Be Yourself
Beginning with a clear understanding of the specific questions you intend to
use Big Data Analytics to address can help guide where and which data
solutions are deployed.
Value
enhancement Delivering future value
• Data-driven decision-making in real time
• Use analytics to develop new
programs/opportunities
• Relies heavily on data supplied by others
• Often struggles to move away from
exclusively intuitive decision-making

Strategic
Enabling strategy and improving
performance
• Use analytics to reduce political
divergence and drive consensus
• Real-time analytics to enable quick
responses to events
• Use data to develop personalized
Tactical services
• Need for more objective and higher
Value quality data
Day to day operations
enablement • Struggle to move from narrow focus on
reactive operations to more proactive,
comprehensive management of daily
operations
• High value for digitization of operational
processes across program units
Operational • Often already proficient in traditional
business intelligence

PwC 19
Step 2: Secure People & Skills
The competencies required of “data scientists” within an analytics
organization or project converge from multiple skill domains.

Subject Area or
Expertise in Domain Deep understanding of industry,
statistical Expertise subject area, or research
techniques, tools domain to help determine which
and languages questions need answering and
used to run on what frequency, specificity,
analyses that or geography
generate insights
to effectively
determine and Computer
communicate Statistical & Science &
actionable insights Mathematica Programming
l
Comfort in programming
across various languages,
a thorough understanding
of external and internal
data sources, data
gathering, storing, and
Organization-
retrieving methods which
specific
Organization-specific knowledge about Information help combine disparate
data assets – including enterprise Knowledge data sources to generate
“metadata” – their location and unique insights
appropriate business context for use in
advanced analytics
PwC 20
Step 3: Let objectives dictate structure, not vice versa
How analytics efforts or organizations are structured – whether reporting is
vertically or horizontally aligned, how interconnected or autonomous separate
units are, how resources and successes are shared – can influence efficiency
and impact. Distributed Analytics Federated Analytics Centralized Analytics
CENTRAL Analytics CENTRAL Analytics CENTRAL Analytics
Competency Center Competency Center Competency Center

LOCAL Metadata Metadata Metadata


Repository Repository Repository

ETL ETL ETL

Data Data Data


Warehouse LOCAL Warehouse Warehouse

Data Mart Data Mart Data Mart

BI BI BI
Applications Applications Applications

• Adopt previously proven • Subject area-specific innovations • Governance


Objectives practices • Repeatable models • Aligning analytics to organization-
• Highly focused analytics support wide strategy
Data • Deployed locally • Deployed locally • Deployed and managed centrally
Warehouses, • Some data and models shared
Marts, etc. across groups

• Managed locally • Managed locally, but connected • Controlled centrally, with units
Analytics Tools to group framework having access to shared
resources
• Placed within individual units • Placed within individual units • Placed within central analytics
Analytics Staff/
• Skills tailored to specific region or team, available as needed to
Competencies subject matter support individual units

PwC 21
The ‘Hub-Spoke’ operating model often serves as a well-
synchronized, connected system
4 3 2 1
Local Local Centers of Competency Global
Business Central Business
Adoption Excellence Center
Operations Decision Hub Strategy
of Practices (Regional) (‘Standards’)

4 4
Local
Local
‘Spoke
‘Spoke
Local ’

‘Spoke
’ 3
4 4
Sample Hub-Spoke Local
Local
‘Spoke Interaction Model Center of ‘Spoke
Competency Excellenc ’

Center e
3
2
4 (Regional)
Center of Central 4
Local Local
Excellenc Decision
‘Spoke ‘Spoke
’ e 1 Hub

(Regional)
4
Local ‘Standardizatio Local
4

‘Spoke n’ ‘Spoke
’ ’

4
Center of
Local Excellenc Local
4

‘Spoke e ‘Spoke
’ 3 (Regional) ’
PwC 22
Step 4: Invest in Appropriate Infrastructure
Big Data introduces challenges related to data volume and variety, processing
constraints, and new data structures that traditional data infrastructure is not
equipped to support
Objective Considerations Impact
Dictates performance needs along with
Analysis Type data structures and processing
Identify the type of
analysis that will architecture
Analytics be conducted and Analysis
Interface could restrict the ability to
Capabilities define which Flexibility perform analysis ad hoc and restrict ability
analytics Analysis to update
capabilities will be Structures Support for analysis specific data
employed structures can improve performance and
reduce analysis effort
Size
Size of data sets introduce need for
scalable infrastructure and
Data Variety Define the data set Structure performance
that will be used
for the analysis Variability of source data models and
including its Sources data set structure require data model
sources, size, and flexibility
structure Diverse sources will require scalability,
Frequency
model flexibility, and flexible interfaces

Application Speed Frequency of analysis will dictate the


Define the processing architecture (batch or real
timeliness and time)
frequency of the Interfaces
The timeliness of the analysis will
analysis results for
PwC impact the need for scalability and 23
reporting and
performance
downstream
Emerging Infrastructure Options
To harness Big Data, storage solutions must be able to support targeted
analytics capabilities, data diversity and performance needs

Distributed Processing Hadoop and similar solutions that


provide scalable distributed storage and
distributed computation on commodity
hardware
NoSQL Embedded and persisted storage that
implement data models through
document, graph, and dictionary
structures
Cloud Computing scalability and cost management and
enable a cohesive business strategy
Cloud computing can improve flexibility, across a org
Traditional challenges being addressed…

• Scalability Issues • Data storage solutions need to


• Big Data set information extraction provide flexible data models to better
and queries require large volumes of ingest unstructured and semi
processing cycles that can quickly structured data
scale • Need to combine and link multiple
data sources

PwC * For more information on these infrastructure options, see Appendix. 24


PwC
25
Summary: Key Guiding Principles for developing best-
in-class analytics organization

Guiding Principles – Illustrative, May be Customized

1. Establish the Analytics organization as an objective advisor for insight


generation .

2. Ensure responsiveness to business needs by balancing ‘consolidation’ with


‘distribution’ of analytics functionality where it makes sense.

3. Innovate, invest in, and build new analytics capabilities, and gradually push them
out to the business as user sophistication matures (e.g., data visualization).

4. Prioritize strategic business value delivery over tactical outputs.

5. Ensure adequate attention to user experience.

6. Focus on speed, accuracy, and reusability.

7. Optimize and manage work-flow to achieve maximum resource efficiency.

8. Allow distributed analytics where it makes sense, but tightly govern and ensure
cataloguing.

9. Ensure a consistent feedback loop of all outputs that are created.

PwC 26
Case Study
‘Nowcasting’ Economic
Activity in Colombia

04
Situation
In Colombia, the leading economic
indicators used to analyze economic
activity have an average lag of 10 weeks.
This presents challenges for the well-
timed design of economic policy and
monitoring of economic shocks or trends.

The Colombian Ministry of Finance looked


for coincident indicators that could allow
tracking the short-term trends of
economic activity.

Characteristics of Data Needed:


• Real-time
• Highly disaggregated – by sector, geography,
etc.
• Statistically correlated with key economic
trends (consumption, GDP, etc.)
• Robust enough of a sample to be
PwC 28
representative of the economy as a whole
Group Discussion

What Big Data sources could the Colombian Ministry of


Finance potentially use to reliably approximate sectorial
economic activity in real-time?

PwC 29
Brainstorming Breakout

In groups of 3-4, take five-ten minutes to brainstorm how the


Ministry could approach answering the following questions:
• What data should it consider using?
• Is this data the Ministry already has available, or will this require the
Ministry to acquire an entirely new source of data?
• How does the cost of acquiring this data – whether by their own collection
or through external data partnership – compare to the expected benefits of
using it?
• If this data is new to the Ministry, what entities may already have this data
in possession?
• How might the Ministry ensure its staff have the skills necessary to
acquire, manage, and use this data? Is this data uniquely complex such
that it may require more advanced or entirely new skillsets?
• What should the Ministry consider in the way of data storage and
security? How extensively may it be required to overhaul data storage
infrastructure to accommodate using this data?

PwC 30
Solution

Based on web searches performed by Google users, Google Trends (GT)


provides daily information about the query volume for a given search term in
a given geographic region. For Colombia, GT data are available at the
departmental level and also for the largest municipalities.
The Colombian Administrative Department for National Statistics (DANE – for
its acronym in Spanish) combined indexes built using GT data with its own
official economic activity data (both at the aggregate level and at the sectorial
level) – both of which are publicly available – to construct leading indicators
that determine, in real-time, the short-term trend of different economic
sectors, as well as their turning points.
In some sense, the GT data takes the place of traditional consumer-sentiment
surveys. For example, the use of data for a certain keyword (such as the
brand for a certain product) might be justified in the case a drop or surge in
the web searches for that keyword could be linked to a fall or increase in its
demand and, therefore, a lower or higher production for the specific sector
producing that product.
PwC 31
Example: “Ahorro” vs. Unemployment Rate

Ahorro – savings

PwC 32
Example: “Ahorro” vs. Unemployment Rate

These trends were shown to correlate with a high coefficient


of correlation with traditional measures of unemployment.

PwC 33
Example: “Zapatos” vs. Employment Rate

Zapatos – shoes

PwC 34
Example: “Zapatos” vs. Employment Rate

These trends were shown to correlate with a high coefficient


of correlation with traditional measures of employment.

PwC 35
Find Out More
Melanie Thomas Armstrong Jean Young
Leading Partner Managing Director
International Public Sector International Public Sector Data Analytics
+1 (202) 320-7098 +1 (703) 918-1001
[Link]@[Link] [Link]@[Link]

Bill Stephens Mariola Pogacnik


Director Director
International Public Sector Data Analytics United Nations & International Public Sector
+1 (703) 635-0800 +1 (646) 471-5467
[Link]@[Link] [Link]@[Link]

Ashraf Faramawi Jared Nyarumba


Manager Manager
International Public Sector Data Analytics Data Analytics, Africa
+1 (202) 271-5711 +254 710 623 426
[Link]@[Link] [Link]@[Link]

This publication has been prepared for general guidance on matters of interest only, and does not constitute professional advice. You should not act upon the
information contained in this publication without obtaining specific professional advice. No representation or warranty (express or implied) is given as to the
accuracy or completeness of the information contained in this publication, and, to the extent permitted by law, PricewaterhouseCoopers LLP, its members,
employees and agents do not accept or assume any liability, responsibility or duty of care for any consequences of you or anyone else acting, or refraining to
act, in reliance on the information contained in this publication or for any decision based on it.

© 2017 PricewaterhouseCoopers LLP. All rights reserved. In this document, “PwC” refers to PricewaterhouseCoopers LLP which is a member firm of
PricewaterhouseCoopers International Limited, each member firm of which is a separate legal entity.
PwC 36
Appendix
Emerging Data
Storage and
Infrastructure
Options
Building an Analytics Organization: Critical Components
Emerging Infrastructure – Data Storage
Options
Distributed Processing Hadoop and similar solutions that
provide scalable distributed storage and
distributed computation on commodity
hardware

Introduction to Hadoop

• Hadoop is based on work done by


Google in early 2000s (combination of Faster and Lower Cost
Google File System (GFS) and Analysis
MapReduce)
• Useful for analyzing copious amounts of
complex data across multiple data
sources
Linear Scalability
• Distributes data as it is initially stored
in the system
• Applications are written in high-level
code
• Computation happens where data is Greater flexibility
stored, whenever possible
• Data is replicated multiple times on the
PwC 39
system for increased availability and
Distributed Storage and Analytics: Hadoop vs. Traditional
Data Stores
Compared to traditional data stores, Hadoop provides greater
flexibility when it comes to storing data and scaling to meet demand.
Traditional Data
Hadoop vs.
Stores

Supports both structured Supports only structured


Data Structure and unstructured data. data.

Limited depending on
Data Size Unlimited
selected RDBMS.

Supports various
serialization and data Supports a single tabular
Data Formats formats (e.g. text, JSON, data format.
XML, etc)
Distributed scaling from Scaling is possible, but is
the ground up – simply add typically more complex and
Scaling more nodes to increase cannot be performed at a
capacity node level.
Sources: [Link]
PwC 40
Building an Analytics Organization: Critical Components
Emerging Infrastructure – Data Storage
Options
NoSQL Embedded and persisted storage that
implement data models through
document, graph, and dictionary
structures
NoSQL - Storage Types

Key – Value Columnar Document


Graph Store
Store Store Store

Increasing Data Complexity


Pros: Simplicity & Pros: Scalability & Pros: Easy to Use Pros: Graph Joins
Scalability Flexibility Cons: Scalability Cons: Flexibility
Cons: Lack of advanced Cons: Complexity
features/queries
Solution Examples

PwC 41
Building an Analytics Organization: Critical Components
Emerging Infrastructure – Data Storage
Options
Cloud Computing The model is compelling; cloud computing can improve
flexibility, scalability and cost management. Businesses best
able to realize the potential will establish a cohesive business
strategy as cloud computing can transform your entire
organization — people, processes, and systems

Cloud transformation begins at the


infrastructure level and leads to more agile
applications, resulting in faster speed to
market and more flexibility to meet client
needs.

The key benefits, beyond consolidation,


include standardized application and
development environments, resulting in better
controlled and more efficient application
lifecycles.
Source: PwC, “Digital IQ Snapshot: Cloud,”; PwC, “FS Viewpoint: Clouds is
the forecast”
PwC 42
Text Mining and
Natural
Language
Processing
Data Mining, Text Mining, and Natural Language
Processing
What are they and how are they used?
Natural Language
Processing
NLP is a theoretically
motivated range of
computational
Text Mining techniques for
Analysis of large analyzing and
quantities of representing
natural language naturally occurring
Data Mining text and detecting texts
lexical or linguistic at one or more levels
Extraction of of linguistic analysis
usage patterns to
implicit, previously for the purpose of
extract probably
unknown, and
useful information achieving human-like
potentially useful language processing
information
Source: Text Mining, Ian from
Witten, 2004 for a range of tasks or
data applications.
PwC 44
Natural Language Processing and Text Mining
What are they and how are they used?

Natural Language Processing Text Mining


Purpose and Overview Purpose and Overview
NLP (Natural Language Text mining represents a system
Processing) applies statistical or of statistical analysis and
rules based computational classification algorithms that
techniques to evaluate and are employed to explore groups
model texts at various levels of of natural language texts and
linguistic analysis in order to identify useful patterns,
identify key concepts, enable relationships, and knowledge
intelligent processing and draw
Objectives Objectives
inferences
• Deep analysis and structuring of individual • Use of data mining techniques and statistical
texts through phrase identification, part of methods to conduct a shallow analysis of
speech tagging, and word disambiguation groups of documents and make accessible
• Identification of a text’s message or meaning knowledge within structured/semi-structured
through the use of linguistic analysis: texts
• Syntactic - sentence structure or • Develop a structured view of a documents
breakdown contents in order to develop linkages among
• Lexical - meaning of words within the texts for classification, categorization,
context of use knowledge discovery and search
• Semantic - logical meaning of phrases or • Conduct a statistical analysis of the
text word/sentence usage and attributes in order to
• Discourse - connections among sentences identify key phrases, summarize texts, extract
and phrases that define the topic information from groups of texts, and discover
• Generate natural language sentences or texts new knowledge using the information within
PwCas a response to an input/ question using a texts 45
context text or knowledge base
NLP Tools
Tools and APIs that provide capabilities to parse and structure
natural language texts for machine analysis

Tool Description Analysis Type


A machine learning based toolkit for the • Tokenization • Named entity extraction
OpenNLP processing of natural language text. Link • sentence segmentation • Chunking, parsing
• Part-of-speech tagging • Coreference resolution.
A Java suite of tools that can perform natural • Information extraction • Tokenizer
GATE language processing tasks for multiple • Part of speech tagging • Sentence splitter
languages. Link
A suite of libraries and programs for symbolic • Information extraction • Word categorization
NLTK and statistical natural language processing • Part of speech tagging, • Text classification
Python. Link • Tokenizer
Statistical NLP toolkits for various • Including tokenization • Classification
computational linguistics problems that can be • Part-of-speech tagging • Segmentation
Stanford NLP incorporated into applications with human • Named entity • Coreference
language technology needs. Link recognition Resolution
• Parsing
A tool kit for processing text using • Sentiment analysis • Part of speech tagging
computational linguistics. Link • Entity recognition • Sentence detection
LingPipe
• Clustering • Disambiguation
• Topic classification
A suite of libraries and programs for symbolic • Information extraction • Text generation
and statistical natural language processing for • Part of speech tagging • Stemming
MontyLingua
both Python and Java. Link • Tokenizer • Phrase chunking
• Word categorization
A suite of linguistic analysis components that • Language • name matching
Rosetta integrate into applications for mining Identification • name translation
Linguistic unstructured data. Link • Name, places, and key
PwCPlatform concept extraction 46
Text Mining/Analytics Tools
Tool kits that provide capabilities for identifying and analyzing
features within individual or groups of texts

Tool Description Analysis Type


An open source environment for machine • Document • Data mining
learning, data mining, text mining, predictive classification • Traditional analytics
RapidMiner
analytics, and business analytics. Link • Sentiment analysis
• Topic tracking
A suite of text processing and analysis tools. • Text Parsing • Feature Extraction
SAS Text Miner
Link, • Filtering • Topic Clustering
Integrated development environment for • Information extractions • Data Mining
building information extraction systems, natural • Summarization • Document Filtering
VisualText
language processing systems, and text • Categorization • Natural Language
analyzers. Link Search
SAS Sentiment Commercial tool that is dedicated to customer • Customer sentiment • sentiment discovery
Analysis sentiment analysis. Link monitoring
Tool for sorting large amounts of unstructured • Topic modeling, • Document analysis
Textifier text with The Public Comment Analysis Toolkit • Information retrieval • Social media analysis
(PCAT). Link
System for automatically preparing and • Term frequency • Customization of stop
transforming unstructured text attributes into a • Term frequency inverse words
Infinite
structured representation. Link • Document frequency • Stemming rules
Insight
• Root word coding • Concepts merging
• synonym identification
Software for grouping related documents into • Document clustering
Clustify clusters, providing an overview of the document
set and aiding with categorization. Link

PwC 47
Text Mining/Analytics Tools Cont.
Tool kits that provide capabilities for identifying and analyzing
features within individual or groups of texts

Tool Description Analysis Type


Customer analytics applications that help • Unstructured • consumer profiling
Attensity analyze high volumes of customer conversations communication
Analyze across multiple channels. Link analysis
• sentiment analysis

A program that automatically identifies and • Information extraction • Topic Linking


ReVerb extracts binary relationships from English • Topic Identification
sentences. Link
Open text Open source tool for summarizing texts. Link • Document
summarizer summarization
Web based API that is used to analyze content • Attribute/feature • Fact identification
Open Calais
and extract topics or information. Link extraction
Knowledge Family of techniques tools for searching and • Semantic Analysis
Search organizing large data collections. Link
A free software for Quantitative Content • Text Parsing • Network analysis
KH Coder
Analysis or Text Mining Link • document search

PwC 48
Resources
Tutorials, Tools, Applications, and Research Groups

Link and Description


Text Mining Overview
Text Mining Activites

Text Mining Tutorial


Tutorials and Text Mining Process
Overviews NLP Introduction
NLP Overview
NLP Overview
NLP Concepts
[Link]

Research Groups [Link]


and Papers [Link]
[Link]
NLP Toolkit List
NLP Tools
Tools and Data Sets Text Mining Tools
Tools by Function
Text Mining Tools

PwC 49
DeepQA, Image
Analytics, and
Audio Analytics
DeepQA
Overview and Introduction

What is DeepQA?
• DeepQA forms that core of Watson, the
open domain question analysis and
answering system
• The DeepQA stack is comprised of set of
search, NLP, learning, and scoring
algorithms
• DeepQA operates on a distributed
computing infrastructure that leverages
Map Reduce and the Unstructured
What is the target problem set?
Information Management Architecture
• Understanding the meaning and context
of human language
• Searching and retrieving information
from large library of unstructured
information
• Identifying accurate and precise answers
to questions that are complex and must
sourced from a large knowledge set

PwC 51
DeepQA Infrastructure Technology
Data Management and Search

Technology Links
Unstructured
Information UIMA Link
Architecture
MySQL Link
SQL Server
Apache Derby Link
Java Natural Open NLP Link
Language
Toolkit Stanford NLP Link

Map/Reduce Apache Hadoop Link


Commonsense OpenCYC Link
Knowledgebas
e Open Mind Common Sense Link

Apache Jena Link


Triple Store
OpenAnzo Link
Lucene Link
Text Search
Open FTS Link

PwC 52
DeepQA Infrastructure Technology
Platform and Administration

Technology Links

Web Server Apache Link

VMWare Link
Virtualization
Host
Zen Link

Apache Hadoop Link


Distributed
File System
OpenAFS Link

File
Management/ rSync Link
Archival

OS Fedora Link

Extreme Cloud Administration Link


Cloud
Management
Open Nebula Link

PwC 53
Business Applications
DeepQA provides capabilities that can facilitate knowledge
discovery, improve customer interaction, and uncover hidden facts
Overview Objectives
• Identify information about a subject through deep
Search internal and external analysis of internal and external information
Knowledge unstructured/structured sources
Discovery information assets to uncover • Answer questions about a business problem or
previously unknown knowledge trend that may be difficult to analyze within
traditional data sources
Search documents and • Identify business topics and trends within
communications to uncover communication and documents
E-Discovery
relevant information associated • Search for non compliance activities within internal
with a specific topic and external data sources
Search through single or • Identify key facts or issues that comprise a contract
Contract multiple contracts to answer or sets of contracts
Evaluations specific questions about the • Identify contracts or legal documents that contain
nature of the contract similar entities or features
Provide the ability to interact • Provide a platform for automatically answering
Relationship with consumers providing consumer questions about products or services
Management precise responses to technical • Reduce reliance on call centers and improve
and open domain questions interaction with consumers
Search consumer • Identify background information about consumers
Consumer communications, social media, • Identify consumer qualities that create risks or
Discovery and sales information to identify represent opportunities
opportunities and demographics
Technical • Utilize unstructured data and communications to
Find answers to technical and identify solutions or root causes to system and
Troubleshootin
PwC process problems through process problems 54
g
Areas for Further Research
Infrastructure/Tools and Search Technologies/Concepts

Topic Research
The tool is used to distribute queries, analysis, and other processing activities
Hadoop across multiple CPUs. Further research is required to understand the tools
Map/Reduce architecture and how to integrate it with other tool kits. OpenNLP, UIMA,
Lucene, etc.
A Java library for NLP tasks. Need to evaluate the tools capabilities and gaps
OpenNLP
as well as how it can be incorporated into the UIMA
Tools An open common sense reasoning platform. Need to better understand the
OpenCYC
tools role as well as how it fits within the other technologies
An architecture for managing unstructured data. Further research is needed
UIMA to understand how to run in parallel and how the SDK can be applied to NLP
activities
A text search platform. Further research is needed to understand the library
Lucene
and how to incorporate it into UIMA
Algorithms are used to score search results based on their alignment with the
Text Search
question. Further research is needed to understand what models and scoring
Scoring
metrics can be applied to search results at various phases of DeepQA.
Triple stores maintain data in a subject-predicate-object structure and is used
Triple Store
for turning around quick facts. Further research is needed to understand the
Search
Search philosophy and technologies behind these data storage mechanisms
Commonsens Research is required to understand the branch of AI, technologies and role
e Reasoning within DeepQA.
Document/ Generate research on information and document retrieval practices.
Information Technologies and algorithms need to be reviewed. Falls within a broader
Retrieval research topic for enterprise search.
PwC 55
Areas for Further Research
Machine Learning and Natural Language Processing
Topic Description
Research the concept and how they are to used evaluate learning models
MetaLearners and assign a confidence score based on the learning models that are used to
rank search results
Machine Question Identify techniques and models that can be employed to analyze and classify
Learning Classification questions
Search Research models are available for ranking search results based on the
Ranking various search and recall techniques that are employed for a question
Models
Logical Form Research how SNA is used to discover logical relationships within text and
Analysis product an understanding about the information within the text
Semantic Identify tools and algorithms that are employed to uncover semantic
Structure relationships within texts/phrases and how these relationships can be
Analysis applied to extract relevant information for question analysis and search
NLP Relationship Research techniques and tools for uncovering temporal, geospatial and
Analysis spatial relationships within a knowledge set
Feature Evaluate tools and algorithms that are used to extract features of entities
Extraction from text and identify methods for structuring the data for search
Phrase Identify algorithms and tools that can be applied to extract key phrases from
Analysis text based on a search context

PwC 56
URLs
Overviews and Applications

Links
The AI Behind Watson
How to build a Watson Jr.

Background Building your own Watson


Documents Algorithms behind Watson
Overview of the technology behind Watson
DeepQA Project Page
Watson and your business

Applications and Understanding the DeepQA Process


Articles The future of DeepQA
DeepQA for e-discovery

PwC 57
Image Analytics Overview
How can we extract insight from images and video?

Overview
• The process of pulling relevant
information from an image or sets of
images for advanced classification and
traditional analysis
• Applies image capture, image
processing, and machine learning
techniques to extract, quantify, and
structure, image information
Advantages
• Provides a method to structure,
organize, and search information that is
stored within images
• Offers an additional data set that can be
applied to understanding consumer
behavior, automating business
processes, and discovering knowledge
enterprise content
58

PwC
Image Analytics Tools
There are few standalone packages that are capable of performing
robust image analysis; however, solutions can be developed using
existing frameworks and analytics toolkits
Image Computer Machine
Tool Overview
Processing Vision Learning
Open source library of computer
OpenCV vision functions that is accessible X X X
via C, Java, and Python
Integrated image analysis
PAXit
Image Analy
platform that provides basic X X
sis feature identification functions
Java based image processing
platform that can be accessed via
ImageJ
an API and expanded with custom X
plugins

PIL Python image processing library X


A modular machine learning
PyBrain
library for Python X

PwC 59
URLs
Tutorials, Tools, Applications, and Research Groups

Link and Description


Tutorial on Image Processing and Analysis

Tutorials Online Book of Algorithms for Computer Vision

Online Machine Vision Book


Computer Vision Group
Research Groups CMU Machine Vision Group
and Papers
Stanford Machine Vision Group
Image Analysis and Mining Framework
Tools and Data Sets
Image Mining Software

PwC 60
Audio Analytics Overview
How can we extract insight from audio and voice media?

Overview
• The process of capturing audio and analyzing
its features as to extract content and context
of an event
• Applies speech analysis and signal
processing principles to structure audio
information for analysis via NLP or traditional
analytics techniques
Advantages
• Provides a method for identifying events or
common patterns within sound bytes
• Offers a way of capturing not only the content
and topics within a conversation, but also the
emotions and context

61

PwC
Audio Analytics: Capabilities and Insights
What data can we capture from sound bites that can be used to
enhance other data or analysis?

Audio Event Information Points


Loudness/Intensity

• Event – audio events are identified as

Power and Intensity of


changes in sound patterns and or the
intensity over time

Sound
• Rate – defines how quickly a sound or a
pattern of sound is occurring and can
be used to evaluate the nature of an
exchange, the state of the sound
Rate
Time source, and context of the topic
• Power and Intensity – measures the
loudness of the sound or event and
Sound or Pitch
provides a way of evaluating the mood
Frequency

or emotion of the sound source


• Sound and Pitch – a measure of the
sound quality and can serve as a tool
for isolating separate audio events or
Time
sources as well as measuring changes
to the sound source
PwC 62
Audio Analytics Applications

Analysis Objectives
• Capture and structure the content of
conversations
Analyze conversations to • Utilize structured speech as an input to text
Voice
capture speech as text based mining and natural language processing
Recognition dialog capabilities
• Combine phone based conversations with other
interaction data sets

• Monitor customer interactions or business


Analyze sound clips to identify operations to capture events in real time
Sound Matching
specific events taking place • Use captured events for comparison,
categorization and analysis with other data points

• Capture the content of the conversation and


Monitor phone calls with conducting sentiment analysis based on word
Sentiment customers to uncover sentiment choice
Analysis towards the experience and/or • Analyze the pitch, loudness, and rate of consumer
products/services speech to identify emotional state during the
conversation and its cause
Monitor customer and job • Analyze prescreen phone conversations to assess job
candidate conversations to candidate personality, interest in job, and fit to job
Employee
extract information from word requirements
/Customer
usage and speech patterns that • Analyze customer conversations to assess level of risk
Screening can inform or improve a and honestly when applying for a product or filing
screening process claim/complaints

PwC 63
Audio Analytics Tools
There are few tools on the market that provide a broad range of
audio analysis capabilities. However, basic audio analysis and
natural language tool kits can be combined for robust analytics
Information
Tool Overview Audio Processing
Retrieval
A C++ library that provides varying level
Clam of audio processing and information X X
retrieval capabilities
A tool that is capable of translating calls
to a more structured text data set and
CallMiner
combining with other communication X
forms
Logs calls and structures audio for text
Nuance
based search and retrieval X
Aduio feature extraction toolkit with
yaafe
wrappers for several languages X
PRAAT Multiple platform audio analysis toolkit X

PwC 64
URLs
Tutorials, Tools, Applications, and Research Groups

Link and Description


Overview of audio features for seniment analysis

Tutorials Lecture on Audio Features and Information


Overview of audio analysis

Research Groups National Center for Voice and Speech

and Papers
Tools and Data Audio analysis package

Sets Audio Mining Software

PwC 65
Social Network
Analysis
Applications
Analyze organizational structures to identify opportunities that can
improve communication, productivity, and collaborations
Analysis Objectives
Evaluate team structures , • Identify team structures that are not effective
information flows among team
Collaboration • Identify informal organizational structures
members, and information
Analysis exchanges with other teams to • Identify individuals/roles or groups that are
improve working structures influential to collaborative work environments

Evaluate how knowledge or • Improve content and knowledge distribution


Content/
content is diffused and • Identify content bottlenecks, open
Knowledge
accessed within an communication flows, and establish channels
Management organization • Explore impact of new communication methods
• Improved structures for key organizational
Identify groups or informal functions.
teams that share knowledge, • Improved information flows
Community
communicate frequently, solve • Identify potential bottlenecks for organizational
Mining problems, or work together to functions
perform specific tasks • Identify cultural patterns to build other
communities
Explore formal and informal • Improve hierarchy and structure of organization
organization structures and to better align with the informal practices
Organization
how individuals work with one
Development another to improve the design • Identify team members that are effective leaders
of the organization and would impact the organization if promoted

PwC 67
Applications
Analyze network structures, communication channels, and
information flows to identify operational enhancements
Analysis Objectives
Assess organizational • Identify communication improvements to disaster
structures and communication recovery teams
Disaster recovery patterns as they relate to the • Identify weak links among functional groups to
planning groups that play a role in improve collaboration during recovery plan
disaster recovery plans execution
Assess how data points or • Identify overlapping information sets and
Data/ information sets originate or bottlenecks for information dissemination
Information are distributed across the • Assess how organization structures or information
Dissemination enterprise to their intended architecture impact the flow of information to its
targets targets
Assess the organization or • Identify network agents that collaborate with
external network to identify known fraudulent agents
Fraud
communication or • Identify activities that align with known fraudulent
Detection /
collaboration patterns that behavior
prevention align with known fraudulent
activity
Analyze the organization • Identify process improvements through discovery
structure and communication of hidden process steps, communication flows , and
Process patterns to uncover process actors
Discovery / improvements or identify new • Discover undocumented or informal processes that
Improvement processes are hidden within frequent collaboration and
communication paths
Evaluate the structure of a • Identify communication gaps that could impact
supply network and the dependent process or operations
PwC Supply Chain interactions among the entities • Identify strategic relationships to optimize the 68
Analysis that comprise the network to supply network
Applications
Analyze social media networks and consumer feedback to improve
product offerings and market interactions
Analysis Objectives
Observe how a specific topic, • Assess how target consumers/market will react to
Novelty/
news articles or sentiment a piece of news or campaign
Sentiment diffuses through a consumer • Evaluate how long news, data, or sentiment will
Diffusion network be retained within a system and how far it will
Analysis spread
Monitor and analyze • Identify individuals or groups that influence
Market connections within social markets and adoption
Influencer media networks to identify • Identify untapped markets
Identification markets or consumers that are
• Identify market segments as targets for ad
influential within communities
campaigns to improve product/service adoption
Analyze the connections and • Improve product or service offerings based on
consumer attributes within the attributes that connect the consumer market
Consumer target market to discover • Develop strategies to target new or existing
Segmentation communities or groups with consumers based on identified segmentation
common characteristics characteristics
Analyze the flow of • Identify segments or individuals that will be likely
Product or communication or ideas early adopters
Brand Diffusion through a market segment to • Identify incentives or campaigns that will improve
Analysis evaluate how a product may product/service adoption
diffuse
Analyze consumer network • Identify new feature sets for products and services
connections and common • Assess new markets for selling similar or new
Recommendatio features among consumers to products
n Systems develop recommendations
PwC • Target consumers with specific products or 69
services
Tools
Social network analysis plug-ins and APIs for development/scripting
languages and data analysis tools

Network Network Network


Tool Overview
Analysis Visual Manipulation
A general purpose network analysis and
SNAP
graph mining library for C++ . Link X X
A package for R that provides capabilities for
Statnet
social network statistical analysis. Link X
libSNA, Python libraries for network analysis and
graphTool, manipulation. libSNA, networkX, graphTool X X
networkX
Java package for network analysis and
JUNG
modeling. Link X X X
Excel plug-in that provides an easy to use
NodeXL and interactive interface to explore and X X
visualize networks Link

PwC 70
Tools
Proprietary and open source social network analysis interactive
application suites

Network Network Network


Tool Overview
Analysis Visual Manipulation
Interactive open source platform for network
GEPHI
analysis and visualization. Gephi X X X
Commercial social network analysis tool with
Ucinet
separate visualization component. Link X X
Open source graph visualization package.
Graphviz
Link X
Proprietary package that provides the ability
NetMiner to develop and implement custom algorithms X X X
link
Network analysis package that provides
kxen SNA predictive analytics and customer MDM X X X
integration. Link
Open source package for mining business
ProM
process networks. Link X X X
Open source tool for network modeling, and
Cytoscape analysis. Can connect to external data X X X
sources Link
Large-Scale Network Analysis, Modeling and
Network
Workbench
Visualization Toolkit for Biomedical, Social X X X
Science and Physics Research. Link

PwC 71
Resources
Tutorials, Tools, Applications, and Research Groups

Link and Description


Introduction for Beginners

Introductory Lecture

Paper on Business Applications


Tutorials
Network Analysis Process
Online Introductory Book
Introduction to Network Analysis Application and Theory (Open source book)
SNA Group at Stanford: Tool, Lectures, and Papers
Complex Networks and Systems Reasearch Collaboration
Research Groups
SNA Group and Indiana University: Lectures, Papers, and Tools
and Papers
Reality Mining at MIT
Papers from International Conference on Advances in Social Networks Analysis and Minin
g
Wiki List of Social Network Analysis Software
Review of 100+ Social Network Analysis Tools
List of Tools from The SAGE Handbook of Social Network Analysis
Tools and Data Sets More Tool Reviews
Twitter Data Sets
Web/Blog Data Sets
Facebook Data Sets

PwC 72
Additional Case
Studies
Example 1

Advanced natural language processing and deep


question-answering technology are being applied to
address clinical decision-making
Memorial Sloan-Kettering Cancer Center

• Memorial Sloan-Kettering Cancer Center is applying DeepQA technology


(technology that relies on advanced analytics powered by IBM’s Watson)
to develop a decision-support application for cancer treatment
• Doctors will be able to generate and evaluate hypothesis on evidence
and treatment and the Cancer Center will be able to better identify and
personalize cancer therapies for individual patients

WellPoint and Cedars-Sinai


• WellPoint and the Cedars-Sinai Samuel Oschin Comprehensive Cancer Institute
will work together to help improve patient care and support physicians in their
efforts to make the most informed, personalized treatment decisions possible.
• It is estimated that new clinical research and medical information doubles every
five years, and nowhere is this knowledge advancing more quickly than in the
complex area of cancer care.
• The WellPoint health care solutions will use DeepQA technology to draw from
vast libraries of information including medical evidence-based scientific and
health care data, and clinical insights from institutions like Cedars-Sinai.
Source: Memorial Sloan-Kettering Cancer Institute Press Release March 2012, WellPoint Press Release, December 2011;
PwC 74
Example 2

Large volumes of real-time sensor data are empowering


individuals to take more control of their health

Quantified Health – P4 Medicine (Predictive, Preventive, Personalized,


Participatory)

• Non-invasive wearable sensors are creating a new


‘Quantified Health’ movement and one of the fastest
growing sectors in the tech industry, let alone in the
field of Big Data Analytics
• The number of connected industrial and medical
devices is projected to reach 16 billion by 2015
• The mHealth market is estimated to reach a value of
$23 billion by 2017

Source: Bruce Bigelow, Big Data, Big Biology, and the ‘Tipping Point’ in Quantified Health: Takeaways from Xconomy’s On-the-Record Dinner,
Xconomy, April 26, 2012

PwC 75
Example 3

Advanced machine learning and visualization


techniques are being used to model drug interactions

Modeling Adverse Drug Reactions

When biological and phenotypic features were integrated alongside chemical


structures to predict adverse drug reactions, prediction accuracy increased
from 0.9054 to 0.9524.

Source: “Liu M, Wu Y, Chen Y, et al . Large-scale prediction of adverse drug reactions by integrating chemical, biological, and phenotypic
properties of drugs. J Am Med Inform Assoc 2012;19:e28–35.

PwC 76
Other Examples

Companies in other sectors are also pursuing various


applications of ‘Big Data’ and ‘Smart Analytics’.

Satellite Data
Hartford Steam Boiler
Allianz Locatio
n Hartford Steam Boiler is using
Allianz is ‘mashing’ satellite sensors and real-time sensor data
data, third-party street-level to monitor assets, reduce losses
data, images, and other and manage risks better
internal data to better
understand risk Map Data Property-Specific Hartford Steam Boiler has been
concentrations and manage Data able to manage concentration
concentration risk in risks and reduce losses, having
commercial property one of the lowest combined ratios
insurance for a commercial insurer

Proctor & Gamble

Proctor & Gamble is investing in analytics talent for quicker


decision making, with the CIO planning to increase fourfold the
number of staff with expertise in business analytics

Executives are currently using big data to uncover what is currently going on in
their business, to understand why, to predict future performance and to understand
what actions P&G should take
Source: “Procter & Gamble – Business Sphere and Decision Cockpits”, Ravi Kalakota, Pratical Analytics Wordpress, Feb. 2012,
[Link]/cancer-care; [Link], Healthcare IT News, IBM Watson to Aid Sloan-Kettering With Cancer Research, March
2012
PwC 77
Big Data
Analytics
Technology &
Vendor
Mappings
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Private EMC Private EMC
Cloud

HP Private Cloud HP

Teradata Private Teradata


Cloud
Dell Private Dell
Cloud
Public Azure SQL Microsoft
1.
Infrastructur Cloud
e Amazon Web Amazon
Services
Google Cloud Google
Platform
Hybrid EMC Hybrid EMC
Cloud

HP Helion HP

IBM Hybrid IBM


Cloud

PwC 79
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Batch/Micro Apache Kafka Apache Software
Foundation

Fluentd Open Source

Sqoop Apache Software


Foundation
Rabbit MQ Rabbit MQ

3. AWS Kinesis Amazon Web Services


Data
Data
Acquisitio
Ingestion &
n
Integration Apache Spark Apache Software
Foundation
Real time/ Apache Storm Apache Software
Streaming Foundation
Apache Spark Apache Software
Streaming Foundation
Samza Apache Software
Foundation
NiFi Apache Software
Foundation

PwC 80
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Data *Need assistance
Profiling/Cl in locating
e-ansing
Data *Need assistance
Data Matching/D in locating
Quality -uplication
Standardiza *Need assistance
ti-on/ in locating
Normaliz-
ation
ETL/ELT Hadoop Apache Hadoop
3.
Data Talend Talend
Ingestion &
Hive Apache Software
Integration
Foundation
Drill Apache Software
Data Foundation
Integratio Staging *Need assistance
n in locating
Persistent *Need assistance
Staging in locating
File *Need assistance
Exchange in locating
File Storage *Need assistance
in locating
PwC 81
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Custom *Need assistance
Compliers in locating
Batch MapReduce Apache Hadoop
MapReduce

Execution Spark Apache Software


/Data Foundation
Processin AWS EMR Amazon Web Services
g
Tez Apache Software
Foundation
3.5.
In-Memory Spark Apache Software
Execution/
Processing Foundation
Data
Processing Computing *Need assistance
Framework in locating
Cluster YARN Apache Hadoop
Managemen
t
Resource Mesos Apache Software
Managem Foundation
-ent
Zookeeper Apache Software
Foundation

Oozie Apache Software


Foundation

PwC 82
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Workflow Hue Open Source
Managemen
3.5. t Ambari Apache Software
Resource Foundation
Execution/
Managem
Data Lipstick Netflix
-ent
Processing
Ganglia The Ganglia Project

Traditional SQL Server Microsoft


Database
Oracle 10g Oracle

Parallel Teradata Teradata


Database
Relational
Data HP Vertica HP
Database
Appliances
IBM BigInsights IBM
4. EMC Greenplum EMC
Data
Repositories NewSQL ClustrixDB Clustrix
Mem SQL Memsql

Hadoop HDFS Apache Hadoop


Distribute DFS
d File AWS Amazon Web Services
System Packaged
Solutions Tachyon Tachyon Project
Operational
ODS
Data Store
PwC 83
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Relational/ MySQL Open Source
NewSQL
PostgreSQL Open Source

AWS RDS Amazon Web Services

In- Columnar Cassandra Apache Software


Memory DB Foundation
4.
Data Hbase Apache Hadoop
Repositories
AWS Redshift Amazon Web Services

NoSQL Hazelcast Open Source


Aerospike Aerospike
*Need
Metadata
assistance
Storage
in locating

PwC 84
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Key Value Redis Open Source

Riak Basho
AWS DynamoDB Amazon Web Services

Column Cassandra Apache Software


Store Foundation
Hbase Apache Hadoop
AWS Redshift Amazon Web Services
4.
Data
NoSQL Graph Neo4j Neo Technology
Repositories
Database
OrientDB Orient Tehcnologies

ArangoDB Open Source

Document MongoDB MongoDB, Inc.


Database
Elastic Elastic

Couchbase Couchbase

PwC 85
Big Data Analytics – Technology & Vendor Mappings

Layer L1 L2 Technology Vendor Logos


Reporting *Need Microstrategy
& assistance
Dashboar- in locating Datameer
ds
Visualizat *Need Qlik Sense Qlick
i-on tools/ assistance
Interactiv in locating
e Visual Tableau Tableau
Analytics
*Need *Need assistance
Real-time assistance in locating
6. Alerts in locating
Presentation/
Data
*Need D3 Open Source
Visualization
assistance
in locating Angular JS Google

Website Flask Open Source


Front-end
Highcharts Highcharts
Django Django Software
Foundation
*Need *Need assistance
API assistance in locating
in locating

PwC 86

You might also like