Chapter 1
Types of Digital Data
Big Data and Analytics by and
Copyright 2015, WILEY INDIA PVT. LTD.
Learning Objectives and Learning Outcomes
Learning Objectives Learning Outcomes
Introduction to digital data
and its types
1. Structured data: Sources of a) To differentiate between
structured data, ease with structured, semi-
structured data, etc. structured and
unstructured data.
2. Semi-Structured data:
Sources of semi-structured b) To understand the need
data, characteristics of to integrate structured,
semi-structured data. semi-structured and
unstructured data.
3. Unstructured data:
Sources of unstructured
data, issues with
terminology, dealing with
unstructured data.
Agenda
Types of Digital Data
Structured
Sources of structured data
Ease with structured data
Semi-Structured
Sources of semi-structured data
Unstructured
Sources of unstructured data
Issues with terminology
Dealing with unstructured data
Classification of Digital Data
Digital data is classified into the following categories:
Structured data
Semi-structured data
Unstructured data
Approximate Percentage Distribution of Digital Data
Approximate percentage distribution of digital data
Structured Data
Structured Data
This is the data which is in an organized form (e.g., in rows and
columns) and can be easily used by a computer program.
Relationships exist between entities of data, such as classes and
their objects.
Data stored in databases is an example of structured data.
Sources of Structured Data
Databases such as
Oracle, DB2,
Teradata, MySql,
PostgreSQL, etc
Structured data Spreadsheets
OLTP Systems
Ease with Structured Data
Input / Update /
Delete
Security
Ease with Structured data Indexing /
Searching
Scalability
Transaction
Processing
Semi-structured
Data
Semi-structured Data
This is the data which does not conform to a data model but has
some structure. However, it is not in a form which can be used easily
by a computer program.
Example, emails, XML, markup languages like HTML, etc. Metadata
for this data is available but is not sufficient.
Sources of Semi-structured Data
XML (eXtensible Markup Language)
Other Markup Languages
JSON (Java Script Object Notation)
Semi-Structured Data
Characteristics of Semi-structured Data
Inconsistent Structure
Self-describing
(lable/value pairs)
Semi-structured data
Often Schema information is
blended with data values
Data objects may have different
attributes not known beforehand
Unstructured Data
Unstructured Data
This is the data which does not conform to a data model or is not in a
form which can be used easily by a computer program.
About 80–90% data of an organization is in this format.
Example: memos, chat rooms, PowerPoint presentations, images,
videos, letters, researches, white papers, body of an email, etc.
Sources of Unstructured Data
Web Pages
Images
Free-Form
Text
Audios
Unstructured data
Videos
Body of
Email
Text
Messages
Chats
Social
Media data
Word
Document
Issues with terminology – Unstructured Data
Structure can be implied despite not being
formerly defined.
Data with some structure may still be labeled
Issues with terminology
unstructured if the structure doesn’t help with
processing task at hand
Data may have some structure or may even be
highly structured in ways that are unanticipated
or unannounced.
Dealing with Unstructured Data
Data Mining
Natural Language Processing (NLP)
Dealing with Unstructured Data Text Analytics
Noisy Text Analytics
Answer Me
Which category (structured, semi-structured, or unstructured) will you
place a Web Page in?
Which category (structured, semi-structured, or unstructured) will you
place Word Document in?
State a few examples of human generated and machine-generated
data.
Next Agenda
Definition of Big Data
Volume
Velocity
Variety
Challenges of Big Data
Other Characteristics of Data Which are Not Definitional Traits of
Big Data
Why Big Data?
Traditional Business Intelligence (BI) versus Big Data
A Typical Data Warehouse Environment
A Typical Hadoop Environment
Coexistence of Big Data and Data Warehouse
A Vocabulary for Measuring Information
If a Grain of Sand were One Byte of Information . . .
1 Megabyte =
1 million bytes
a tablespoon of sand
1 Gigabyte =
1 billion bytes
patch of sand—
9” square, 1’ deep
1 Terabyte =
1 trillion bytes
a sandbox—
24’ square, 1’ deep
1 Petabyte =
1,000 terabytes
a mile long beach—
100’ wide , 1’ deep
A New Vocabulary for Measuring Information
If a Grain of Sand were One Byte of Information . . .
1 Exabyte =
1 Megabyte = 1,000 petabytes
1 million bytes the same beach—
a tablespoon of sand from Maine to North Carolina
1 Gigabyte = 1 Zetabyte =
1 billion bytes 1,000 exabytes
patch of sand— the same beach—
9” square, 1’ deep along the entire US coast
1 Terabyte = 1 Yottabyte =
1 trillion bytes 1,000 zetabytes
a sandbox— enough info to bury the entire
24’ square, 1’ deep US under 296 feet of sand
1 Petabyte =
1,000 terabytes
a mile long beach—
100’ wide , 1’ deep
Characteristics of Big Data
Key Requirements for Data Center Elements
Variety
Structured data: example: traditional transaction processing
systems and RDBMS, etc.
Semi-structured data: example: Hyper Text Markup Language
(HTML), eXtensible Markup Language (XML).
Unstructured data: example: unstructured text documents, audio,
video, email, photos, PDFs, social media, etc.
Velocity
Batch Periodic Near real time Real-time processing
Veracity
How accurate or truthful a data set may be.
it’s not just the quality of the data itself but how trustworthy the data source,
type, and processing of it is.
Removing things like bias, abnormalities or inconsistencies, duplication, and
volatility are just a few aspects that factor into improving the accuracy of big
data.
Volatility
the rate of change and lifetime of the data.
An example of
highly volatile data includes social media, where sentiments and trending
topics change quickly and often.
Less volatile data would look something more like weather trends that
change less frequently and are easier to predict and track
Information Lifecycle Management
Protect
New Process Deliver Warranty
order order order claim
Time
Value
Fulfilled Aged Warranty
order data Voided
Create Access Migrate Archive Dispose
Information Lifecycle Management Process
• Policy-based Alignment of Storage Infrastructure with Data Value
Benefits of Implementing ILM
o Improved utilization
o Tiered storage platforms
o Simplified management
o Processes, tools and automation
o Simplified backup and recovery
o A wider range of options to balance the need for business continuity
o Maintaining compliance
o Knowledge of what data needs to be protected for what length of time
o Lower Total Cost of Ownership
o By aligning the infrastructure and management costs with information
value
Definition of Big Data
Definition of Big Data
High-volume
Big Data is high-volume, high-
High-velocity
High-variety
velocity, and high-variety
information assets that
demand cost effective,
innovative forms of
information processing for
enhanced insight and decision
Cost-effective, innovative forms
of information processing making.
Source: Gartner IT
Glossary
Enhanced insight &
decision making
Challenges with Big Data
Challenges with Big Data
Capture
Storage
Curation
Challenges with Big Data
Search
Analysis
Transfer
Visualization
Privacy
Violations
Sources of Big Data
Sources of Big Data
Why Big Data?
Why Big Data?
More Data
More Accurate Analysis
More Confidence in decision making
Greater operational efficiencies, Cost reduction,
Time reduction, New product development, Optimized offerings, etc.
Traditional Business Intelligence (BI) versus
Big Data
A Typical Data Warehouse Environment
Reporting /
ERP
Dashboarding
CRM OLAP
Legacy Data Warehouse Ad hoc querying
3rd party Apps Modeling
A Typical Hadoop Environment
Web Logs HDFS
Hadoop
Operational
Systems
Images and Videos
Data Warehouse
Social Media
(Twitter, Facebook, etc.)
MapReduce
Data Marts
Docs & PDFs ODS
Co-existence of Big Data and Data Warehouse
Web Logs HDFS
Hadoop
Operational
Systems
Images and Videos
Data Warehouse
Data Warehouse
Social Media
(Twitter, Facebook, etc.)
MapReduce
Data Marts
Docs & PDFs ODS
Answer a few quick
questions …
Fill in the blanks
Big data is high-volume, high-velocity, and high-variety
information assets that demand ---------------------,
------------------------- forms of information processing for
enhanced ----------------------- and ---------
-------------------.
Answer Me
Share your understanding of Big Data.
How is traditional BI environment different from the Big Data
environment?
Share your experience as a customer on an e-commerce site.
Comment on the big data that gets created on a typical e-commerce
site.
What do you Think ?
If you were to send a 1.1 MB e-mail to 4 of your friends, by the time
your 4 friend receive that e-mail, What is your contribution to the
digital Universe ( how many Mb’s )
a) < 5 MB
b) 10 – 30 MB
c) 30 – 50 MB
d) > 50 MB
How Much Data do YOU Create?
START 1.1 MB SENT TO FOUR
Original MB
COLLEAGUES Document + 1.0
E-mail with Attachment
E-mail Text 0.1
Local E-mail Copy 1.1
E-mail with 2.2 MB E-mail Server 1.1
doc 1.1 MB Desktop Backup 1.0
Redundant Server 2.1
Document E-mail with
E-mail with 2.2 MB Tape Archive 4.2
doc 1.1 MB 9.5
1 MB doc 1.1 MB
Copies (4)
E-mail with E-mail Local Copies 4.4
doc 1.1 MB 2.2 MB
Server Copies 4.4
1.1
MB 2.1 MB Server Backup 4.4
1.0 E-mail with 2.2 MB Tape Archive 8.8
MB doc 1.1 MB 22.0
E-mail Server/ Redundant
2.1
Desktop Backup Backup
MB
Backup Transient Overhead20.0
4.2 MB 8.8 MB TOTAL 51.5
FINISH 51.5 MB!
Tape Back-up Tape Back-up
Thank you