0% found this document useful (0 votes)
11 views47 pages

Understanding Digital Data Types

The document provides an overview of digital data types, categorizing them into structured, semi-structured, and unstructured data, along with their characteristics and sources. It discusses the significance of big data, its definition, challenges, and the differences between traditional business intelligence and big data environments. Additionally, it highlights the importance of information lifecycle management and the benefits of implementing it in data management.

Uploaded by

Raj Shah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views47 pages

Understanding Digital Data Types

The document provides an overview of digital data types, categorizing them into structured, semi-structured, and unstructured data, along with their characteristics and sources. It discusses the significance of big data, its definition, challenges, and the differences between traditional business intelligence and big data environments. Additionally, it highlights the importance of information lifecycle management and the benefits of implementing it in data management.

Uploaded by

Raj Shah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Chapter 1

Types of Digital Data

Big Data and Analytics by and


Copyright 2015, WILEY INDIA PVT. LTD.
Learning Objectives and Learning Outcomes

Learning Objectives Learning Outcomes


Introduction to digital data
and its types

1. Structured data: Sources of a) To differentiate between


structured data, ease with structured, semi-
structured data, etc. structured and
unstructured data.
2. Semi-Structured data:
Sources of semi-structured b) To understand the need
data, characteristics of to integrate structured,
semi-structured data. semi-structured and
unstructured data.
3. Unstructured data:
Sources of unstructured
data, issues with
terminology, dealing with
unstructured data.
Agenda

Types of Digital Data


 Structured
 Sources of structured data
 Ease with structured data

 Semi-Structured
 Sources of semi-structured data

 Unstructured
 Sources of unstructured data
 Issues with terminology
 Dealing with unstructured data
Classification of Digital Data

Digital data is classified into the following categories:

 Structured data

 Semi-structured data

 Unstructured data
Approximate Percentage Distribution of Digital Data

Approximate percentage distribution of digital data


Structured Data
Structured Data

 This is the data which is in an organized form (e.g., in rows and


columns) and can be easily used by a computer program.

 Relationships exist between entities of data, such as classes and


their objects.

 Data stored in databases is an example of structured data.


Sources of Structured Data

Databases such as
Oracle, DB2,
Teradata, MySql,
PostgreSQL, etc

Structured data Spreadsheets

OLTP Systems
Ease with Structured Data

Input / Update /
Delete

Security

Ease with Structured data Indexing /


Searching

Scalability

Transaction
Processing
Semi-structured
Data
Semi-structured Data

 This is the data which does not conform to a data model but has
some structure. However, it is not in a form which can be used easily
by a computer program.

 Example, emails, XML, markup languages like HTML, etc. Metadata


for this data is available but is not sufficient.
Sources of Semi-structured Data

XML (eXtensible Markup Language)

Other Markup Languages

JSON (Java Script Object Notation)


Semi-Structured Data
Characteristics of Semi-structured Data

Inconsistent Structure

Self-describing
(lable/value pairs)
Semi-structured data
Often Schema information is
blended with data values

Data objects may have different


attributes not known beforehand
Unstructured Data
Unstructured Data

 This is the data which does not conform to a data model or is not in a
form which can be used easily by a computer program.

 About 80–90% data of an organization is in this format.

 Example: memos, chat rooms, PowerPoint presentations, images,


videos, letters, researches, white papers, body of an email, etc.
Sources of Unstructured Data
Web Pages

Images

Free-Form
Text

Audios
Unstructured data

Videos

Body of
Email

Text
Messages

Chats

Social
Media data

Word
Document
Issues with terminology – Unstructured Data

Structure can be implied despite not being


formerly defined.

Data with some structure may still be labeled


Issues with terminology
unstructured if the structure doesn’t help with
processing task at hand

Data may have some structure or may even be


highly structured in ways that are unanticipated
or unannounced.
Dealing with Unstructured Data

Data Mining

Natural Language Processing (NLP)

Dealing with Unstructured Data Text Analytics

Noisy Text Analytics


Answer Me

 Which category (structured, semi-structured, or unstructured) will you


place a Web Page in?

 Which category (structured, semi-structured, or unstructured) will you


place Word Document in?

 State a few examples of human generated and machine-generated


data.
Next Agenda

 Definition of Big Data


 Volume
 Velocity
 Variety
 Challenges of Big Data
 Other Characteristics of Data Which are Not Definitional Traits of
Big Data
 Why Big Data?
 Traditional Business Intelligence (BI) versus Big Data
 A Typical Data Warehouse Environment
 A Typical Hadoop Environment
 Coexistence of Big Data and Data Warehouse
A Vocabulary for Measuring Information
If a Grain of Sand were One Byte of Information . . .

1 Megabyte =
1 million bytes
a tablespoon of sand
1 Gigabyte =
1 billion bytes
patch of sand—
9” square, 1’ deep
1 Terabyte =
1 trillion bytes
a sandbox—
24’ square, 1’ deep
1 Petabyte =
1,000 terabytes
a mile long beach—
100’ wide , 1’ deep
A New Vocabulary for Measuring Information
If a Grain of Sand were One Byte of Information . . .

1 Exabyte =
1 Megabyte = 1,000 petabytes
1 million bytes the same beach—
a tablespoon of sand from Maine to North Carolina
1 Gigabyte = 1 Zetabyte =
1 billion bytes 1,000 exabytes
patch of sand— the same beach—
9” square, 1’ deep along the entire US coast
1 Terabyte = 1 Yottabyte =
1 trillion bytes 1,000 zetabytes
a sandbox— enough info to bury the entire
24’ square, 1’ deep US under 296 feet of sand
1 Petabyte =
1,000 terabytes
a mile long beach—
100’ wide , 1’ deep
Characteristics of Big Data
Key Requirements for Data Center Elements
Variety
 Structured data: example: traditional transaction processing
systems and RDBMS, etc.

 Semi-structured data: example: Hyper Text Markup Language


(HTML), eXtensible Markup Language (XML).

 Unstructured data: example: unstructured text documents, audio,


video, email, photos, PDFs, social media, etc.

Velocity
Batch  Periodic  Near real time  Real-time processing
Veracity
 How accurate or truthful a data set may be.

 it’s not just the quality of the data itself but how trustworthy the data source,
type, and processing of it is.

 Removing things like bias, abnormalities or inconsistencies, duplication, and


volatility are just a few aspects that factor into improving the accuracy of big
data.

Volatility
 the rate of change and lifetime of the data.

 An example of
 highly volatile data includes social media, where sentiments and trending
topics change quickly and often.
 Less volatile data would look something more like weather trends that
change less frequently and are easier to predict and track
Information Lifecycle Management

Protect

New Process Deliver Warranty


order order order claim
Time
Value

Fulfilled Aged Warranty


order data Voided

Create Access Migrate Archive Dispose


Information Lifecycle Management Process

• Policy-based Alignment of Storage Infrastructure with Data Value


Benefits of Implementing ILM

o Improved utilization
o Tiered storage platforms
o Simplified management
o Processes, tools and automation
o Simplified backup and recovery
o A wider range of options to balance the need for business continuity
o Maintaining compliance
o Knowledge of what data needs to be protected for what length of time
o Lower Total Cost of Ownership
o By aligning the infrastructure and management costs with information
value
Definition of Big Data
Definition of Big Data

High-volume
Big Data is high-volume, high-
High-velocity
High-variety
velocity, and high-variety
information assets that
demand cost effective,
innovative forms of
information processing for
enhanced insight and decision
Cost-effective, innovative forms
of information processing making.
Source: Gartner IT
Glossary

Enhanced insight &


decision making
Challenges with Big Data
Challenges with Big Data
Capture

Storage

Curation

Challenges with Big Data


Search

Analysis

Transfer

Visualization

Privacy
Violations
Sources of Big Data
Sources of Big Data
Why Big Data?
Why Big Data?

More Data

More Accurate Analysis

More Confidence in decision making

Greater operational efficiencies, Cost reduction,


Time reduction, New product development, Optimized offerings, etc.
Traditional Business Intelligence (BI) versus
Big Data
A Typical Data Warehouse Environment

Reporting /
ERP
Dashboarding

CRM OLAP

Legacy Data Warehouse Ad hoc querying

3rd party Apps Modeling


A Typical Hadoop Environment

Web Logs HDFS

Hadoop
Operational
Systems
Images and Videos

Data Warehouse

Social Media
(Twitter, Facebook, etc.)
MapReduce
Data Marts

Docs & PDFs ODS


Co-existence of Big Data and Data Warehouse

Web Logs HDFS

Hadoop
Operational
Systems
Images and Videos

Data Warehouse
Data Warehouse
Social Media
(Twitter, Facebook, etc.)
MapReduce
Data Marts

Docs & PDFs ODS


Answer a few quick
questions …
Fill in the blanks

Big data is high-volume, high-velocity, and high-variety


information assets that demand ---------------------,
------------------------- forms of information processing for
enhanced ----------------------- and ---------
-------------------.
Answer Me

 Share your understanding of Big Data.

 How is traditional BI environment different from the Big Data


environment?

 Share your experience as a customer on an e-commerce site.


Comment on the big data that gets created on a typical e-commerce
site.
What do you Think ?

 If you were to send a 1.1 MB e-mail to 4 of your friends, by the time


your 4 friend receive that e-mail, What is your contribution to the
digital Universe ( how many Mb’s )

a) < 5 MB

b) 10 – 30 MB

c) 30 – 50 MB

d) > 50 MB
How Much Data do YOU Create?

START 1.1 MB SENT TO FOUR


Original MB
COLLEAGUES Document + 1.0
E-mail with Attachment
E-mail Text 0.1
Local E-mail Copy 1.1
E-mail with 2.2 MB E-mail Server 1.1
doc 1.1 MB Desktop Backup 1.0
Redundant Server 2.1
Document E-mail with
E-mail with 2.2 MB Tape Archive 4.2
doc 1.1 MB 9.5
1 MB doc 1.1 MB
Copies (4)
E-mail with E-mail Local Copies 4.4
doc 1.1 MB 2.2 MB
Server Copies 4.4
1.1
MB 2.1 MB Server Backup 4.4
1.0 E-mail with 2.2 MB Tape Archive 8.8
MB doc 1.1 MB 22.0
E-mail Server/ Redundant
2.1
Desktop Backup Backup
MB
Backup Transient Overhead20.0

4.2 MB 8.8 MB TOTAL 51.5


FINISH 51.5 MB!
Tape Back-up Tape Back-up
Thank you

You might also like