Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Lecture One
Introduction to Big Data
1. Introduction
Big Data is a field dedicated to the analysis, processing, and storage of
large collections of data that frequently originate from disparate sources.
Big Data solutions and practices are typically required when traditional
data analysis, processing and storage technologies and techniques are
insufficient. Specifically, Big Data addresses distinct requirements, such
as the combining of multiple unrelated datasets, processing of large
amounts of unstructured data and harvesting of hidden information in a
time-sensitive manner.
Although Big Data may appear as a new discipline, it has been developing
for years. The management and analysis of large datasets has been a long-
standing problem—from labor intensive approaches of early census efforts
to the actuarial science behind the calculations of insurance premiums. Big
Data science has evolved from these roots.
In addition to traditional analytic approaches based on statistics, Big Data
adds newer techniques that leverage computational resources and
approaches to execute analytic algorithms. This shift is important as
datasets continue to become larger, more diverse, more complex and
streaming-centric. While statistical approaches have been used to
approximate measures of a population via sampling since Biblical times,
advances in computational science have allowed the processing of entire
datasets, making such sampling unnecessary.
The analysis of Big Data datasets is an interdisciplinary endeavor that
blends mathematics, statistics, computer science and subject matter
expertise. This mixture of skillsets and perspectives has led to some
confusion as to what comprises the field of Big Data and its analysis, for
the response one receives will be dependent upon the perspective of
whoever is answering the question. The boundaries of what constitutes a
Big Data problem are also changing due to the ever-shifting and advancing
landscape of software and hardware technology. This is due to the fact that
the definition of Big Data takes into account the impact of the data’s
characteristics on the design of the solution environment itself. Thirty years
ago, one gigabyte of data could amount to a Big Data problem and require
special purpose computing resources. Now, gigabytes of data are
commonplace and can be easily transmitted, processed and stored on
consumer-oriented devices.
1
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Data within Big Data environments generally accumulates from being
amassed within the enterprise via applications, sensors and external
sources. Data processed by a Big Data solution can be used by enterprise
applications directly or can be fed into a data warehouse to enrich existing
data there. The results obtained through the processing of Big Data can
lead to a wide range of insights and benefits, such as:
• operational optimization
• actionable intelligence
• identification of new markets
• accurate predictions
• fault and fraud detection
• more detailed records
• improved decision-making
• scientific discoveries
Evidently, the applications and potential benefits of Big Data are broad.
However, there are numerous issues that need to be considered when
adopting Big Data analytics approaches. These issues need to be
understood and weighed against anticipated benefits so that informed
decisions and plans can be produced.
2. What is Data?
Data means the quantities, characters, or symbols on which operations are
performed by a computer, which may be stored and transmitted in the form
of electrical signals and recorded on magnetic, optical, or mechanical
recording media.
While, Dataset is the collections or groups of related data are generally
referred to as datasets. Each group or dataset member (datum) shares the
same set of attributes or properties as others in the same dataset. Some
examples of datasets are:
• tweets stored in a flat file
• a collection of image files in a directory
• an extract of rows from a database table stored in a CSV formatted file
• historical weather observations that are stored as XML files
Figure 1.1 shows three datasets based on three different data formats.
2
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
In fact there are two important terms should be describe briefly: Data
Analysis and Data Analytic:
2.1 Data analysis
It is the process of examining data to find facts, relationships, patterns,
insights and/or trends. The overall goal of data analysis is to support better
decision making. A simple data analysis example is the analysis of ice
cream sales data in order to determine how the number of ice cream cones
sold is related to the daily temperature. The results of such an analysis
would support decisions related to how much ice cream a store should
order in relation to weather forecast information. Carrying out data analysis
helps establish patterns and relationships among the data being analyzed.
Figure 1.2 shows the symbol used to represent data analysis.
2.2 Data Analytics
Data analytics is a broader term that encompasses data analysis. Data
analytics is a discipline that includes the management of the complete data
lifecycle, which encompasses collecting, cleansing, organizing, storing,
analyzing and governing data. The term includes the development of
analysis methods, scientific techniques and automated tools. In Big Data
3
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
environments, data analytics has developed methods that allow data
analysis to occur through the use of highly scalable distributed
technologies and frameworks that are capable of analyzing large volumes
of data from different sources. Figure 1.3 shows the symbol used to
represent analytics.
The Big Data analytics lifecycle generally involves identifying, procuring,
preparing and analyzing large amounts of raw, unstructured data to extract
meaningful information that can serve as an input for identifying patterns,
enriching existing enterprise data and performing large-scale searches.
Different kinds of organizations use data analytics tools and techniques in
different ways. For example, these three sectors:
• In business-oriented environments, data analytics results can lower
operational costs and facilitate strategic decision-making.
• In the scientific domain, data analytics can help identify the cause of a
phenomenon to improve the accuracy of predictions.
• In service-based environments like public sector organizations, data
analytics can help strengthen the focus on delivering high-quality
services by driving down costs.
Data analytics enable data-driven decision-making with scientific backing
so that decisions can be based on factual data and not simply on past
experience or intuition alone. There are four general categories of analytics
that are distinguished by the results they produce:
4
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
• Descriptive analytics
• Diagnostic analytics
• Predictive analytics
• Prescriptive analytics
The different analytics types leverage different techniques and analysis
algorithms. This implies that there may be varying data, storage and
processing requirements to facilitate the delivery of multiple types of
analytic results. Figure 1.4 depicts the reality that the generation of high
value analytic results increases the complexity and cost of the analytic
environment.
Figure 1.4 Value and complexity increase from descriptive to
prescriptive analytics.
2.2.1 Descriptive Analytics
Descriptive analytics are carried out to answer questions about events that
have already occurred. This form of analytics contextualizes data to
generate information. Sample questions can include:
• What was the sales volume over the past 12 months?
• What is the number of support calls received as categorized by severity
and geographic location?
• What is the monthly commission earned by each sales agent?
It is estimated that 80% of generated analytics results are descriptive in
nature. Value wise, descriptive analytics provide the least worth and
require a relatively basic skillset.
Descriptive analytics are often carried out via ad-hoc reporting or
dashboards, as shown in Figure 1.5. The reports are generally static in
5
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
nature and display historical data that is presented in the form of data grids
or charts. Queries are executed on operational data stores from within an
enterprise, for example a Customer Relationship Management system
(CRM) or Enterprise Resource Planning (ERP) system.
Figure 1.5 The operational systems, pictured left, are queried via descriptive
analytics tools to generate reports or dashboards, pictured right.
2.2.2 Diagnostic Analytics
Diagnostic analytics aim to determine the cause of a phenomenon that
occurred in the past using questions that focus on the reason behind the
event. The goal of this type of analytics is to determine what information
is related to the phenomenon in order to enable answering questions that
seek to determine why something has occurred. Such questions include:
• Why were Q2 sales less than Q1 sales?
• Why have there been more support calls originating from the Eastern
region than from the Western region?
• Why was there an increase in patient re-admission rates over the past
three months?
Diagnostic analytics provide more value than descriptive analytics but
require a more advanced skillset. Diagnostic analytics usually require
collecting data from multiple sources and storing it in a structure that lends
itself to performing drill-down and roll-up analysis, as shown in Figure 1.6.
Diagnostic analytics results are viewed via interactive visualization tools
that enable users to identify trends and patterns. The executed queries are
more complex compared to those of descriptive analytics and are
performed on multidimensional data held in analytic processing systems.
6
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Figure 1.6 Diagnostic analytics can result in data that is suitable for
performing drilldown and roll-up analysis.
2.2.3 Predictive Analytics
Predictive analytics are carried out in an attempt to determine the outcome
of an event that might occur in the future. With predictive analytics,
information is enhanced with meaning to generate knowledge that conveys
how that information is related. The strength and magnitude of the
associations form the basis of models that are used to generate future
predictions based upon past events. It is important to understand that the
models used for predictive analytics have implicit dependencies on the
conditions under which the past events occurred. If these underlying
conditions change, then the models that make predictions need to be
updated.
Questions are usually formulated using a what-if rationale, such as the
following:
• What are the chances that a customer will default on a loan if they have
missed monthly payment?
• What will be the patient survival rate if Drug B is administered instead
of Drug A?
• If a customer has purchased Products A and B, what are the chances
that they will also purchase Product C?
Predictive analytics try to predict the outcomes of events, and predictions
are made based on patterns, trends and exceptions found in historical and
current data. This can lead to the identification of both risks and
opportunities.
This kind of analytics involves the use of large datasets comprised of
internal and external data and various data analysis techniques. It provides
greater value and requires a more advanced skillset than both descriptive
and diagnostic analytics. The tools used generally abstract underlying
7
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
statistical intricacies by providing user-friendly front-end interfaces, as
shown in Figure 1.7.
Figure 1.7 Predictive analytics tools can provide user-friendly front-end interfaces.
2.3.4 Prescriptive Analytics
Prescriptive analytics build upon the results of predictive analytics by
prescribing actions that should be taken. The focus is not only on which
prescribed option is best to follow, but why. In other words, prescriptive
analytics provide results that can be reasoned about because they embed
elements of situational understanding. Thus, this kind of analytics can be
used to gain an advantage or mitigate a risk.
Sample questions may include:
• Among three drugs, which one provides the best results?
• When is the best time to trade a particular stock?
Prescriptive analytics provide more value than any other type of analytics
and correspondingly require the most advanced skillset, as well as
specialized software and tools. Various outcomes are calculated, and the
best course of action for each outcome is suggested. The approach shifts
from explanatory to advisory and can include the simulation of various
scenarios.
This sort of analytics incorporates internal data with external data. Internal
data might include current and historical sales data, customer information,
product data and business rules. External data may include social media
data, weather forecasts and government produced demographic data.
Prescriptive analytics involve the use of business rules and large amounts
of internal and external data to simulate outcomes and prescribe the best
course of action, as shown in Figure 1.8.
8
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Figure 1.8 Prescriptive analytics involves the use of business rules and internal and/or
external data to perform an in-depth analysis.
2.3.5 Business Intelligence (BI)
BI enables an organization to gain insight into the performance of an
enterprise by analyzing data generated by its business processes and
information systems. The results of the analysis can be used by
management to steer the business in an effort to correct detected issues or
otherwise enhance organizational performance. BI applies analytics to
large amounts of data across the enterprise, which has typically been
consolidated into an enterprise data warehouse to run analytical queries.
As shown in Figure 1.9, the output of BI can be surfaced to a dashboard
that allows managers to access and analyze the results and potentially
refine the analytic queries to further explore the data.
Figure 1.9 BI can be used to improve business applications, consolidate data in
data warehouses and analyze queries via a dashboard.
2.3.6 Key Performance Indicators (KPI)
A KPI is a metric that can be used to gauge success within a particular
business context. KPIs are linked with an enterprise’s overall strategic
goals and objectives. They are often used to identify business performance
problems and demonstrate regulatory compliance. KPIs therefore act as
quantifiable reference points for measuring a specific aspect of a business’
overall performance. KPIs are often displayed via a KPI dashboard, as
shown in Figure 1.10. The dashboard consolidates the display of multiple
9
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
KPIs and compares the actual measurements with threshold values that
define the acceptable value range of the KPI.
Figure 1.10 A KPI dashboard acts as a central reference point for gauging business
3. What is Big Data?
Big Data is a collection of data that is huge in volume, yet growing
exponentially with time. It is a data with so large size and complexity that
none of traditional data management tools can store it or process it
efficiently. Big data is also a data but with huge size. Moreover, Big Data
applies to information that can’t be processed or analyzed using traditional
processes or tools. Big data is not a new phenomenon, but one that is part
of a long evolution of data collection and analysis.
Big data is “the ability of society to harness information in novel ways to
produce useful insights or goods and services of significant value” and
“things one can do at a large scale that cannot be done at a smaller one, to
extract new insights or create new forms of value.” Big Data is the ocean
of information we swim in every day – vast zettabytes of data flowing from
our computers, mobile devices, and machine sensors. This data is used by
organizations to drive decisions, improve processes and policies, and
create customer-centric products, services, and experiences. Big Data is
defined as “big” not just because of its volume, but also due to the variety
and complexity of its nature. Typically, it exceeds the capacity of
traditional databases to capture, manage, and process it. And, Big Data can
come from anywhere or anything on earth that we’re able to monitor
digitally. Weather satellites, Internet of Things (IoT) devices, traffic
cameras, social media trends – these are just a few of the data sources being
mined and analyzed to make businesses more resilient and competitive.
Below are some the Big Data examples-
11
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
The New York Stock Exchange is an example of Big Data
that generates about one terabyte of new trade data per day.
Social Media, the statistic shows that 500+terabytes of new data get
ingested into the databases of social media site Facebook, every day.
This data is mainly generated in terms of photo and video uploads,
message exchanges, putting comments etc.
A single Jet engine can generate 10+terabytes of data in 30
minutes of flight time. With many thousand flights per day,
generation of data reaches up to many Petabytes.
11
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
4. Types of Big Data
The data processed by Big Data solutions can be human-generated or
machine-generated although it is ultimately the responsibility of machines
to generate the analytic results. Human-generated data is the result of
human interaction with systems, such as online services and digital
devices. Figure 1.16 shows examples of human-generated data.
Machine-generated data is generated by software programs and hardware
devices in response to real-world events. For example, a log file captures
an authorization decision made by a security service, and a point-of-sale
system generates a transaction against inventory to reflect items purchased
by a customer. From a hardware perspective, an example of machine-
generated data would be information conveyed from the numerous sensors
in a cellphone that may be reporting information, including position and
cell tower signal strength. Figure 1.17 provides a visual representation of
different types of machine generated data.
12
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
As demonstrated, human-generated and machine-generated data can come
from a variety of sources and be represented in various formats or types.
This section examines the variety of data types that are processed by Big
Data solutions. The primary types of data are:
• structured data
• unstructured data
• semi-structured data
These data types refer to the internal organization of data and are
sometimes called data formats. Apart from these three fundamental data
types, another important type of data in Big Data environments is metadata.
Each will be explored in turn.
13
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
4.1 Structured Data
Structured data conforms to a data model or schema and is often stored in
tabular form. It is used to capture relationships between different entities
and is therefore most often stored in a relational database. Structured data
is frequently generated by enterprise applications and information systems
like ERP and CRM systems. Due to the abundance of tools and databases
that natively support structured data, it rarely requires special consideration
in regards to processing or storage. Examples of this type of data include
banking transactions, invoices, and customer records. Figure 1.18 shows
the symbol used to represent structured data.
In other words, Structured data: This kind of data is the simplest to
organize and search. It can include things like financial data, machine logs,
and demographic details. An Excel spreadsheet, with its layout of pre-
defined columns and rows, is a good way to envision structured data. Its
components are easily categorized, allowing database designers and
administrators to define simple algorithms for search and analysis. Even
when structured data exists in enormous volume, it doesn’t necessarily
qualify as Big Data because structured data on its own is relatively simple
to manage and therefore doesn’t meet the defining criteria of Big Data.
Traditionally, databases have used a programming language called
Structured Query Language (SQL) in order to manage structured data. SQL
was developed by IBM in the 1970s to allow developers to build and
manage relational (spreadsheet style) databases that were beginning to take
off at that time.
Any data that can be stored, accessed and processed in the form of fixed
format is termed as a ‘structured’ data. Over the period of time, talent in
computer science has achieved greater success in developing techniques
for working with such kind of data (where the format is well known in
advance) and also deriving value out of it. However, nowadays, we are
foreseeing issues when a size of such data grows to a huge extent, typical
sizes are being in the rage of multiple zettabytes.
Examples Of Structured Data
An ‘Employee’ table in a database is an example of Structured Data
14
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
4.2 Unstructured Data
Data that does not conform to a data model or data schema is known as
unstructured data. It is estimated that unstructured data makes up 80% of
the data within any given enterprise. Unstructured data has a faster growth
rate than structured data. Figure 1.19 illustrates some common types of
unstructured data. This form of data is either textual or binary and often
conveyed via files that are self-contained and non-relational. A text file
may contain the contents of various tweets or blog postings. Binary files
are often media files that contain image, audio or video data. Technically,
both text and binary files have a structure defined by the file format itself,
but this aspect is disregarded, and the notion obeing unstructured is in
relation to the format of the data contained in the file itself.
Special purpose logic is usually required to process and store unstructured
data. For example, to play a video file, it is essential that the correct codec
(coder-decoder) is available. Unstructured data cannot be directly
processed or queried using SQL. If it is required to be stored within a
relational database, it is stored in a table as a Binary Large Object (BLOB).
Alternatively, a Not-only SQL (NoSQL) database is a non-relational
database that can be used to store unstructured data alongside structured
data.
15
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
In other words, Unstructured data: This category of data can include
things like social media posts, audio files, images, and open-ended
customer comments. This kind of data cannot be easily captured in
standard row-column relational databases. Traditionally, companies that
wanted to search, manage, or analyze large amounts of unstructured data
had to use laborious manual processes. There was never any question as to
the potential value of analyzing and understanding such data, but the cost
of doing so was often too exorbitant to make it worthwhile. Considering
the time it took, results were often obsolete before they were even
delivered. Instead of spreadsheets or relational databases, unstructured data
is usually stored in data lakes, data warehouses, and NoSQL databases.
Any data with unknown form or the structure is classified as unstructured
data. In addition to the size being huge, un-structured data poses multiple
challenges in terms of its processing for deriving value out of it. A typical
example of unstructured data is a heterogeneous data source containing a
combination of simple text files, images, videos etc. Now day
organizations have wealth of data available with them but unfortunately,
they don’t know how to derive value out of it since this data is in its raw
form or unstructured format.
Examples Of Un-structured Data
The output returned by ‘Google Search’
16
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
4.3 Semi-structured Data
Semi-structured data has a defined level of structure and consistency, but
is not relational in nature. Instead, semi-structured data is hierarchical or
graph-based. This kind of data is commonly stored in files that contain text.
For instance, Figure 1.20 shows that XML and JSON files are common
forms of semi-structured data. Due to the textual nature of this data and its
conformance to some level of structure, it is more easily processed than
unstructured data.
Examples of common sources of semi-structured data include electronic
data interchange (EDI) files, spreadsheets, RSS feeds and sensor data.
Semi-structured data often has special pre-processing and storage
requirements, especially if the underlying format is not text-based. An
example of pre-processing of semi-structured data would be the validation
of an XML file to ensure that it conformed to its schema definition.
Metadata
Metadata provides information about a dataset’s characteristics and
structure. This type of data is mostly machine-generated and can be
appended to data. The tracking of metadata is crucial to Big Data
processing, storage and analysis because it provides information about the
pedigree of the data and its provenance during processing. Examples of
metadata include:
• XML tags providing the author and creation date of a document
• attributes providing the file size and resolution of a digital photograph
Big Data solutions rely on metadata, particularly when processing semi-
structured and unstructured data. Figure 1.21 shows the symbol used to
represent metadata.
17
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
In other words, Semi-structured data: As it sounds, semi-structured data
is a hybrid of structured and unstructured data. E-mails are a good example
as they include unstructured data in the body of the message, as well as
more organizational properties such as sender, recipient, subject, and date.
Devices that use geo-tagging, time stamps, or semantic tags can also
deliver structured data alongside unstructured content. An unidentified
smartphone image, for instance, can still tell you that it is a selfie, and the
time and place where it was taken. A modern database running AI
technology can not only instantly identify different types of data, it can also
generate algorithms in real time to effectively manage and analyze the
disparate data sets involved.
Semi-structured data can contain both the forms of data. We can see semi-
structured data as a structured in form but it is actually not defined with
e.g. a table definition in relational DBMS. Example of semi-structured data
is a data represented in an XML file.
Examples Of Semi-structured Data Personal data stored in an XML file-
18
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Data Growth over the years
5. Big Data Characteristics
For a dataset to be considered Big Data, it must possess one or more
characteristics that require accommodation in the solution design and
architecture of the analytic environment. Most of these data characteristics
were initially identified by Doug Laney in early 2001 when he published
an article describing the impact of the volume, velocity and variety of e-
commerce data on enterprise data warehouses. To this list, veracity has
been added to account for the lower signal-to-noise ratio of unstructured
data as compared to structured data sources. Ultimately, the goal is to
conduct analysis of the data in such a manner that high-quality results are
delivered in a timely manner, which provides optimal value to the
enterprise.
The five main Big Data characteristics that can be used to help differentiate
data categorized as “Big” from other forms of data will explain. The five
Big Data traits shown in Figure 1.22 are commonly referred to as the Five
Vs:
• volume
• velocity
• variety
• veracity
• value
19
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Figure 1.22 The Five Vs of Big Data.
(i) Volume – The name Big Data itself is related to a size which is
enormous. Size of data plays a very crucial role in determining value out
of data. Also, whether a particular data can actually be considered as a Big
Data or not, is dependent upon the volume of data. Hence, ‘Volume’ is one
characteristic which needs to be considered while dealing with Big Data
solutions.
The anticipated volume of data that is processed by Big Data solutions is
substantial and ever-growing. High data volumes impose distinct data
storage and processing demands, as well as additional data preparation,
duration and management processes. Figure 1.23 provides a visual
representation of the large volume of data being created daily by
organizations and users world-wide.
Figure 1.23 Organizations and users world-wide create over 2.5 EBs of
data a day.
21
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
As a point of comparison, the Library of Congress currently holds more
than 300 TBs of data.
Typical data sources that are responsible for generating high data volumes
can include:
• online transactions, such as point-of-sale and banking
• scientific and research experiments, such as the Large Hadron Collider
and Atacama Large Millimeter/Submillimeter Array telescope
• sensors, such as GPS sensors, RFIDs, smart meters and telematics
• social media, such as Facebook and Twitter.
Volume: While volume is by no means the only component that makes
Big Data “big,” it is certainly a primary feature. To fully manage and utilize
Big Data, advanced algorithms and AI-driven analytics are required. But
before any of that can happen, there needs to be a secure and reliable means
of storing, organizing, and retrieving the many terabytes of data that are
held by large companies
(ii) Variety – Variety refers to heterogeneous sources and the nature of
data, both structured and unstructured. During earlier days, spreadsheets
and databases were the only sources of data considered by most of the
applications. Nowadays, data in the form of emails, photos, videos,
monitoring devices, PDFs, audio, etc. are also being considered in the
analysis applications. This variety of unstructured data poses certain issues
for storage, mining and analyzing data.
Data variety refers to the multiple formats and types of data that need to be
supported by Big Data solutions. Data variety brings challenges for
enterprises in terms of data integration, transformation, processing, and
storage. Figure 1.24 provides a visual representation of data variety, which
includes structured data in the form of financial transactions, semi-
structured data in the form of emails and unstructured data in the form
of images.
Figure 1.24 Examples of high-variety Big Data datasets include
structured, textual, image, video, audio, XML, JSON, sensor data and
metadata
21
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
In other words, Variety: Data sets that are comprised solely of structured
data are not necessarily Big Data, regardless of how voluminous they are.
Big Data is typically comprised of combinations of structured,
unstructured, and semi-structured data. Traditional databases and data
management solutions lack the flexibility and scope to manage the
complex, disparate data sets that make up Big Data.
(iii) Velocity – The term ‘velocity’ refers to the speed of generation of data.
How fast the data is generated and processed to meet the demands,
determines real potential in the data. Big Data Velocity deals with the speed
at which data flows in from sources like business processes, application
logs, networks, and social media sites, sensors, Mobile devices, etc. The
flow of data is massive and continuous.
In Big Data environments, data can arrive at fast speeds, and enormous
datasets can accumulate within very short periods of time. From an
enterprise’s point of view, the velocity of data translates into the amount
of time it takes for the data to be processed once it enters the enterprise’s
perimeter. Coping with the fast inflow of data requires the enterprise to
design highly elastic and available data processing solutions and
corresponding data storage capabilities.
Depending on the data source, velocity may not always be high. For
example, MRI scan images are not generated as frequently as log entries
from a high-traffic webserver. As illustrated in Figure 1.25, data velocity
is put into perspective when considering that the following data volume
can easily be generated in a given minute: 350,000 tweets, 300 hours of
video footage uploaded to YouTube, 171 million emails and 330 GBs of
sensor data from a jet engine.
22
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Figure 1.25 Examples of high-velocity Big Data datasets produced every
minute include tweets, video, emails and GBs generated from a jet
engine.
Moreover, in the past, any data that was generated had to later be entered
into a traditional database system – often manually – before it could be
analyzed or retrieved. Today, Big Data technology allows databases to
process, analyze, and configure data while it is being generated –
sometimes within milliseconds. For businesses, that means real-time data
can be used to capture financial opportunities, respond to customer needs,
thwart fraud, and address any other activity where speed is critical.
(iv) Veracity y – This refers to the inconsistency which can be shown by
the data at times, thus hampering the process of being able to handle and
manage the data effectively. While, modern database technology makes it
possible for companies to amass and make sense of staggering amounts
and types of Big Data, it’s only valuable if it is accurate, relevant, and
timely. For traditional databases that were populated only with structured
data, syntactical errors and typos were the usual culprits when it came to
data accuracy. With unstructured data, there is a whole new set of veracity
challenges. Human bias, social noise, and data provenance issues can all
have an impact upon the quality of data.
Veracity refers to the quality or fidelity of data. Data that enters Big Data
environments needs to be assessed for quality, which can lead to data
processing activities to resolve invalid data and remove noise. In relation
to veracity, data can be part of the signal or noise of a dataset. Noise is data
23
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
that cannot be converted into information and thus has no value, whereas
signals have value and lead to meaningful information. Data with a high
signal-to-noise ratio has more veracity than data with a lower ratio. Data
that is acquired in a controlled manner, for example via online customer
registrations, usually contains less noise than data acquired via
uncontrolled sources, such as blog postings. Thus the signalto-noise ratio
of data is dependent upon the source of the data and its type.
(v) Value
Value is defined as the usefulness of data for an enterprise. The value
characteristic is intuitively related to the veracity characteristic in that the
higher the data fidelity, the more value it holds for the business. Value is
also dependent on how long data processing takes because analytics results
have a shelf-life; for example, a 20 minute delayed stock quote has little to
no value for making a trade compared to a quote that is 20 milliseconds
old.
As demonstrated, value and time are inversely related. The longer it takes
for data to be turned into meaningful information, the less value it has for
a business. Stale results inhibit the quality and speed of informed decision-
making. Figure 1.26 provides two illustrations of how value is impacted by
the veracity of data and the timeliness of generated analytic results.
Figure 1.26 Data that has high veracity and can be analyzed quickly has
more value to a business.
Apart from veracity and time, value is also impacted by the following
lifecycle-related concerns:
24
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
• How well has the data been stored?
• Were valuable attributes of the data removed during data cleansing?
• Are the right types of questions being asked during data analysis?
• Are the results of the analysis being accurately communicated to the
appropriate decision-makers?
6. Why big data is important?
The importance of big data doesn’t revolve around how much data you
have, but what you do with it. You can take data from any source and
analyze it to find answers that enable 1) cost reductions, 2) time reductions,
3) new product development and optimized offerings, and 4) smart
decision making. When you combine big data with high-
powered analytics, you can accomplish business-related tasks such as:
Determining root causes of failures, issues and defects in near-real
time.
Generating coupons at the point of sale based on the customer’s
buying habits.
Recalculating entire risk portfolios in minutes.
Detecting fraudulent behavior before it affects your organization.
9. Sources of Big Data
Big data is used by organizations for the sole purpose of analytics.
However, before companies can set out to extract insights and valuable
information from big data, they must have the knowledge of several big
data sources available. Data, as we know, is massive and exists in various
forms. If it is not classified or sourced well, it can end up wasting precious
time and resources. In order to achieve success with big data, it is important
that companies have the know-how to sift between the various data sources
available and accordingly classify its usability and relevance.
25
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
The range of data-generating things is growing at a phenomenal rate – from
drone satellites to toasters. But for the purposes of categorization, data
sources are generally broken down into three types:
Social data
As is sounds, social data is generated by social media comments, posts,
images, and, increasingly, video. And with the growing global ubiquity of
4G and 5G cellular networks, it is estimated that the number of people in
the world who regularly watch video content on their smartphones will rise
to 2.72 billion by 2023. Although trends in social media and its usage tend
to change quickly and unpredictably, what does not change is its steady
growth as a generator of digital data.
26
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
Machine data
IoT devices and machines are fitted with sensors and have the ability to
send and receive digital data. IoT sensors help companies collect and
process machine data from devices, vehicles, and equipment across the
business. Globally, the number of data-generating things is rapidly
growing – from weather and traffic sensors to security surveillance. The
IDC estimates that by 2025 there will be over 40 billion IoT devices on
earth, generating almost half the world’s total digital data.
Transactional data
This is some of the world’s fastest moving and growing data. For example,
a large international retailer is known to process over one million customer
transactions every hour. And when you add in all the world’s purchasing
and banking transactions, you get a picture of the staggering volume of
data being generated. Furthermore, transactional data is increasingly
comprised of semi-structured data, including things like images and
comments, making it all the more complex to manage and process.
MEDIA AS A BIG DATA SOURCE
Media is the most popular source of big data, as it provides valuable
insights on consumer preferences and changing trends. Since it is self-
broadcasted and crosses all physical and demographical barriers, it is the
fastest way for businesses to get an in-depth overview of their target
audience, draw patterns and conclusions, and enhance their decision-
making. Media includes social media and interactive platforms, like
Google, Facebook, Twitter, YouTube, Instagram, as well as generic media
like images, videos, audios, and podcasts that provide quantitative and
qualitative insights on every aspect of user interaction.
27
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
CLOUD AS A BIG DATA SOURCE
Today, companies have moved ahead of traditional data sources by shifting
their data on the cloud. Cloud storage accommodates structured and
unstructured data and provides business with real-time information and on-
demand insights. The main attribute of cloud computing is its flexibility
and scalability. As big data can be stored and sourced on public or private
clouds, via networks and servers, cloud makes for an efficient and
economical data source.
THE WEB AS A BIG DATA SOURCE
The public web constitutes big data that is widespread and easily
accessible. Data on the Web or ‘Internet’ is commonly available to
individuals and companies alike. Moreover, web services such as
Wikipedia provide free and quick informational insights to everyone. The
enormity of the Web ensures for its diverse usability and is especially
beneficial to start-ups and SME’s, as they don’t have to wait to develop
their own big data infrastructure and repositories before they can leverage
big data.
IOT AS A BIG DATA SOURCE
Machine-generated content or data created from IoT constitute a valuable
source of big data. This data is usually generated from the sensors that are
connected to electronic devices. The sourcing capacity depends on the
ability of the sensors to provide real-time accurate information. IoT is now
gaining momentum and includes big data generated, not only from
computers and smartphones, but also possibly from every device that can
emit data. With IoT, data can now be sourced from medical devices,
vehicular processes, video games, meters, cameras, household appliances,
and the like.
DATABASES AS A BIG DATA SOURCE
Businesses today prefer to use an amalgamation of traditional and modern
databases to acquire relevant big data. This integration paves the way for a
hybrid data model and requires low investment and IT infrastructural costs.
Furthermore, these databases are deployed for several business intelligence
purposes as well. These databases can then provide for the extraction of
insights that are used to drive business profits. Popular databases include a
variety of data sources, such as MS Access, DB2, Oracle, SQL, and
Amazon Simple, among others.
The process of extracting and analyzing data amongst extensive big data
sources is a complex process and can be frustrating and time-consuming.
28
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
These complications can be resolved if organizations encompass all the
necessary considerations of big data, take into account relevant data
sources, and deploy them in a manner which is well tuned to their
organizational goals.
9. How Big Data works
Big Data works when its analysis delivers relevant and actionable insights
that measurably improve the business. In preparation for Big Data
transformation, businesses should ensure that their systems and processes
are sufficiently ready to gather, store, and analyze Big Data.
The three main steps involved in using Big Data
1. Gather Big Data. Much of Big Data is comprised of massive sets of
unstructured data, flooding in from disparate and inconsistent sources.
Traditional disk-based databases and data integration mechanisms are
simply not equal to the task of handling this. Big Data management
requires the adoption of in-memory database solutions and software
solutions specific to Big Data acquisition.
29
@ Asst. Prof. Dr. Lahieb Mohammed Jawad
Computer Networks Engineering Department,
Information Engineering College, Al-Nahrain University
[Link]., Big Data Analysis
2. Store Big Data. By its very name, Big Data is voluminous. Many
businesses have on-premise storage solutions for their existing data and
hope to economize by repurposing those repositories to meet their Big Data
processing needs. However, Big Data works best when it is unconstrained
by size and memory limitations. Businesses that fail to incorporate cloud
storage solutions into their Big Data models from the beginning often
regret this a few months down the road.
3. Analyze Big Data. Without the application of AI and machine learning
technologies to Big Data analysis, it is simply not feasible to realize its full
potential. One of the five V’s of Big Data is “velocity.” For Big Data
insights to be actionable and valuable, they must come quickly. Analytics
processes have to be self-optimizing and able to learn from experience on
a regular basis – an outcome which can only be achieved with AI
functionality and modern database technologies
31
@ Asst. Prof. Dr. Lahieb Mohammed Jawad