UNIT 1
Data Management: Design Data Architecture and manage the data for analysis,
understand various sources of Data like Sensors/Signals/GPS etc. Data Management,
Data Quality(noise, outliers, missing values, duplicate data) and Data Processing &
Processing.
Design Data Architecture and Manage the Data for Analysis
In the beginning times of computers and Internet, the data used was not as much of as it is
today, The data then could be so easily stored and managed by all the users and business
enterprises on a single computer, because the data never exceeded to the extent of 19
exabytes but now in this era, the data has increased about 2.5 quintillions per [Link] of the
data is generated from social media sites like Facebook, Instagram, Twitter, etc, and the other
sources can be e-business, e-commerce transactions, hospital, school, bank data, etc. This
data is impossible to manage by traditional data storing techniques. So Big-Data came into
existence for handling the data which is big and [Link] Data is the field of collecting the
large data sets from various sources like social media, GPS, sensors etc and analyzing them
systematically and extract useful patterns using some tools and techniques by enterprises.
Before analyzing and determining the data, the data architecture must be designed by the
architect.
Data architecture Design and Data Management:
Data architecture design is set of standards which are composed of certain policies, rules,
models and standards which manages, what type of data is collected, from where it is
collected, the arrangement of collected data, storing that data, utilizing and securing the data
into the systems and data warehouses for further [Link] is one of the essential pillars
of enterprise architecture through which it succeeds in the execution of business
[Link] architecture design is important for creating a vision of interactions occurring
between data systems, like for example if data architect wants to implement data integration,
so it will need interaction between two systems and by using data architecture the visionary
model of data interaction during the process can be achieved.
Data architecture also describes the type of data structures applied to manage data and it
provides an easy way for data preprocessing. The data architecture is formed by dividing into
three essential models and then are combined :
1. Conceptual model –It is a business model which uses Entity Relationship (ER)
model for relation between entities and their attributes.
2. Logical model –It is a model where problems are represented in the form of logic
such as rows and column of data, classes, xml tags and other DBMS techniques.
3. Physical model –Physical models holds the database design like which type of
database technology will be suitable for architecture.
A data architect is responsible for all the design, creation, manage, deployment of data
architecture and defines how data is to be stored and retrieved, other decisions are made by
internal bodies.
Factors that influence Data Architecture :
Few influences that can have an effect on data architecture are business policies, business
requirements, Technology used, economics, and data processing needs.
1. Business requirements –These include factors such as the expansion of business, the
performance of the system access, data management, transaction management,
making use of raw data by converting them into image files and records, and then
storing in data warehouses. Data warehouses are the main aspects of storing
transactions in business.
2. Business policies –The policies are rules that are useful for describing the way of
processing data. These policies are made by internal organizational bodies and other
government agencies.
3. Technology in use –This includes using the example of previously completed data
architecture design and also using existing licensed software purchases, database
technology.
4. Business economics –The economical factors such as business growth and loss,
interest rates, loans, condition of the market, and the overall cost will also have an
effect on design architecture.
5. Data processing needs –These include factors such as mining of the data, large
continuous transactions, database management, and other data preprocessing needs.
Data Management:
Data management is the process of managing tasks like extracting data, storing data,
transferring data, processing data, and then securing data with low-cost [Link]
motive of data management is to manage and safeguard the people’s and organization data in
an optimal way so that they can easily create, access, delete, and update the [Link] data
management is an essential process in each and every enterprise growth, without which the
policies and decisions can’t be made for business advancement. The better the data
management the better productivity in [Link] volumes of data like big data are harder
to manage traditionally so there must be the utilization of optimal technologies and tools for
data management such as Hadoop, Scala, Tableau, AWS, etc. Which can further used for big
data analysis in achieving improvements in [Link] management can be achieved by
training the employees necessarily and maintenance by DBA, data analyst, and data
architects.
Understanding various sources of Data like Sensors/Signals/GPS etc.
Data collection is the process of acquiring, collecting, extracting, and storing the voluminous
amount of data which may be in the structured or unstructured form like text, video, audio,
XML files, records, or other image files used in later stages of data analysis. In the process of
big data analysis, “Data collection” is the initial step before starting to analyze the patterns or
useful information in data. The data which is to be analyzed must be collected from different
valid [Link] data which is collected is known as raw data which is not useful now but
on cleaning the impure and utilizing that data for further analysis forms information, the
information obtained is known as “knowledge”. Knowledge has many meanings like business
knowledge or sales of enterprise products, disease treatment, etc. The main goal of data
collection is to collect information-rich data. Data collection starts with asking some
questions such as what type of data is to be collected and what is the source of collection.
Most of the data collected are of two types known as “qualitative data“ which is a group of
non-numerical data such as words, sentences mostly focus on behavior and actions of the
group and another one is “quantitative data” which is in numerical forms and can be
calculated using different scientific tools and sampling data.
The actual data is then further divided mainly into two types known as:
1. Primary data
2. Secondary data
[Link] data:
The data which is Raw, original, and extracted directly from the official sources is known as
primary data. This type of data is collected directly by performing techniques such as
questionnaires, interviews, and surveys. The data collected must be according to the demand
and requirements of the target audience on which analysis is performed otherwise it would be
a burden in the data processing. Few methods of collecting primary data:
1. Interview method:
The data collected during this process is through interviewing the target audience by a person
called interviewer and the person who answers the interview is known as the interviewee.
Some basic business or product related questions are asked and noted down in the form of
notes, audio, or video and this data is stored for processing. These can be both structured and
unstructured like personal interviews or formal interviews through telephone, face to face,
email, etc.
2. Survey method:
The survey method is the process of research where a list of relevant questions are asked and
answers are noted down in the form of text, audio, or video. The survey method can be
obtained in both online and offline mode like through website forms and email. Then that
survey answers are stored for analyzing data. Examples are online surveys or surveys through
social media polls.
3. Observation method:
The observation method is a method of data collection in which the researcher keenly
observes the behavior and practices of the target audience using some data collecting tool and
stores the observed data in the form of text, audio, video, or any raw formats. In this method,
the data is collected directly by posting a few questions on the participants. For example,
observing a group of customers and their behaviour towards the products. The data obtained
will be sent for processing.
4. Experimental method:
The experimental method is the process of collecting data through performing experiments,
research, and investigation. The most frequently used experiment methods are CRD, RBD,
LSD, FD.
CRD- Completely Randomized design is a simple experimental design used in data
analytics which is based on randomization and replication. It is mostly used for
comparing the experiments.
RBD- Randomized Block Design is an experimental design in which the experiment
is divided into small units called blocks. Random experiments are performed on each
of the blocks and results are drawn using a technique known as analysis of variance
(ANOVA). RBD was originated from the agriculture sector.
LSD – Latin Square Design is an experimental design that is similar to CRD and
RBD blocks but contains rows and columns. It is an arrangement of NxN squares with
an equal amount of rows and columns which contain letters that occurs only once in a
row. Hence the differences can be easily found with fewer errors in the experiment.
Sudoku puzzle is an example of a Latin square design.
FD- Factorial design is an experimental design where each experiment has two
factors each with possible values and on performing trail other combinational factors
are derived.
2. Secondary data:
Secondary data is the data which has already been collected and reused again for some valid
purpose. This type of data is previously recorded from primary data and it has two types of
sources named internal source and external source.
Internal source:
These types of data can easily be found within the organization such as market record, a sales
record, transactions, customer data, accounting resources, etc. The cost and time consumption
is less in obtaining internal sources.
External source:
The data which can’t be found at internal organizations and can be gained through external
third party resources is external source data. The cost and time consumption is more because
this contains a huge amount of data. Examples of external sources are Government
publications, news publications, Registrar General of India, planning commission,
international labor bureau, syndicate services, and other non-governmental publications.
Other sources:
Sensors data: With the advancement of IoT devices, the sensors of these devices
collect data which can be used for sensor data analytics to track the performance and
usage of products.
Satellites data: Satellites collect a lot of images and data in terabytes on daily basis
through surveillance cameras which can be used to collect useful information.
Web traffic: Due to fast and cheap internet facilities many formats of data which is
uploaded by users on different platforms can be predicted and collected with their
permission for data analysis. The search engines also provide their data through
keywords and queries searched mostly.
Data Management
Data management is the process of managing tasks like extracting data, storing data,
transferring data, processing data, and then securing data with low-cost [Link]
motive of data management is to manage and safeguard the people’s and organization data in
an optimal way so that they can easily create, access, delete, and update the [Link] data
management is an essential process in each and every enterprise growth, without which the
policies and decisions can’t be made for business advancement. The better the data
management the better productivity in [Link] volumes of data like big data are harder
to manage traditionally so there must be the utilization of optimal technologies and tools for
data management such as Hadoop, Scala, Tableau, AWS, etc. Which can further used for big
data analysis in achieving improvements in [Link] management can be achieved by
training the employees necessarily and maintenance by DBA, data analyst, and data
architects.
Data Quality(Noise,Outliers,Missing values,Duplicate data)
Real-world data tend to be incomplete, noisy, and inconsistent. Data cleaning (or data
cleansing) routines attempt to fill in missing values, smooth out noise while identifying
outliers, and correct inconsistencies in the data.
Noise
Noise is a random error or variance in a measured variable. Given a numeric attribute such as,
say, price, smooth out the data to remove the [Link] data smoothing techniques are:
[Link]: Binning methods smooth a sorted data value by consulting its “neighborhood,”
that is, the values around it. The sorted values are distributed into a number of “buckets,” or
bins.
a) Smoothing by bin means, each value in a bin is replaced by the mean value of the bin.
For example, the mean of the values 4, 8, and 15 in Bin 1 is 9. Therefore, each
original value in this bin is replaced by the value 9.
b) Smoothing by bin medians, in which each bin value is replaced by the bin median.
c) Smoothing by bin boundaries, the minimum and maximum values in a given bin are
identified as the bin boundaries. Each bin value is then replaced by the closest
boundary value.
2. Regression: Data smoothing can also be done by regression, a technique that conforms
data values to a function. Linear regression involves finding the “best” line to it two attributes
(or variables) so that one attribute can be used to predict the other. Multiple linear regression
is an extension of linear regression, where more than two attributes are involved and the data
are fit to a multidimensional surface.
3. Outlier analysis: Outliers may be detected by clustering, for example, where similar
values are organized into groups, or “clusters.” Intuitively, values that fall outside of the set
of clusters may be considered outliers.
Outliers are the values that fall outside of the cluster sets.
Outliers
An outlier is a data object that deviates significantly from the rest of the objects, as if it were
generated by a different mechanism.
Types of Outliers
Outliers can be classified into three categories, namely global outliers, contextual (or
conditional) outliers, and collective outliers.
[Link] Outliers
In a given data set, a data object is a global outlier if it deviates significantly from the rest of
the data set. Global outliers are sometimes called point anomalies, and are the simplest type
of outliers.
Global outlier detection is important in many applications. Consider intrusion detection in
computer [Link] the communication behavior of a computer is very different from the
normal patterns (e.g., a large number of packages is broad cast in a short time), this behavior
may be considered as a global outlier and the corresponding computer is a suspected victim
of hacking. As another example, in trading transaction auditing systems, transactions that do
not follow the regulations are considered as global outliers and should be held for further
examination.
[Link] Outliers
In a given data set, a data object is a contextual outlier if it deviates significantly with respect
to a specific context of the object. Contextual outliers are also known as conditional outliers
because they are conditional on the selected context. Generally, in contextual outlier
detection, the attributes of the data objects are divided into two groups:
Contextual attributes: The contextual attributes of a data object define the object’s context. In
the temperature example, the contextual attributes may be date and location.
Behavioral attributes: These define the object’s characteristics, and are used to evaluate
whether the object is an outlier in the context to which it belongs. In the temperature
example, the behavioral attributes may be the temperature, humidity, and pressure.
Example: A configuration of behavioral attribute values may be considered an outlier in one
context (e.g., 28C is an outlier for a Toronto winter), but not an outlier in another context
(e.g., 28C is not an outlier for a Toronto summer).
[Link] Outliers
Given a data set, a subset of data objects forms a collective outlier if the objects as a whole
deviate significantly from the entire data set. Importantly, the individual data objects may not
be outliers.
Example: If 100 orders are delayed on a single day, those 100 orders as a whole form an
outlier, although each of them may not be regarded as an outlier if considered individually.
Collective outlier detection has many important applications. For example, in intrusion
detection, a denial-of-service package from one computer to another is considered normal,
and not an outlier at all. However, if several computers keep sending denial-of-service
packages to each other, they as a whole should be considered as a collective outlier. The
computers involved may be suspected of being compromised by an attack.
Missing values
Many tuples have no recorded value for several attributes such as customer income. The
following methods helps in filling the missing values for the attribute.
1. Ignore the tuple: This is usually done when the class label is missing. This method is
not very effective, when the percentage of missing values per attribute varies
considerably.
2. Fill in the missing value manually: In general, this approach is time consuming and
may not be feasible given a large data set with many missing values.
3. Use a global constant to fill in the missing value: Replace all missing attribute
values by the same constant such as a label like “Unknown”. If missing values are
replaced by, say, “Unknown,” then the mining program may mistakenly think that
they form an interesting concept, since they all have a value in common that of
Unknown.
4. Use a measure of central tendency for the attribute (e.g., the mean or median) to
fill in the missing value:
For example, suppose that the data distribution regard ing the income of
AllElectronics customers is symmetric and that the mean income is $56,000. Use this
value to replace the missing value for income.
5. Use the attribute mean or median for all samples belonging to the same class as
the given tuple:
For example, if classifying customers according to credit risk, we may replace the
missing value with the mean income value for customers in the same credit risk
category as that of the given tuple. If the data distribution for a given class is skewed,
the median value is a better choice.
6. Use the most probable value to fill in the missing value:
This may be determined with regression, inference-based tools using a Bayesian
formalism, or decision tree induction.
Duplicate data/Redundant Data
Duplication should be detected at the tuple level(where there are 2 or more identical tuples
for a give unique data entry case).The use of denormalized tables is another source of data
redundancy .Inconsistencies arise between various duplicates, due to inaccurate data entry or
updating some but not all data occurrences. For example if a purchase order database
contains attributes for the purchasers name and address instead of a key to this information in
a purchaser database, discrepancies can occur, such as the same purchasers name appearing
with different addresses within the purchase order database.
Data Preprocessing
The major steps involved in data preprocessing are data cleaning, data integration, data
reduction, and data transformation.
Data cleaning routines work to “clean” the data by filling in missing values, smooth ing
noisy data, identifying or removing outliers, and resolving inconsistencies. been applied.
Furthermore, dirty data can cause confusion for the mining procedure, resulting in unreliable
output.
Data integration involve integrating multiple databases, data cubes, or files from multiple
sources. Some attributes representing a given concept may have different names in different
databases, causing inconsistencies and redundancies.
Forexample,the attribute for customer identification may be referred to as customer_id in one
data store and cust_id in another. Having a large amount of redundant data may slow down or
confuse the knowledge discovery process.
Data reduction obtains a reduced representation of the data set that is much smaller in
volume, yet produces the same (or almost the same) analytical results. Data reduction
strategies include dimensionality reduction and numerosity reduction.
In dimensionality reduction, data encoding schemes are applied so as to obtain a reduced or
“compressed” representation of the original data. Examples include
data compression techniques (e.g., wavelet transforms and principal components
analysis),
attribute subset selection (e.g., removing irrelevant attributes), and
attribute construction (e.g., where a small set of more useful attributes is derived from
the original set).
In numerosity reduction, the data are replaced by alternative, smaller representations using
parametric models (e.g., regression or log-linear models) or
nonparametric models (e.g., histograms, clusters, sampling, or data aggregation).
For example, consider the customer data that contain the attributes age and annual salary. The
annual salary attribute usually takes much larger values than age. Therefore, if the attributes
are left unnormalized, the distance measurements taken on annual salary will generally
outweigh distance measurements taken on age.
Data Transformation:
Discretization and concept hierarchy generation are powerful tools for data mining in that
they allow data mining at multiple abstraction levels. The raw data are replaced by higher
level concepts such as youth, adult or senior.