0% found this document useful (0 votes)
30 views99 pages

CST 322 Module 2: Data Analytics Intro

Module 2 of CST 322 Data Analytics introduces the fundamentals of data analytics, including the analytics process model, data collection, and the data analytics life cycle. It emphasizes the importance of defining business problems, identifying data sources, and cleaning and transforming data for analysis. The module also outlines the requirements for analytical models and the phases of the data analytics lifecycle, from discovery to model building and effectiveness measurement.

Uploaded by

backupanji2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
30 views99 pages

CST 322 Module 2: Data Analytics Intro

Module 2 of CST 322 Data Analytics introduces the fundamentals of data analytics, including the analytics process model, data collection, and the data analytics life cycle. It emphasizes the importance of defining business problems, identifying data sources, and cleaning and transforming data for analysis. The module also outlines the requirements for analytical models and the phases of the data analytics lifecycle, from discovery to model building and effectiveness measurement.

Uploaded by

backupanji2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CST 322 DATA ANALYTICS

Module -2 (Introduction to Data Analytics)

Ojus Thomas Lee


CE Kidangoor

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 1 / 99
CO’s

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 2 / 99
CO - Module Mapping

Mod 1 Mod 2 Mod 3 Mod 4 Mod 5


CO1 X
CO2 X
CO3 X
CO4 X
CO5 X
CO6 X

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 3 / 99
PO’s

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 4 / 99
CO-PO-Mapping

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 5 / 99
Syllabus - Module -2 (Introduction to Data Analytics)

Introduction to Data Analysis - Analytics, Analytics Process Model,


Analytical Model Requirements. Data Analytics Life Cycle overview.
Basics of data collection, sampling, preprocessing and dimensionality
reduction

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 6 / 99
PO’s

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 7 / 99
CO-PO-Mapping

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 8 / 99
1
Introduction to Data Analysis

Data analytics is a multidisciplinary field that employs a wide range


of analysis techniques, including math, statistics, and computer
science, to draw insights from data sets.
Data analytics is a broad term that includes everything from simply
analyzing data to theorizing ways of collecting data and creating the
frameworks needed to store it.

1
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 9 / 99
3
Introduction to Data Analysis
Data analytics is the process of collecting, transforming, and
organizing data in order to draw conclusions, make predictions, and
drive informed decision making.

2
Figure 1:
2
[Link]
3
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 10 / 99
4
Introduction to Data Analysis

Analysis is the detailed examination of the elements or structure of


something.
Analytics is the systematic computational analysis of data or
statistics.
Data Analytics is a wide area involving handling data with a lot of
necessary tools to produce helpful decisions with useful predictions for
a better output,
Data Analysis is a subset of Data Analytics which helps us to
understand the data by questioning and to collect useful insights from
the already available.

4
[Link]
are-they-similar/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 11 / 99
5
Introduction to Data Analysis

Data Analytics is the process of exploring the data from the past to
make appropriate decisions in the future by using valuable insights.
Data Analysis helps in understanding the data and provides required
insights from the past to understand what happened so far.

5
[Link]
are-they-similar/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 12 / 99
6
Analytics Vs Analysis

6
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 13 / 99
7
Analytics Process Model

7
[Link]
talking-about-the-analytics-process-model/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 14 / 99
8
Analytics Process Model - Identify Business Problem

As a first step, a thorough definition of the business problem to be


addressed is needed.
The objective of applying analytics needs to be unambiguously
defined.
Eg : customer segmentation of a mortgage portfolio, retention
modeling for a postpaid Telco subscription, or fraud detection for
credit cards.
Defining the perimeter of the analytical modeling exercise requires a
close collaboration between the data scientists and business experts.

8
[Link]
talking-about-the-analytics-process-model/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 15 / 99
9
Analytics Process Model - Identify Data Sources

All source data that could be of potential interest need to be


identified.
The golden rule here is: the more data, the better!
The analytical model itself will later decide which data are relevant
and which are not for the task at hand.

9
[Link]
talking-about-the-analytics-process-model/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 16 / 99
10
Analytics Process Model - Data Collection

You need to collect transactional business data and customer-related


information from the past few years to address the problems your
business is facing.
The data can have information about the total units that were sold
for a product, the sales, and profit that were made, and also when
was the order placed.
Past data plays a crucial role in shaping the future of a business.

10
[Link]
analytics
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 17 / 99
11
Analytics Process Model - Data Cleaning

All the data you collect will often be disorderly, messy, and contain
unwanted missing values.
Such data is not suitable or relevant for performing data analysis.
Hence, you need to clean the data to remove unwanted, redundant,
and missing values to make it ready for analysis.

11
h[Link]
analytics
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 18 / 99
12
Analytics Process Model - Data Transformation

Data transformation is the process of converting data from a source


format to a destination format.
This can include cleansing data by changing data types, deleting nulls
or duplicates, aggregating data, enriching the data, or other
transformations.
For example, ”India” can be transformed to ”IN” to match the
destination format.

12
h[Link]
analytics
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 19 / 99
13
Analytics Process Model - Data Analysis

Depending on the business objective and the exact task at hand, a


particular analytical technique will be selected and implemented by
the data scientist.
Once the results are obtained, they will be interpreted and evaluated
by the business experts.
Results may be clusters, rules, patterns, or relations, among others,
all of which will be called analytical models resulting from applying
analytics.

13
h[Link]
analytics
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 20 / 99
14
Analytics Process Model - Deploy the model

Once the analytical model has been appropriately validated and


approved, it can be put into production as an analytics application
(e.g., decision support system, scoring engine).

14
h[Link]
analytics
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 21 / 99
Analytical Model Requirements.15

Business relevance
Statistical performance.
Interpretable and Justifiable
Operationally efficient and Economic cost
International Regulation and Legislation

15
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 22 / 99
Analytical Model Requirements.16

Business relevance
The analytical model should actually solve the business problem for
which it was developed.
In order to achieve business relevance, it is of key importance that the
business problem to be solved is appropriately defined, qualified, and
agreed upon by all parties involved at the outset of the analysis.

16
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 23 / 99
Analytical Model Requirements.17

Statistical performance.
The model should have statistical significance and predictive power.
How this can be measured will depend upon the type of analytics
considered.
For example, in a classification setting (churn, fraud), the model should
have good discrimination power.

17
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 24 / 99
Analytical Model Requirements.18

Interpretable and Justifiable.


Interpretability refers to the fact that the analytical model should be
comprehensible or understandable to the decision maker (e.g. marketer,
fraud analyst, credit expert).
Justifiability indicates that the model is in accordance with the
expectations and business knowledge of the expert.
Both interpretability and justifiability are subjective and depend on the
knowledge and experience of the decision maker.

18
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 25 / 99
Analytical Model Requirements.19

Operationally efficient and Economic cost.


Operational efficiency relates to the effort that is needed to evaluate,
monitor, backtest or rebuild the model.
From this perspective,
In settings like credit card fraud detection, operational efficiency is very
important because a decision should be made within a few seconds
after the credit card transaction was initiated.

19
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 26 / 99
Analytical Model Requirements.20

International Regulation and Legislation


This refers to the extent to which the model is compliant with
regulation and legislation.
In a credit risk modeling setting, it is important that the models are
compliant with the Basel II and III regulations.
In an analytical insurance setting, the Solvency II accord must be
respected.

20
[Link]
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 27 / 99
21
Data Analytics Life Cycle overview

21
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 28 / 99
22
Data Analytics Life Cycle overview

Data Analytics Lifecycle defines the roadmap of how data is


generated, collected, processed, used, and analyzed to achieve
business goals.
It offers a systematic way to manage data for converting it into
information that can be used to fulfill organizational and project goals.

22
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 29 / 99
23
Data Analytics Lifecycle Phases

Data Discovery and Formation


Data Preparation and Processing
Design a Model
Model Building
Result Communication and Publication
Measuring of Effectiveness

23
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 30 / 99
24
Data Analytics Lifecycle Phases

Data Discovery and Formation


This phase is all about defining the data’s purpose and how to achieve
it by the end of the data analytics lifecycle.
The stage consists of identifying critical objectives a business is trying
to discover by mapping out the data.
In this phase, the team also evaluates technology, people, data, and
time.
For example, while dealing with a small dataset, the team can use
Excel.
However, heftier tasks demand more rigid tools for data preparation
and exploration like Python, R, Tableau Desktop or Tableau Prep, and
other data cleaning tools.

24
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 31 / 99
25
Data Analytics Lifecycle Phases

Data Preparation and Processing


In this phase, the experts’ focus shifts from business requirements to
information requirements.
One of the essential aspects of this phase is ensuring data availability
for processing.
The stage encompasses the collection, processing, and cleansing of the
accumulated data.

25
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 32 / 99
26
Data Analytics Lifecycle Phases

Design a Model
The team explores the data,
Identifies relations between data points to select the key variables, and
Eventually devises a suitable model.

26
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 33 / 99
27
Data Analytics Lifecycle Phases

Model Building
In this phase, the team develops testing, training, and production
datasets.
Further, the team builds and executes models meticulously as planned
during the model planning phase.
They test data and try to find out answers to the given objectives.

27
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 34 / 99
28
Data Analytics Lifecycle Phases

Result Communication and Publication


This phase aims to determine whether the project results are a success
or failure and start collaborating with significant stakeholders.
The team identifies the vital findings of their analysis, measures the
associated business value

28
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 35 / 99
29
Data Analytics Lifecycle Phases

Measuring of Effectiveness
In this final phase, the team presents an in-depth report with coding,
briefing, key findings, and technical documents and papers to the
stakeholders.
Besides this, the data is moved to a live environment and monitored to
measure the analysis’s effectiveness.
If the findings are in line with the objective, the results and reports are
finalized.
On the other hand, if they deviate from the set intent, the team moves
backward in the lifecycle to any previous phase to change the input and
get a different outcome.

29
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 36 / 99
30
Basics of data collection

TYPES OF DATA SOURCES


Transactions and Database
Text Files
CSV files
Cloud Data warehouses / Cloud Databases
Miscellaneous Sources - Multimedia, Social Media API

30
bart-baesens-analytics-in-a-big-data-world.-the-essential-guide-to-data-science-and-
its-applications-wiley-2014-Chapter
2
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 37 / 99
31
Basics of data collection - TYPES OF DATA SOURCES

Transactions and Database


Transactional data consist of structured, low-level, detailed information
capturing the key characteristics of a customer transaction (e.g.,
purchase, claim, cash transfer, credit card payment).
This type of data is usually stored in massive online transaction
processing (OLTP) relational databases.

31
bart-baesens-analytics-in-a-big-data-world.-the-essential-guide-to-data-science-and-
its-applications-wiley-2014-Chapter
2
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 38 / 99
32
Basics of data collection - TYPES OF DATA SOURCES

Text Files
The most basic method for data storage is using a text file.
The content in the text file is structured and will follow a specific
format.
The most common usage of text file appears in logging.
The log entries are stored in a specific format which can be read and
extracted using a programming language.

32
[Link]
overview-and-usage-cdbf7e86dbbd
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 39 / 99
33
Basics of data collection - TYPES OF DATA SOURCES

CSV files
One of the most common ways in which data is stored is in a CSV file.
CSV file consists of data that are comma-separated. When opened in
software like Excel, CSV displays like an excel sheet, where data is
stored column-wise and row-wise.
CSV files can be easily accessed and processed using programming
language like Python using libraries like Pandas.

33
[Link]
overview-and-usage-cdbf7e86dbbd
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 40 / 99
34
Basics of data collection - TYPES OF DATA SOURCES

Cloud Data warehouses / Cloud Databases


Data science often correlates with cloud platforms.
The major storage solutions offered by the cloud are data warehouses
and cloud database.
Warehouses are used to store large amounts of incoming data for
analytics purposes,
while cloud database stores the usual customer data in the cloud.
Both of these can be accessed from an application using its respective
APIs.

34
[Link]
overview-and-usage-cdbf7e86dbbd
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 41 / 99
35
Basics of data collection - TYPES OF DATA SOURCES

Miscellaneous Sources - Multimedia, Social Media API


Certain data will be present in multimedia forms like Images or Audio.
These are present as it is in folders.
For such type of data, libraries like OpenCV is used to read image and
convert into an array.
Audios are generally converted to an image (spectrogram) which again
comes back to an image processing problem.

35
[Link]
overview-and-usage-cdbf7e86dbbd
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 42 / 99
36
Basics of data collection - TYPES OF DATA SOURCES

Miscellaneous Sources - Multimedia, Social Media API


Certain data science problems require you to connect to an API or
another platform like social media to obtain specific data.
Data is fetched from Twitter hashtags and sentiment analysis is
performed on it for natural language processing.
APIs provide live streaming of data (Eg: COVID data or Election
results data).
This helps perform live analytics and dashboard generation.

36
[Link]
overview-and-usage-cdbf7e86dbbd
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 43 / 99
37
What is Sampling?

It is the practice of selecting an individual group from a population in


order to study the whole population.
sampling type comes under two broad categories:
Probability sampling - Random selection techniques are used to select
the sample.
Non-probability sampling - Non-random selection techniques based on
certain criteria are used to select the sample.

37
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 44 / 99
38
What is Sampling?

38
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 45 / 99
39
What is Sampling?

Probability Sampling Techniques


Simple Random Sampling
Systematic Sampling
Stratified Sampling
Cluster

39
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 46 / 99
40
What is Sampling?

Probability Sampling Techniques - Simple Random Sampling


In simple random sampling, the researcher selects the participants
randomly.
There are a number of data analytics tools like random number
generators and random number tables used that are based entirely on
chance.
Example: The researcher assigns every member in a company database
a number from 1 to 1000 (depending on the size of your company) and
then use a random number generator to select 100 members.

40
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 47 / 99
41
What is Sampling?

Probability Sampling Techniques - Systematic Sampling


In systematic sampling, every population is given a number as well like
in simple random sampling. However, instead of randomly generating
numbers, the samples are chosen at regular intervals.
Example: The researcher assigns every member in the company
database a number.
Instead of randomly generating numbers, a random starting point (say
5) is selected.
From that number onwards, the researcher selects every, say, 10th
person on the list (5, 15, 25, and so on) until the sample is obtained.

41
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 48 / 99
42
What is Sampling?

Probability Sampling Techniques - Stratified Sampling


In stratified sampling, the population is subdivided into subgroups,
called strata, based on some characteristics (age, gender, income, etc.).
After forming a subgroup, you can then use random or systematic
sampling to select a sample for each subgroup.
This method allows you to draw more precise conclusions because it
ensures that every subgroup is properly represented.
Example: If a company has 500 male employees and 100 female
employees, the researcher wants to ensure that the sample reflects the
gender as well. So the population is divided into two subgroups based
on gender.

42
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 49 / 99
What is Sampling? - Stratified-Sampling43

43
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 50 / 99
What is Sampling? - Stratified-Sampling44

44
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 51 / 99
45
What is Sampling?

Probability Sampling Techniques - Cluster Sampling


In cluster sampling, the population is divided into subgroups, but each
subgroup has similar characteristics to the whole sample.
Instead of selecting a sample from each subgroup, you randomly select
an entire subgroup. This method is helpful when dealing with large and
diverse populations.
Example: A company has over a hundred offices in ten cities across the
world which has roughly the same number of employees in similar job
roles. The researcher randomly selects 2 to 3 offices and uses them as
the sample.

45
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 52 / 99
What is Sampling? - ClusterSampling46

46
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 53 / 99
47
What is Sampling?

Non-Probability Sampling
Non-Probability Sampling Techniques is one of the important types of
Sampling techniques.
In non-probability sampling, not every individual has a chance of being
included in the sample.
This sampling method is easier and cheaper but also has high risks of
sampling bias.
It is often used in exploratory and qualitative research with the aim to
develop an initial understanding of the population.

47
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 54 / 99
48
What is Sampling?

Non-Probability Sampling Techniques


Convenience Sampling
Voluntary Response Sampling
Purposive Sampling
Snowball Sampling

48
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 55 / 99
49
What is Sampling?

Non-Probability Sampling Techniques - Convenience Sampling


In this sampling method, the researcher simply selects the individuals
which are most easily accessible to them.
This is an easy way to gather data, but there is no way to tell if the
sample is representative of the entire population.
The only criteria involved is that people are available and willing to
participate.

49
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 56 / 99
50
What is Sampling?

Non-Probability Sampling Techniques - Voluntary response sampling


Voluntary response sampling is similar to convenience sampling, in the
sense that the only criterion is people are willing to participate.
However, instead of the researcher choosing the participants, the
participants volunteer themselves.

50
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 57 / 99
51
What is Sampling?

Non-Probability Sampling Techniques - Purposive Sampling


In purposive sampling, the researcher uses their expertise and judgment
to select a sample that they think is the best fit.
It is often used when the population is very small and the researcher
only wants to gain knowledge about a specific phenomenon rather than
make statistical inferences.

51
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 58 / 99
52
What is Sampling?

Non-Probability Sampling Techniques - Snowball Sampling


In snowball sampling, the research participants recruit other
participants for the study.
It is used when participants required for the research are hard to find.
It is called snowball sampling because like a snowball, it picks up more
participants along the way and gets larger and larger.

52
[Link]
article#what is sampling
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 59 / 99
53
Data Preprocessing

Data preprocessing is the process of transforming raw data into an


understandable format.
It is also an important step in data mining as we cannot work with
raw data.
The quality of the data should be checked before applying machine
learning or data mining algorithms.

53
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 60 / 99
54
Data Preprocessing

54
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 61 / 99
55
Data Preprocessing - Data cleaning

Data cleaning is the process to remove incorrect data, incomplete


data and inaccurate data from the datasets, and it also replaces the
missing values.
There are some techniques in data cleaning
Handling Missing values
Handling Noisy data

55
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 62 / 99
56
Data Preprocessing - Data cleaning

Handling Missing values


Standard values like “Not Available” or “NA” can be used to replace
the missing values.
Missing values can also be filled manually but it is not recommended
when that dataset is big.
The attribute’s mean value can be used to replace the missing value
when the data is normally distributed wherein in the case of
non-normal distribution median value of the attribute can be used.
While using regression or decision tree algorithms the missing value can
be replaced by the most probable value.

56
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 63 / 99
57
Handling Missing values

57
Analytics in a Big Data World: The Essential Guide to Data Science and its
Applications, Chapter 2; Bart Baesens
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 64 / 99
58
Handling Missing values

Replace (impute)
Implies replacing the missing value with a known value
One could impute the missing credit bureau scores with the average or
median of the known values.
For marital status, the mode can then be used.
One could also apply regression-based imputation whereby a regression
model is estimated to model a target variable (e.g., credit bureau
score) based on the other information available (e.g., age, income).

58
Analytics in a Big Data World: The Essential Guide to Data Science and its
Applications, Chapter 2; Bart Baesens
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 65 / 99
59
Handling Missing values

Delete.
The most straightforward option and consists of deleting observations
or variables with lots of missing values.
Assumes that information is missing at random and has no meaningful
interpretation and/or relationship to the target.

59
Analytics in a Big Data World: The Essential Guide to Data Science and its
Applications, Chapter 2; Bart Baesens
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 66 / 99
60
Handling Missing values

Keep.
Missing values can be meaningful (e.g., a customer did not disclose his
or her income because he or she is currently unemployed).
Obviously, this is clearly related to the target (e.g., good/bad risk or
churn) and needs to be considered as a separate category.

60
Analytics in a Big Data World: The Essential Guide to Data Science and its
Applications, Chapter 2; Bart Baesens
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 67 / 99
61
Data Preprocessing - Data cleaning

Noisy Data handling


Noisy generally means random error or containing unnecessary data
points.
Binning: This method is to smooth or handle noisy data. First, the
data is sorted then and then the sorted values are separated and stored
in the form of bins.
There are three methods for smoothing data in the bin.
1 Smoothing by bin mean method: In this method, the values in the
bin are replaced by the mean value of the bin;
2 Smoothing by bin median: In this method, the values in the bin are
replaced by the median value;
3 Smoothing by bin boundary: In this method, the using minimum and
maximum values of the bin values are taken and the values are replaced
by the closest boundary value.

61
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 68 / 99
62
Data Preprocessing

62
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 69 / 99
63
Data Preprocessing - Data cleaning

Noisy Data handling


Noisy generally means random error or containing unnecessary data
points.
Regression: This is used to smooth the data and will help to handle
data when unnecessary data is present.
For the analysis, purpose regression helps to decide the variable which
is suitable for our analysis.
Clustering : This is used for finding the outliers and also in grouping
the data. Clustering is generally used in unsupervised learning.

63
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 70 / 99
64
Data Preprocessing - Data integration

The process of combining multiple sources into a single dataset. The


Data integration process is one of the main components in data
management.
While doing data integration you have to be care full about
Schema integration: Integrates metadata(a set of data that describes
other data) from different sources.
Entity identification problem: Identifying entities from multiple
databases. For example, the system or the user should know student id
of one database and student name of another database belongs to the
same entity.
Detecting and resolving data value concepts: The data taken from
different databases while merging may differ.
Like the attribute values from one database may differ from another
database. For example, the date format may differ like
“MM/DD/YYYY” or “DD/MM/YYYY”.
64
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 71 / 99
65
Data Preprocessing - Data reduction

This process helps in the reduction of the volume of the data which
makes the analysis easier yet produces the same or almost the same
result. This reduction also helps to reduce storage space.
techniques in data reduction are
Dimensionality reduction,
Numerosity reduction,
Data compression.

65
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 72 / 99
66
Data Preprocessing - Data reduction

Numerosity reduction
In numerosity reduction, data volume is reduced by choosing
alternating, smaller forms of data representation.
In numerosity reduction, regression and log-linear models can be used
to approximate the given data.
In linear regression, the data are modeled to fit a straight line.

66
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 73 / 99
67
Data Preprocessing - Data reduction

Data compression
The compressed form of data is called data compression.
This compression can be lossless or lossy.
When there is no loss of information during compression it is called
lossless compression.
Whereas lossy compression reduces information but it removes only the
unnecessary information.

67
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 74 / 99
68
Data Preprocessing - Data reduction

Dimensionality reduction
This process is necessary for real-world applications as the data size is
big.
In this process, the reduction of random variables or attributes is done
so that the dimensionality of the data set can be reduced.
Combining and merging the attributes of the data without losing its
original characteristics.
This also helps in the reduction of storage space and computation time.

68
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 75 / 99
69
Data Preprocessing - Data Transformation

The change made in the format or the structure of the data is called
data transformation.
Methods of Data transformation are
Smoothing
Aggregation
Discretization
Normalization

69
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 76 / 99
70
Data Preprocessing - Data Transformation

Smoothing:
With the help of algorithms, we can remove noise from the dataset and
helps in knowing the important features of the dataset.
By smoothing we can find even a simple change that helps in
prediction.
Aggregation:
In this method, the data is stored and presented in the form of a
summary.
The data set which is from multiple sources is integrated into with
data analysis description.
This is ideal when you want to perform statistical analysis over your
data as you might want to aggregate your data over a specific time
period and provide statistics such as average, sum, minimum and
maximum.

70
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 77 / 99
71
What is Aggregation?

71
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 78 / 99
72
Data Preprocessing - Data Transformation

Discretization:
The continuous data here is split into intervals.
Discretization reduces the data size.
For example, rather than specifying the class time, we can set an
interval like (3 pm-5 pm, 6 pm-8 pm).
Normalization:
It is the method of scaling the data so that it can be represented in a
smaller range. Example ranging from -1.0 to 1.0.

72
[Link]
mining-a-hands-on-guide/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 79 / 99
73
What is Normalization?

Min-Max Normalization
Min-Max is the most commonly used transformation.
This transforms the numerical variable into a new range, for example,
0 to 1.
It is calculated by the formula given below.

73
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 80 / 99
74
What is Normalization?

74
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 81 / 99
75
What is Normalization?

Standardization (or Z-score normalization) is that the features will be


rescaled to ensure the mean and the standard deviation to be 0 and 1,
respectively.

75
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 82 / 99
76
What is Normalization?

mean is 33.75 and Standard Deviation is 24.95

76
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 83 / 99
77
What is Normalization?

j is the number of digits in v


77
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 84 / 99
78
What is Normalization?

78
[Link]
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 85 / 99
79
Dimensionality Reduction

The number of input variables or features for a dataset is referred to


as its dimensionality or dimensions (d).
Dimensionality reduction refers to techniques that reduce the
number of input variables in a dataset.
More input features often make a predictive modeling task more
challenging to model, more generally referred to as the curse of
dimensionality.
Dimensionality reduction techniques are used in machine learning to
simplify a classification or regression dataset in order to better fit a
predictive model.

79
[Link]
learning/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 86 / 99
Dimensionality Reduction - Problem With Many Input
Variables 80

High-dimensionality might mean hundreds, thousands, or even


millions of input variables.
Fewer input dimensions often mean correspondingly fewer parameters
or a simpler structure in the machine learning model, referred to as
Degrees of freedom.
In machine learning, the degrees of freedom may refer to the number
of parameters in the model, such as the number of coefficients in
a linear regression model or the number of weights in a deep
learning neural network.
A model with too many degrees of freedom is likely to overfit the
training dataset and therefore may not perform well on new data.

80
[Link]
learning/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 87 / 99
Dimensionality Reduction - Problem With Many Input
Variables 81
The performance of machine learning algorithms can degrade with
too many input variables.
If your data is represented using rows and columns, such as in a
spreadsheet, then the input variables are the columns that are fed as
input to a model to predict the target variable.
Input variables are also called features.
We can consider the columns of data representing dimensions on an
n-dimensional feature space and the rows of data as points in that
space.
Having a large number of dimensions in the feature space can mean
that the volume of that space is very large, and
in turn, the points that we have in that space (rows of data) often
represent a small and non-representative sample.
81
[Link]
learning/
Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 88 / 99
Dimensionality Reduction - Methods

There are two main methods for reducing dimensionality:


feature selection
feature extraction

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 89 / 99
Dimensionality Reduction - Methods

In feature selection, we are interested in


finding k of the d dimensions that give us the most information and
we discard the other (d − k) dimensions.
In feature extraction, we are interested in finding a new set of k
dimensions that are combinations of the original d dimensions.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 90 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

In subset selection, we are interested in finding the best subset of


the set of features.
The best subset contains the least number of dimensions that most
contribute to accuracy.
We discard the remaining, unimportant dimensions.
Using a suitable error function, this can be used in both regression
and classification problems.
There are 2d possible subsets of d variables, but we cannot test for all
of them unless d is small and
we employ heuristics to get a reasonable (but not optimal) solution in
reasonable (polynomial) time.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 91 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

There are two approaches:


Forward selection
Backward selection

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 92 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

In forward selection
We start with no variables and add them one by one,
at each step adding the one that decreases the error the most, until
any further addition does not decrease the error (or decreases it only
sightly).

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 93 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

In backward selection
We start with all variables and remove them one by one,
at each step removing the one that decreases the error the most (or
increases it only slightly), until
any further removal increases the error significantly.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 94 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

Let us denote by F , a feature set of input dimensions,


F = {xi , such that , i = 1, ..., d.
E (F ) denotes the error incurred on the validation sample when only
the inputs in F are used.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 95 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

sequential forward selection


Start with no features, F = φ.
At each step, for all possible xi , we train our model on the training set
and calculate E (F ∪ xj ) on the validation set.
Then, we choose that input xj that causes the least error
j = arg min E (F ∪ xi )
add xj to F if E (F ∪ xj ) < E (F )

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 96 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

sequential forward selection


We stop if adding any feature does not decrease E .
We may even decide to stop earlier if the decrease in error is too small,
There can be a user-defined threshold that depends on the application
constraints, trad- ing off the importance of error and complexity.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 97 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

sequential backward selection


Let F = { all features }
We remove one attribute from F as opposed to adding to it, and
we remove the one that causes the least error
j =arg min E (F − xj ) and we
remove xj from F if E (F − xj ) < E (F )

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 98 / 99
Dimensionality Reduction - Feature selection - Subset
Selection

sequential backward selection


We stop if removing a feature does not decrease the error.
To decrease complexity, we may decide to remove a feature if its
removal causes only a slight increase in error.

Ojus Thomas Lee CE Kidangoor CST 322 DATA ANALYTICS Module -2 (Introduction to Data Analytics) 99 / 99

You might also like