0% found this document useful (0 votes)
8 views111 pages

Chapter 4 Data Analytics

Chapter 4 focuses on Data Analytics within the ICT Management course, covering topics such as data visualization and machine learning. The course aims to provide a broad overview of the data analytics landscape and its business applications, emphasizing the importance of data-driven decision-making. Practical sessions are scheduled for hands-on experience with data analytics exercises.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views111 pages

Chapter 4 Data Analytics

Chapter 4 focuses on Data Analytics within the ICT Management course, covering topics such as data visualization and machine learning. The course aims to provide a broad overview of the data analytics landscape and its business applications, emphasizing the importance of data-driven decision-making. Practical sessions are scheduled for hands-on experience with data analytics exercises.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 4.

Information Management

Boje Deforce
Prof. Dr. Estefanía Serral Asensio
ICT Management

MODULE I: MODULE II: MODULE III:


Introduction Modelling Advanced
to ICT technologies
Management

Introduction Process and Data Analytics,


to ICT Data IoT
Perspective
Chapter 1 Chapters 4, 5
Chapters 2, 3
Chapter 4

Data Analytics
Theory session
Who am I?
• Boje Deforce
• PhD student at LIRIS with:
• Prof. Dr. Estefanía Serral Asensio
• Prof. Dr. Bart Baesens
• Prof. Dr. Jan Diels
• I work on Machine/Deep Learning for:
• IoT sensor data (i.e. time-series)
• Forecasting 📈
• Anomaly detection ❌
• Other:
• Worked at Deloitte as data scientist for ~2 years
• Have a start-up ([Link])

3 FEB, Campus Brussels


Course Overview

1. Introduction to ICT Management

2. Business Process Management

3. Information Management

4. Data Analytics

5. The Internet of Things

4 FEB, Campus Brussels


Chapter 4: Overview

4.1 Introduction to data analytics


4.2 Data visualization
4.3 Machine learning

5 FEB, Campus Brussels


Course Material

• Course material for this chapter:


• Slides
• Exercises

6 FEB, Campus Brussels


Chapter Planning

Date Topic Modality


Friday, April 26th Data Analytics LIVE / ON CAMPUS

Monday, April 29th Data Analytics exercises (G1) LIVE / ON CAMPUS

Friday, May 3rd Data Analytics exercises (G2) LIVE / ON CAMPUS

7 FEB, Campus Brussels


Some practical details
• Primary goal of the course:
• Provide a broad overview of the data analytics landscape
• Give you insight in how analytics is beneficial for business
• Provide you with a basic toolbox to perform analytics
• Practical:
• Today - 26th of April: theory session
• Monday – 29th of April: Practice session Group 1
• Friday – 3rd of May: Practice session Group 2
• Content:
• Some slides are based on Prof. dr. Van den Broucke’s and Dr. Jari
Peeperkorn’s slides

8 FEB, Campus Brussels


4.1 Introduction
The what/why/how of data analytics

9 FEB, Campus Brussels


2024: The AI race is on

• Live demo: [Link]

[Link]
10 FEB, Campus Brussels
The real (business) value of LLMs

11 FEB, Campus Brussels


Data Analytics – what is it?

• Data Analytics is the management of data for all


uses (operational and analytical) and the
analysis of data to drive business processes
and improve business outcomes through
more effective decision making and
enhanced customer experiences.

12 [Link] FEB, Campus Brussels


13 FEB, Campus Brussels
Analytics is everywhere…

…and so much more

14 FEB, Campus Brussels


Some analytics examples

Added
Some Analytics
business
Information process value

Forecast

Price
Prediction

15 FEB, Campus Brussels


Some analytics examples

Added
Some Analytics
business
Information process value

Hire a
candidate:
yes/no

16 FEB, Campus Brussels


Data science lifecycle

Source: datacamp

80% 20%
What makes a “good” data analyst/scientist?
• Traits
• “Statistical thinking”
• “Problem solving skills”
• “Creativity”
• “Communication and storytelling skills”
• “Business intuition”
• …

• Knowledge
• Dashboarding
• Technical knowledge (DS)
• (Basic) data engineering knowledge Source: Verbeke et al. (2018)

• Programming (DS)
• …
4.1 Introduction
How to get from raw data to added business value?

19 FEB, Campus Brussels


Data-analytical thinking
• How to extract useful knowledge from data to improve business decision-
making?

• Data should be considered an asset and we need to consider how to invest

• Analytics from a business perspective


• Business problem → goal(s) + constraints on its solution(s)
• Data and domain knowledge → raw materials
• Data analytics → frameworks/processes for decomposing the problem into
subproblems, as well as tools and techniques for solving them
Data-analytical thinking
• Framework to systematically extract useful knowledge from data to improve business
decision-making: CRISP-DM

• CRoss Industry Standard Process for Data Mining = CRISP-DM


• Structured thinking
• Fundamental concepts and principles
• Important not only for data scientists themselves, but for anyone working with
data scientists, employing data scientists, investing in data-heavy ventures, or
directing the application of analytics in an organization (i.e., spotting DS
opportunities)
CRISP-DM

[Link], B. van, Herhausen, D. & Fahse, T. Overcoming the pitfalls and perils of algorithms: A *no need to know the
classification of machine learning biases and mitigation methods. J Bus Res 144, 93–106 (2022). biases

22 [Link] FEB, Campus Brussels


CRISP-DM alternative
• Article to read: Getting started with AI? Start here! – by Cassie Kozyrkov
• [Link]
Getting started with AI? Start here! – Cassie Kozyrkov
• Step 1: Figure out who’s in charge
• Decision skills and deep understanding of the business

• Step 2: Identify the use case


• Focus on the outputs
• “Imagine your ML/AI system is operational and ask yourself if you’re happy you sunk
company resources into making it. No? Keep brainstorming. Better to discover no one
needs your application before several PhDs waste their lives on it.”
• Now is not the time for inputs
• Reason 1: Missed opportunities
• “To some people, data is data. It’s all the same.”
• Reason 2: Tacit agreement
• “So here’s the tragicomedy: when you’ve spent the last 6 hours arguing with your buddies about
whether or not variable x (raw, standardized, or normalized?) is a good input with suitable
logging for predicting output y, you’ve, ahem, normalized the idea that y is worth pursuing. You
stop questioning the point of working on y in the first place and end up building things that don’t
need to be built.”
Getting started with AI? Start here! – Cassie Kozyrkov

“Still struggling to find a use case? Consider pausing ML/AI in favor of analytics for a while. The
goal of analytics is to generate inspiration for the decision maker.”

• DS = data-driven decision making

• Type 1 problem – exploratory analysis


• Insights based on “discoveries” → tactical and strategic decision making
• Type 2 problem – supporting operational decisions
• Decisions that repeat, especially at a massive scale, and so decision-making can
benefit from even small increases in decision-making accuracy based on data
analytics
• Supporting vs. automating
Getting started with AI? Start here! – Cassie Kozyrkov

• Step 3: Do some reality checks


• Data about this business problem?
• Do we have a plan to get the data soon?

• Step 4: Craft a performance metric wisely


• Figure out how to trade off various results on one single output
• Loss function ≠ performance metric
• “You’re free to optimize using a standard loss function that moves in the same
direction as the function your leader’s imagination just spawned.”
Getting started with AI? Start here! – Cassie Kozyrkov

• Step 5: Set testing criteria to overcome human biases


• Define your population of interest
• Decide on the minimum performance you’re willing to sign off on
• “When humans invest time and effort into something, we fall in love with what we have
made… even if it is a pile of poisonous rubbish.”
• “Better than human?” vs. “Is it good enough to be useful?”
Analytics in Business
[Link]

• Analytics should not just be used because it sounds cool (e.g. everything is AI
nowadays)
→ It should serve/improve the business outcome
It should be used to discover patterns in the data which are:
• Valid
• Useful
• Unexpected
• Understandable

28 FEB, Campus Brussels


[Link]
by-using-chatgpt
Analytics in Business
• Valid: hold on new data with some certainty, i.e. generalizable (see more later)
• Over time, seasonal effects, overfitting, sub-groups, regional differences…
• Useful: should be possible to act on the item, i.e. actionable
• Business question, implementation, maintenance costs, ease-of-use…
• Unexpected: non-obvious to the system, i.e. interesting
• Balance between trust and discovery… Big and “weird” data
• Understandable: humans should be able to interpret the pattern
• Black box vs. white box, trust, validity…

29 FEB, Campus Brussels


Analytics, data science, decision science, AI?

• The goal matters:


• Exploratory analytics
• Descriptive analytics
• Predictive analytics
• Prescriptive analytics
• Depending on who you ask, especially outside of academia:
• Data Mining ≈ Big Data ≈ Data Analytics ≈ Data Science ≈ Knowledge
Discovery ≈ Artificial Intelligence ≈ Deep Learning ≈ Decision Science ≈
Machine Learning

30 FEB, Campus Brussels


What’s the goal?
• Exploratory analytics: a bird’s-eye view of your data, think: “what do we have here”
• Examples: plots (outliers?), distributions, quick charts, basic correlations… very
visual
• Descriptive analytics: describe the data, relations between the data, …
• Examples: mean/median (!), variance, standard deviation, percentiles, value count
• Also: basic linear regression, decision trees (see later)
• Predictive analytics: predict a target measure of interest
• Examples: regression, classification, … → predicting stock prices, bank fraud,
sentiment analysis, …
• Also consider whether your goal is really predictive?
• Prescriptive analytics: “what should I do?”
• Examples: What-if analysis on a trained supervised model
• Or using “good old” operations research

31 FEB, Campus Brussels


Purpose is key

• “Can data analysis help in solving the business question/problem?”


• “Which type of analysis will best solve the business question/problem?”

[Link]

32 FEB, Campus Brussels


Data comes in all types and shapes

Structured data Unstructured data

33

(Semi-structured data)

33 FEB, Campus Brussels


Structured data

34 FEB, Campus Brussels


Structured data
• A tabular data set (“structured data”):
• Has instances (examples, rows, observations, customers, cases, …)
• And features (attributes, fields, variables, predictors, covariates,
explanatory variables, regressors, independent variables)
• These features can be:
• Numeric (continuous)
• Categorical (discrete, factor), either nominal (binary as a special case) or ordinal
• New features can be created through “featurization”:
• I.e. creating new features from existing features
• Target (label, class, dependent variable, response variable) can also be
present
• Numeric, categorical, …

35 FEB, Campus Brussels


Unstructured data

• Basically, all data that has no tabular structure to it


• (except for semi-structured data)
• Can also be turned into structured data through featurization
• For example:

Contains
Twitter_ID Text_length picture #_likes #_retweets #_comments
pickover 12yes 8500 1500 97
… … … … … …

36 FEB, Campus Brussels


From here onwards

• Explore the surface-layers of:


4.2 Data Visualization through dashboards (often referred to as business
intelligence)
~ Exploratory analytics
~ Descriptive analytics
4.3 Machine Learning
~ Predictive analytics
~ Prescriptive analytics

37 FEB, Campus Brussels


4.2 Data Visualization
Dashboards

38 FEB, Campus Brussels


Dashboards - purpose

• Often, visualization can already get you a longways (cf. Rule #1 from Google)
• A very interactive way to understand the data that a business has available

39 FEB, Campus Brussels


Example

40 FEB, Campus Brussels


Dashboards (≈ Business intelligence)

• “Business intelligence is an umbrella term that includes the applications,


infrastructure and tools, and best practices that enable access to and analysis
of information to improve and optimize decisions and performance.” (Gartner)
• Objective: to transform data into meaningful information.
Enable interactive access to data and provide business managers and data
analysts the ability to conduct data analyses and derive better business
decisions

41 FEB, Campus Brussels


Some problems to solve with dashboards

• How much have we sold the past year?


• Who is the most profitable customer?
• What are the area coverage levels?
• …
• What is our estimated sales for next month?
• “What if” we lose our biggest customer

42 FEB, Campus Brussels


4.3 Machine Learning
An exploration of the surface

43 FEB, Campus Brussels


What to know as a BBA-student?

• A 3rd year BBA-student should at least be able to understand the importance


of the following elements in order not to be fooled by data scientists. That is:
• What is good data, which data do I need? “In Phoenix, the problem was particularly acute. 9
in 10 homes Zillow bought were put up for sale
• How does ML work? at a lower price than the company originally
bought them”
• How to identify a good model?
• From model to business value
• Importance of model monitoring
• Failed ML projects
[Link]
business-worlds-love-affair-ai/

44 FEB, Campus Brussels


Machine Learning - What is it?
• Machine Learning (and Artificial Intelligence) is about letting the machine
learn from data. We distinguish between different learning paradigms:
• Supervised learning
• Target variable available
• Relate input variables (X) to target (Y)
by learning a mapping (f)
• Unsupervised learning
• No target variable (Y) available
• Find structure, patterns in data
• (semi-supervised learning)
• (self-supervised learning)
• (reinforcement learning)

45 FEB, Campus Brussels


Supervised learning - example

[Link]
46 FEB, Campus Brussels
learning
Supervised learning - example

𝑋 𝑌
Learn mapping
𝑓 𝑋
Parking occupancy
Full
Learning algorithm Half-full
Empty
Empty
Half-full

Feedback loop

238 78 67 239 35 198


238
238 78
78 67
67239
239 35
35198
198
48 41 137 7 93 181
48
48 41
41137
137 77 9393181
181
92 160 212 161 21 220
92 160 212 161 21 220
92 160 212 161 21 220
62 240 195 26 115 54
62
62240
240195
195 26
26115
115 54
54
69 239 165 50 202 6
69
69 239 165 50 202 66
239 165 50 202
21 188 173 65 182 69
21
21188
188173
173 65
65182
182 69
69
What the computer sees

47 FEB, Campus Brussels


Supervised learning – What is it?
• “In machine learning and artificial intelligence, supervised learning refers to a
class of systems and algorithms that determine a predictive model using
data points with known outcomes. The model is learned by training through
an appropriate learning algorithm that typically works through some
optimization routine to minimize a loss or error function.”

• class of systems and algorithms: It’s an umbrella term


• data points with known outcomes: You need an input linked to an output
• appropriate learning algorithm: An algorithm that can learn this link
• works through some optimization routine: there is a feedback loop in the
process

48 FEB, Campus Brussels


Supervised learning – more detailed
• Regression: continuous label
• Classification: categorical label
• For classification:
• Binary classification (positive/negative outcome)
• Multiclass classification (more than two possible outcomes but case only belongs
to one)
• Ordinal classification (target is ordinal)
• Multilabel classification (multiple outcomes to which a case can belong)
• For regression:
• Absolute values
• Delta values
• Quantiles regression
• Single versus multi-output models is possible as well

49 FEB, Campus Brussels


Supervised learning types - examples

• Other problems where supervised learning can offer a solution:


• Predict who will default on a loan
• Predict whether a video contains
Classification problem
political/harmful content or not
• Predict whether will go up or down

• Predict how much a house would sell for


• Predict the temperature from sensor-data Regression problem

• Predict how much revenue will produce

50 FEB, Campus Brussels


Supervised learning - algorithms

• Linear regression
• Logistic regression
• K-Nearest neighbors
• Decision trees
• Random forest
• Naïve Bayes
• Neural networks
• …

51 FEB, Campus Brussels


Supervised learning - algorithms

• Linear regression
Will be covered in the Msc
• Logistic regression
• K-Nearest neighbors
This course
• Decision trees
• Random forest
• Naïve Bayes A bit more advanced

• Neural networks
• …

52 FEB, Campus Brussels


Supervised learning – lifecycle

2. Data wrangling & exploratory data analysis (EDA) ~ exploratory/descriptive analytics

3. Training
4. Evaluation
5. Testing
6. Evaluation
7. Deployment
8. Monitoring

53 FEB, Campus Brussels


Exploratory/descriptive analysis

Supervised learning – (2) Data wrangling


• Clean the data (e.g. age: 540) “80% of data analysis is data wrangling”
• Understand data distributions ~ bias in data? (e.g. 90% men vs. 10% women)
• Look at correlations between target and input variables
• We want correlated variables
• …
• Split the data into:
• training data → used to train the model on
• validation data → used to assess different models/model configurations
• test data → acts like real world unseen data, how will our model perform in the
open world?
• For simplicity, we will only work with training and validation data in the examples,
just know that in real-life examples, having a final test set is important!

54 FEB, Campus Brussels


Live interactive

• [Link]

55 FEB, Campus Brussels


Predictive analysis

Supervised learning – (3) Training

3. Training:
• Learn a mapping 𝑓 that can map an input 𝑋 to
an output 𝑌 using a learning algorithm
• In practice, typically multiple learning algorithms are applied and the best one is
selected
• What is the best one? For that, we need a formal evaluation (see next slide)
• “Learning” is achieved by minimizing some objective function through an
optimization routine (typically iterative)
• See more in the example

56 FEB, Campus Brussels


Predictive analysis

Supervised learning – (4) Evaluation


• Evaluation metrics allow for formal evaluation of a learning algorithm
• Choosing an evaluation metric depends on the use case
• We want to know how well our model did on the training data, but more importantly, how well it can
generalize to unseen data (cf. validation and test data)
• Common evaluation metrics:
• For classification:
𝑇𝑃+ 𝑇𝑁 𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑙𝑦 𝑐𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑒𝑑
• 𝐴𝑐𝑐 = =
𝑃+𝑁 𝐴𝑙𝑙 𝑑𝑎𝑡𝑎
• Beware of imbalanced data (see exercise session)
• Confusion matrix (see further + exercise session)
• ROC-curve
• …
• (For regression:)
1 2
• 𝑀𝑆𝐸 = σ𝑛𝑖=1 𝑌𝑖 − 𝑌෡𝑖
𝑛
1
• 𝑀𝐴𝐸 = σ𝑛𝑖=1 |𝑌𝑖 − 𝑌෡𝑖 |
𝑛
• …
[Link]

57 FEB, Campus Brussels


Predictive analysis

Supervised learning – (4) Evaluation


• It’s “easy” to obtain a very high score on the training data
• Some algorithms can even just “memorize” the training data
• What we’re really interested in, is performance on unseen data
• As such, we need to strike the balance between underfitting & overfitting

E.g. learning a
classifier that can
separate between
the red and blue
class

Live

58 FEB, Campus Brussels


Predictive analysis

Supervised learning – (5) Testing + (6) final evaluation


• Does my model really do what I want it to do on unseen data?
• Am I sure that I am not accidentally overfitting on the validation data?
• Continuously checking out performance of a model on the validation data
can cause one to overfit on validation data too!
• This is key to useful ML and AI

59 FEB, Campus Brussels


~Prescriptive analysis

Supervised learning – (7) Deployment


• Represent model output in a user-friendly
way (even in a dashboard is possible)
• Integrate with existing business tools and
decision engines
• In real-world ML systems, only a small
fraction is comprised of actual ML code
• There is a vast array of surrounding
infrastructure and processes to support
their evolution
• Technical debt: data dependencies, model
complexity, reproducibility, changing world
(e.g. COVID), changing decisions, …

60 FEB, Campus Brussels


Supervised learning – (8) Monitoring

• Continuously monitor model output


• Contrast model output with observed numbers
~
• Concept drift: the target definition (and statistical properties thereof) changes
• Data drift: Statistical properties of input data change

• Technically missing from CRISP-DM (also because it’s a more recent phenomenon)

61 FEB, Campus Brussels


1. Get data (see previous classes)
2. Data wrangling & EDA
3. Training
4. Evaluation
5. Testing
6. Evaluation
7. Deployment
8. Monitoring

Full example

62 FEB, Campus Brussels


Let’s look at an example
• Imagine, we own a wine business and need index Alcohol % color_intensity good_wine
to rate wine. Can we use ML for this? 0
1
14.23
13.20
5.64
4.38
0
1
2 13.16 5.68 0
• What type of problem? 3 14.37 7.80 0
4 13.24 ? 0
→ classification problem ...
125
...
12.07
...
2.76
...
1

• Find a function that maps the input to some


126 12.43 3.94 0
127 11.79 3.00 0
128 225.36 2.12 1
output 129 12.04 2.60 1

• Input: alcohol % and color_intensity of


the wine
• Output: good wine (1) or bad wine (0)

Data obtained and adapted from:


63 Lichman, M. (2013). UCI Machine Learning Repository [[Link] Irvine, CA: University of California, School of Information and FEB, Campus Brussels
Computer Science. – inspired by Google AI lectures from Cassie Kozyrkov
2. Data wrangling & EDA

Data wrangling

64 FEB, Campus Brussels


Exploratory/descriptive analysis

Let’s look at an example


• Data wrangling & EDA: index alcohol color_intensity good_wine

• Missing values: 0
1
14.23
13.20
5.64
4.38
0
0
2 13.16 5.68 0
• “Impute” with: 3 14.37 7.80 0
4 13.24 ? 0
• Mean: if data is not skewed ... ... ... ...
• Median: if data is skewed (more robust) 125 12.07 2.76 1
126 12.43 3.94 1
• Mode: for categorical data 127 11.79 3.00 1
128 225.36 2.12 1
• Extreme values: 129 12.04 2.60 1

• Investigate cause:
• Keep if valid?
• Else remove
• Or impute
• Explore data:

65 FEB, Campus Brussels


Exploratory/descriptive analysis

Let’s look at an example


• Data wrangling & EDA: index alcohol color_intensity good_wine

• Missing values: 0
1
14.23
13.20
5.64
4.38
0
0
2 13.16 5.68 0
• “Impute” with: 3 14.37 7.80 0
4 13.24 ? 0
• Mean: if data is not skewed ... ... ... ...
• Median: if data is skewed (more robust) 125 12.07 2.76 1
126 12.43 3.94 1
• Mode: for categorical data 127 11.79 3.00 1
128 225.36 2.12 1
• Extreme values: 129 12.04 2.60 1

• Investigate cause:
• Keep if valid
• Else remove
• Or impute
• Explore data:

66 FEB, Campus Brussels


Predictive analysis

Let’s look at an example

Train-validation-(test) split:

80% of data used for training

20 % of data used for validation

(typically also a final test set)

67 FEB, Campus Brussels


Predictive analysis

Let’s look at an example

Train-validation-(test) split:

80% of data used for training

20 % of data used for validation

(typically also a final test set)

68 FEB, Campus Brussels


Predictive analysis

Let’s look at an example

Train-validation-(test) split:

80% of data used for training

20 % of data used for validation

(typically also a final test set)

69 FEB, Campus Brussels


3. Training
4. Evaluation
5. Testing
6. Evaluation

Train-test cycle

70 FEB, Campus Brussels


Let’s look at an example

Algorithm selection:

K-nearest neighbors

Decision tree

71 FEB, Campus Brussels


Let’s look at an example

Algorithm selection:

K-nearest neighbors

Decision tree

72 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm - intuition

• Basic intuition:
• If we have a new wine with alcohol %=12.0 and color_intensity=2.6
it is very likely that it belongs to the neighboring class

73 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm - intuition

• Basic intuition:
• If we have a new wine with alcohol %=12.0 and color_intensity=2.6
it is very likely that it belongs to the neighboring class
3. Classify
Training 2. Compute
Unseen
instances Distance
instance

1. Choose k (e.g. 5)
of the “nearest” 74

instances
74 FEB, Campus Brussels
Predictive analysis

K-nearest neighbors algorithm - intuition

• What if we have a new wine with alcohol %=12.8 and


color_intensity=3.6? 3. Classify
Unseen
instance
based on
Training 2. Compute majority
instances Distance vote

1. Choose k (e.g. 5)
of the “nearest” 75

instances
75 FEB, Campus Brussels
Predictive analysis

K-nearest neigbors algorithm - training


• Problems at the boundaries:
• For example, set k=3:
• Look at 3 nearest neighbors with a
specified distance metric
(commonly Euclidean distance)
• Blue: 2
• Orange: 1
→Classify unseen instance as Good
K=3 wine
K=5
• Observe that e.g. k=5 yields the
opposite result
• What if k=2?
• Choosing optimal K requires training

76 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm – different values for K

77 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm – different values for K


Overfit Just right

Underfit

78 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm – different values for K

• For K=1 we perform well on the


training data but not so well on
unseen data (i.e. validation data)
• As K increases (we take into
account more
neighbors/information), training
accuracy goes down and validation
accuracy goes up
• Find the sweet spot (i.e. k=4)

79 FEB, Campus Brussels


Predictive analysis

K-nearest neighbors algorithm – final thoughts


• To further improve:
• Normalize the data (especially when features are on very different scales)
• Remove outliers
• …
• Advantages:
• Lazy learner: KNN actually does not learn anything. It simply stores the data and
computes distances to the data at prediction time. Making “training” very fast
• Can easily add new data
• Easy algorithm: only two parameters, namely K (and the distance function)
• Disadvantages:
• Lazy learning becomes expensive for large datasets
• Suffers from curse of dimensionality
• Sensitive to noise in the data

80 FEB, Campus Brussels


Let’s look at an example

Algorithm selection:

K-nearest neighbors

Decision tree: live

81 FEB, Campus Brussels


Predictive analysis

Decision trees - intuition

• Basic intuition:
• Partition the data into leafs that are “pure”. I.e. can we split our data such
that we get subgroups that only contain one class?
1. Select feature to split
on (e.g. at random)
2. Split on feature
3. Calculate measure of
“purity” in resulting
subsets
4. Restart from 1. until
all subsets are “pure”
or until stopping
criterion is reached 82

82 FEB, Campus Brussels


Predictive analysis

Decision trees - intuition

• Basic intuition:
• Partition the data into leafs that are “pure”. I.e. can we split our data such
that we get subgroups that only contain one class?
1. Select feature to split
on (e.g. at random)
Alcohol < 12.8? 2. Split on feature
yes no

3. Calculate measure of
Color_intensity
Good wine < 5? “purity” in resulting
subsets
4. Restart from 1. until
{Bad: 13, Good: 8} → Bad Bad wine
all subsets are “pure”
or until stopping
criterion is reached 83

83 FEB, Campus Brussels


Decision trees – some terminology
• Root Node: represents the entire population or sample and this further gets divided
into two or more homogeneous sets.
• Splitting: a process of dividing a node into two (or more) sub-nodes, typically through
binary splits.
• Decision Node: When a sub-node splits into further sub-nodes, then it is called the
decision node.
• Leaf / Terminal Node: Nodes that do not further split are called leaf or terminal
nodes.
• Pruning: When we remove sub-nodes of a decision node, this process is called
pruning.
• Branch / Sub-Tree: A subsection of the entire tree is called branch or sub-tree.
• Parent and Child Node: A node, which is divided into sub-nodes is called a parent
node of sub-nodes whereas sub-nodes are the child of a parent node.

84 [Link] FEB, Campus Brussels


Decision trees – some terminology

85 [Link] FEB, Campus Brussels


Predictive analysis

Decision trees – it’s (almost) all about the split

• Selecting the best split


• Degree of impurity:
• “The more pure leaf nodes of a split become, the better”
• The lower the impurity of a split, the better
• Two popular impurity measures*
• Gini impurity:

• Entropy:

*not to know by heart for exam

86 FEB, Campus Brussels


Predictive analysis

Decision tree - training

• For KNN we only had 2 hyperparameters (only 1 really) to set (i.e. K &
distance metric)
• For decision trees a lot more hyperparameters are available. Below a selection
of the most important ones:
• Splitting criterion: the impurity measure used to split on (we will use Gini)
• Max depth: how deep we let the tree grow*
• Min_samples_split: how many samples a leaf should contain for it to be
split further
•…

*this is also often called a stopping criterion since the tree stops splitting after the max depth has been reached

87 FEB, Campus Brussels


Predictive analysis

Decision trees – classifying an unseen case

• If we have a new wine with alcohol %=12.0 and color_intensity=2.6,


we just follow the learned tree

Alcohol < 12.8?


yes no

Color_intensity
Good wine < 5?

{Bad: 13, Good: 8} → Bad Good wine

88 FEB, Campus Brussels


Predictive analysis

Decision tree - training

• We’ll use the “Gini impurity” as splitting criterion

2 2
48 0
𝐺𝑖𝑛𝑖(𝑇1 ) = 1 − +
Decision Tree 48 48
𝑇2
=0
𝑇1 2 2
8 48
𝐺𝑖𝑛𝑖(𝑇2 ) = 1 − +
56 56
= 0.245
48 56
𝐺𝑖𝑛𝑖(𝑇1 , 𝑇2 ) = 𝐺𝑖𝑛𝑖 𝑇1 + 𝐺𝑖𝑛𝑖(𝑇2 )
104 104
= 𝟎. 𝟏𝟑𝟐

89 FEB, Campus Brussels


Predictive analysis

Decision tree - training

• We’ll use the “Gini impurity” as splitting criterion

2 2
48 3
𝐺𝑖𝑛𝑖(𝑇1 ) = 1 − +
Decision Tree
51 51
= 0.111
𝑇2
2 2
8 45
𝐺𝑖𝑛𝑖(𝑇2 ) = 1 − +
53 53
= 0.256
𝑇1 51 53
𝐺𝑖𝑛𝑖(𝑇1 , 𝑇2 ) = 𝐺𝑖𝑛𝑖 𝑇1 + 𝐺𝑖𝑛𝑖(𝑇2 )
104 104
= 𝟎. 𝟏𝟖𝟓

90 FEB, Campus Brussels


Predictive analysis

Decision tree - training

• We’ll use the “Gini impurity” as splitting criterion

𝑇2
𝑇2
𝑇1
less impure than

So we take the left 𝑇1


split as first split

91 FEB, Campus Brussels


Predictive analysis

Decision tree – different tree depths

• For depth=1 the tree underfits


• As depth increases (the tree can
split the data into “purer” nodes),
training accuracy goes up and
validation acc. goes down (i.e. we
overfit)
• Find the sweet spot (i.e. depth=2)
• Since we “cut back on the
branches”, for trees, we call
this tree pruning

92 FEB, Campus Brussels


Predictive analysis

Decision tree – different tree depths

93 FEB, Campus Brussels


Predictive analysis

Decision tree – different tree depths


Underfit Just right

Overfit

94 FEB, Campus Brussels


Decision TreePredictive analysis

Decision tree – different tree depths


• Tree depth=6
• This tree is overfitting and
performs poorly on new data
• Overfitting: model learns detailed
characteristics that are particular to
the specific training set and – as a
result – the model cannot
generalize well to unseen data

95 FEB, Campus Brussels


Decision Tree

Decision tree – different tree depths


• Tree depth=1
• This tree is underfitting and
performs poorly on training data
(and also on new data)
• Underfitting: model is not
sophisticated enough to learn
patterns present in the data

96 FEB, Campus Brussels


Predictive analysis

Decision tree – different tree depths

Sweet spot

Underfitting Overfitting

97 FEB, Campus Brussels


Predictive analysis

Decision trees – final thoughts

• Advantages:
• Learning and using trees is very efficient
• They tend to have good predictive accuracy (on smaller datasets), even
better when combined into ensembles (on larger datasets, i.e. random
forests - see elsewhere)
• They are interpretable: we can understand what the tree says, and we can
explain its predictions
• Disadvantages:
• Trees are sensitive to high variance in the data
• Harder to learn from large datasets or very noisy datasets

98 FEB, Campus Brussels


Predictive analysis

Supervised learning – (4) Evaluation


• Evaluation metrics allow for formal evaluation of a learning algorithm
• Choosing an evaluation metric depends on the business case (cf. Cassie Kozyrkov)
• We want to know how well our model did on the training data, but more importantly, how well it did
on unseen data (i.e. in the real world)
• Common evaluation metrics:
• For classification:
𝑇𝑃+ 𝑇𝑁 𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑙𝑦 𝑐𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑒𝑑
• 𝐴𝑐𝑐 = =
𝑃+𝑁 𝐴𝑙𝑙 𝑑𝑎𝑡𝑎
• Beware of imbalanced data
• Confusion matrix (see later)
• ROC-curve
• …
• For regression:
1 2
• 𝑀𝑆𝐸 = σ𝑛𝑖=1 𝑌𝑖 − 𝑌෡𝑖
𝑛
1
• 𝑀𝐴𝐸 = σ𝑛𝑖=1 |𝑌𝑖 − 𝑌෡𝑖 |
𝑛
• …
[Link]

99 FEB, Campus Brussels


Predictive analysis

Evaluation - How to choose our final algorithm?

• Classification (regression) problem requires classification (regression)


evaluation metrics
𝑇𝑃+ 𝑇𝑁 𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑙𝑦 𝑐𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑒𝑑
• 𝐴𝑐𝑐 = =
𝑃+𝑁 𝐴𝑙𝑙 𝑑𝑎𝑡𝑎
• Beware of imbalanced data Predicted label

• Confusion matrix:
N P
True False
N Negative Positive
(TN) (FP)
True label
False True
P Negative Positive
(FN) (TP)

• ROC-curve:

[Link]

100 FEB, Campus Brussels


Predictive analysis

Evaluation - How to choose our final algorithm? Predicted label


N P
True False
N Negative Positive
𝑇𝑃+ 𝑇𝑁 𝑐𝑜𝑟𝑟𝑒𝑐𝑡𝑙𝑦 𝑐𝑙𝑎𝑠𝑠𝑖𝑓𝑖𝑒𝑑
• 𝐴𝑐𝑐 = = True label
(TN)
False
(FP)
True
𝑃+𝑁 𝐴𝑙𝑙 𝑑𝑎𝑡𝑎
P Negative Positive
• KNN Accuracy: (FN) (TP)
49+47
• Train-set: 𝐴𝑐𝑐 = 56+48 = 0.92
14+10
• Test-set: 𝐴𝑐𝑐 = = 0.92
15+11
• Decision tree Accuracy:
52+48
• Train-set: 𝐴𝑐𝑐 = = 0.96
56+48
14+11
• Test-set: 𝐴𝑐𝑐 = = 0.96
15+11

• Beware of imbalanced data


• Imagine a dataset with 96 N and 4 P
• We train a classifier with 96% accuracy
• Is this a good classifier?
• (see also exercise session)

101 FEB, Campus Brussels


Predictive analysis

Evaluation - How to choose our final algorithm?


• 𝑅𝑂𝐶-curve and 𝐴𝑈𝐶_𝑅𝑂𝐶:
• Live
• Live 2
• Good tool to compare models
• Model with highest area under the
curve “wins” (most of the time)
• Can be used to adapt a model for
dealing with unbalanced data (see
elsewhere)
• The AUC of a classifier is equivalent to
the probability that the classifier will
rank a randomly chosen positive
instance higher than a randomly
chosen negative instance.

102 FEB, Campus Brussels


Predictive analysis

Note on Cross-Validation

103 [Link] FEB, Campus Brussels


7. Deployment
8. Monitoring

Train-test cycle

104 FEB, Campus Brussels


Pitfalls of ML

• Data leakage: leaking part of your validation/testing data to your training phase
(e.g. providing data from the same year as your training data)
• Bias in data: beware of biased data, the model will exploit such bias without
sorry
• Model drift: [Link]
learning-models/ and [Link]
datasets

105 FEB, Campus Brussels


Extra resources
This is not part of the course material
Just for those interested

106 FEB, Campus Brussels


EXTRA: MLOps

107 FEB, Campus Brussels


EXTRA: Note on unsupervised learning

• If there is no target variable available, we typically use unsupervised learning


to detect patterns in the data.
• K-means
• Use cases

108 FEB, Campus Brussels


EXTRA: word2vec demo

• How to represent words to a machine?


• [Link]

109 FEB, Campus Brussels


The data science team

[Link]

110 FEB, Campus Brussels


Extra study material

• For many years, people made all sorts of lists with “5 free resources to get
better at ML, AI, …”. Recently, tensorflow (Google’s deep learning library)
summarized this into a nice list of how to get “better” in data science with many
free resources from different sources. See here.
• Datacamp also offers many courses (only the first lessons are free)
• [Link]
• Bit more advanced:
• [Link]
• Overview of useful algorithms etc.: [Link]
ml !!

111 FEB, Campus Brussels

You might also like