Module 1 Notes
Module 1 Notes
categorical variable on
* not
rectangular
no y # # #
find proportions (% )
# total
within
classification Model :
predicting classification of entities one
target calegorical variable
·
as
group w/ highest proportion
-
distribution by
distribution of variable a of variable a 2
taller : Most entities
Within 2 no = x fatter -
> Most 2
entities =2
yes'
I D2 yes
F
overall
F ↑ D 2
larger difference
no
n
·
n
in height ofz between
groups =
More
useful in prediction
of classification
between U
x Y x Y
variable a variable a
confusion Matrix use to evaluate classification Model percentage correctly classified (PCC) (%. )
Predicted
sum of correctly classified
(C predicted y total X 100
total entities
actual
# #
x #
conditional proportion :
of what proportion was correctly predicted ?
actual
-
# # -
PCC within a specific subset of entire data
Chapter two :
classification
pin
·
dot plot
·
box
:
00
fairly symmetric
negativeined :
biModal
..
variable s
Mean
in :
·
neglect outliers
if required
trimodal
(IQR) difference between Va & LQ
anaman
interquartile range =
·
at least 50 % of values no more than IQR value apart
&
·
-
-
~
25 %
standard deviation Measure of:
how much values deviate from
11 the Mean
training-testing approach in
developing classification models -
ensures Model generalises to new data (not just memorising training data
determine best (numeric) variable to use to predict classification of entities within the target categorical Variable
eg. determining variable to use to predict if entily time of day (target variable) is night or
day (group/levels)
bad variable (variable x) no/less overlap
lots of
good
:
overlap variable (variable
y)
-
-
:
day day
·
:iiiiii
: G
night
.......................
variable x
·
if variable y is
Y
less than cut-off
determined
- cutoff value
value,Classic
Dis
developing decision rules -
-
decision rule (for binary outcomes) : if value of (numeric variable G
is less than a cutoff value , classify target variable of entity
as A , else classify the entity as B .
(Where A & B =
will be
levels/groups within target categorical variable Misclassiea
I
lesting
to data
....
19
I
day O
· o O
decision rule :
· cutof
classification
o correct
·
if variable y is less than cut-off
night
value , classify time of day as
incorrect
=
o
&
night ,
else , classify as day
6 00
2
15
no of correct
I I I I I 1 PCC =
total no of values
variable
y
-
·
potential sources of bias
-
algorithmic bias =
data/assumptions in algorithm lead to unfair outcomes
-
sources of bias :
unbalanced of
training data unbalanced ratio entities in one group vs another
·
:
of smaller group
-
focuses too much on correctly classifying larger group at expense
; a
variable)
isnay timprintin
can use middle 95 % (of
training data results) as the
>
- Prediction interval for the model
write interval as : (Min value , Max value
Inumpredictheynumberedtue
Predict accuracy percentage correctly calculated
Enumpredicthey
=
Xten
S
0
create a scatter for both variables
e plot
.
· O
!
There appears
·
to be
negative/positive relationship
o
[variable (c]
S o
a between and
· ....
O *
8
[variable u] . [variable (C) increases
As the ,
the [variable y> tends to
O 60 increase/decrease) .
O 00
Y numerical
response variable
variable
(y)
c
↑
explanatory variable (x)
look at correlation : Measure of association + indicate strength (value) & direction (sign : -/ ) +
rank correlation : +
1/-1 =
perfect positive/negative association between variables
Bo y-intercept
=
y =
Bo +
B , EE .
y & x = variable y/variable va l u e s B, =
Slope/gradient (ise) >
-
neg/pos correlation
relative position
>2 5 .
% = Within Middle 95 %: likely/Usual
-
<2 5 %.
=
Outside Middle <5 %: Unlikely/ unusual
perform simulations generate data from null Model + visualise data & chance Variation
many -
randomisation test :
randomly reallocate data 1000 times
·
random allocation :
randomly allocate 2 units/groups
·
display data using simple Modelling tool
- (observed result(
Middle asy .
-
- --
&
P =
- %. (null hypothesis (