0% found this document useful (0 votes)
2 views53 pages

Understanding Statistical Measures

The document provides an overview of statistical measures, focusing on central tendency (mean and median) and variability (range, quartiles, variance, and standard deviation). It explains the differences between mean and median, the calculation of quartiles, and the concepts of variance and standard deviation as measures of spread in data. Additionally, it discusses covariance and correlation, including their significance in understanding relationships between variables and the properties of multivariate Gaussian distributions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views53 pages

Understanding Statistical Measures

The document provides an overview of statistical measures, focusing on central tendency (mean and median) and variability (range, quartiles, variance, and standard deviation). It explains the differences between mean and median, the calculation of quartiles, and the concepts of variance and standard deviation as measures of spread in data. Additionally, it discusses covariance and correlation, including their significance in understanding relationships between variables and the properties of multivariate Gaussian distributions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Statistical

Measures
Statistical Measures

• Center of the data


• Mean
• Median
• Variation
• Range
• Quart
iles
• Varia
nce
• Stand
ard
Mean or Average or
Expectation
• Traditional measure of center
• Sum the values and divide by the number of values
Mean or
Average
1
(5,6)
11
[(1,2)+
(3,4)+
(6,5)
(5,6)+ Mean
(2,4)+ (5,5)
(1,1)+ (2,4) (3,4) (3.3636,3.0909)
(4,2)+
(6,5)+ (5,3)
(3,1)+
(2,1)+ (2,1) (4,2)
(5,3)+
(5,5)] (1,1) (1,2) (3,1)
Median
(M)
• A resistant measure of the data’s center
• At least half of the ordered values are less than or
equal to the median value
• At least half of the ordered values are greater than
or equal to the median value
• If n is odd, the median is the middle ordered value
• If n is even, the median is the average of the two
middle ordered values
Median
(M)
Location of the median: L(M) =
(n+1)/2 , where n = sample size.

Example: If 25 data values are recorded,


the Median would be the
(25+1)/2 = 13th ordered value.
Median

• Example 1 data: 2 4 6
Median (M) = 4

• Example 2 data: 2 4 6 8
Median = 5
(average of 4 and 6)
• Example 3 data: 6 2
4 Median  2
(order the values: 2 4 6 , so Median = 4)
Comparing the Mean &
Median
• Computation of mean is easier.

• Finding median in higher dimension is much complex.

• Mean is prone to noise.

• The mean and median of data from a symmetric


distribution should be close together. The actual
(true) mean and median of a symmetric distribution
are exactly the same.
Spread or
Variability
• If all values are the same, then they all equal to the
mean. There is no variability.
• Eg: 2, 2, 2, 2, 2, 2; mean = 2
• Variability exists when some values are
different from (above or below) the mean.
• Eg: 10, 15,-20,-22,30, 22
• We will discuss the following measures of
spread: range, quartiles, variance, and standard
deviation
Range
• One way to measure spread is to give the smallest
(minimum) and largest (maximum) values in
the data set;
Range = max  min
• Eg: 10,-2,-7,22,0,11; Range = 22-(-7)=28

• The range is strongly affected by outliers


Quartiles

• Three numbers which divide the ordered data into


four equal sized groups.
•Q1 has 25% of the data below it.
•Q2 has 50% of the data below it. (Median)
•Q3 has 75% of the data below it.
Quartiles Uniform
Distribution

1st Qtr Q1 2nd Qtr Q2 3rd Qtr Q3 4th Qtr


Obtaining the
Quartiles
• Order the data.
•For Q2, just find the median.
• For Q1, look at the lower half of the data values,
those to the left of the median location; find
the median of this lower half.
• For Q3, look at the upper half of the data values,
those to the right of the median location; find
the median of this upper half.
Variance and Standard
Deviation
• Recall that variability exists when some values are
different from (above or below) the mean.
• Each data value has an associated deviation
from the mean:

xi  x
Deviations
•what is a typical deviation from
the mean? (standard deviation)
•small values of this typical
deviation indicate small variability
in the data
•large values of this typical
deviation indicate large variability
in the data
Varianc
e
Variance is the average squared deviation from the mean of
a set of data. It is used to find the standard deviation.
Varian
ce
Mean
Variance

2
-
Variance

2
-

2
-
Variance

1
---------------- 2 2
……… + - + - + ………
No. of Data
Points
Variance Formula
Standard
Deviation

𝑛
1
𝜎 ෍(𝑥 𝑖
=


𝑥ҧ)2
𝑖=1
[ standard deviation = square root of the variance ]
standard deviation = square root of the variance
Variance and Standard
Deviation
Metabolic rates of 7 men (cal./24hr.) :
1792 1666 1362 1614 1460 1867 1439

1792 1666 1362 1614 1460


x 
1867 1439
7
11,200

7
 1600
Variance and Standard
Deviation
Observation Deviations Squared
s deviations
xi xi  x x x 2

1792 17921600 = 192 (192)2 36,864
=
1666 1666 1600 66 (66)2 4,356
= =
1362 1362 1600 = - (-238)2 56,644
238 =
1614 1614 1600 14 (14)2 196
= =
1460 1460 1600 = - (-140)2 19,600
140 =
2
Variance and Standard
Deviation
214,870
 
2

30695.71
7
  30695.71  175.20 calories
Variance (2D)
Variance (2D)
Variance (2D)
Variance (2D)
Variance (2D)

Variance doesn’t explore


relationship between variables
Covariance:measures the direction of
the relationship between two
variables.
Covariance
Covarian
ce
Covariance

𝑦ത
𝑦1 − 𝑦ത<0

𝑦1

𝑥 𝑥ҧ
𝑥1 −
1
Covariance

𝑦
𝑦1 − 𝑦ത
>0
1


� 𝑥1 𝑥1 −
𝑥ҧ >0
ҧ

Covariance

(𝑥𝑖 − 𝑥ҧ)(𝑦𝑖
− 𝑦ത)>0
(𝑥𝑖 − 𝑥ҧ)(𝑦𝑖
− 𝑦ത)<0
Positiv
e
Relatio
n (𝑥𝑖 − 𝑥ҧ)(𝑦𝑖
(𝑥𝑖 − 𝑥ҧ)(𝑦𝑖 − 𝑦ത)<0
− 𝑦ത)>0
Covariance Matrix
𝐶𝑜 σ
𝑣 𝑐𝑜𝑣(𝑥1, 𝑐𝑜𝑣(𝑥1, 𝑥2) ⋯ 𝑐𝑜𝑣(𝑥1,
𝑥1) 𝑥𝑛 )
= 𝑐𝑜𝑣(𝑥2, 𝑐𝑜𝑣(𝑥2, 𝑥2) ⋯ 𝑐𝑜𝑣(𝑥2,
𝑥1) 𝑥𝑛 )
⋮ ⋮ ⋮ ⋮
𝑐𝑜𝑣(𝑥𝑛, 𝑐𝑜𝑣(𝑥𝑛, 𝑥2) ⋯ 𝑐𝑜𝑣(𝑥𝑛,
𝑥1) 𝑥𝑛 )
 Diagonal elements are variances, i.e. Cov(𝑥, 𝑥)=𝑣𝑎𝑟
𝑥 .
 Covariance Matrix is symmetric.
 It is a positive semi-definite matrix.
Correlation

Positive relation Negative relation No relation

• Covariance determines whether relation is positive or negative, but it was impossible to


measure the degree to which the variables are related.
• Correlation is another way to determine how two variables are related.
• In addition to whether variables are positively or negatively related, correlation also tells
the degree to which the variables are related each other.
Correlation
𝑐𝑜𝑣(𝑥, 𝑦)
𝜌𝑥𝑦 = 𝐶𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛
𝑣𝑎𝑟(𝑥)
𝑥, 𝑦 =
𝑣𝑎𝑟(𝑦).

−1 ≤ 𝐶𝑜𝑟𝑟𝑒𝑙𝑎𝑡𝑖𝑜𝑛 𝑥, 𝑦
≤ +1
Multivariate Gaussians (or "multinormal
distribution“ or
“multivariate normal distribution”)
Univariate case: single mean  and
variance 

Multivariate case:
Vector of observations x,
vector of means  and covariance matrix

Dimension of x Determinant
Gaussian distribution

The normal distribution (or Gaussian distribution), also referred as bell curve, is very useful due to the
central limit theorem. Normal distribution states which are average of random variables converge in
distribution to the normal and are normally distributed when the number of random variables is large.

The probability density of the normal distribution can be computed as


Multivariate
Gaussians
Univariate case

Multivariate case

do not depend on x
normalization constants
depends on x and positive
The mean
vector

 μ1 
 μ 2 
μ  E(x)  
.  
. 
 μm


Covariance of two random
variables
•Recall for two random variables xi, xj

  Cov(xi , xj )
2
ij

 E[(xi  i )(x j   j )]
 E(xi x j ) 
E(xi )E(x j )
The covariance
matrix transpose operator
  E[(x  μ)(x 
μ) ]
T

  2 1 12 ..  
(x1  μ1 ) 
   21  22
14
 .  .  
 μ )..(x  μ )]  . 24 
  E  . [(x 1 1 n n   . 
   
 
(xm  μm )
  . .. 2 
  m 

m1 .
. m)=Cov(xm, xm)
Var(x

..
An example: 2
variate case

The pdf of the multivariate will be: Covariance matrix

Determinant
An example: 2
variate case

Factorized into two independent Gaussians!


They are independent!
Recall in general case independence implies uncorrelation
but uncorrelation does not necessarily implies independence.

Multivariate Gaussians is a special case where uncorrelation


implies independence as well.
Diagonal covariance
matrix
If all the variables are independent from each other,
The covariance matrix will be an diagonal one.
Reverse is also true:
If the covariance matrix is a diagonal one they
are independent
  21 Diagonal matrix: m matrix where off-diagonal terms are zero
0 
 0  2 2 

  E[(xi  i )(x j   j )] 
2
ij
0i  j
Gaussian Intuitions: Size
of 

Identity matrix

 = [0 0]  = [0 0]  = [0 0]
 = 2I  = =I
0.6I
As  becomes larger, 
Gaussian becomes more spread out
Gaussian Intuitions: Off-
diagonal

As the off-diagonal entries increase, more correlation between value of x and value of y
Gaussian Intuitions: off-diagonal and
diagonal

• Decreasing non-diagonal entries (#1-2)


• Increasing variance of one dimension in diagonal (#3)

You might also like