0% found this document useful (0 votes)
3 views140 pages

Multivariate

The document provides an overview of multivariate analysis, highlighting its importance in understanding complex data involving multiple variables. It discusses various statistical methods, including covariance, correlation, and canonical correlation, as well as the challenges of visualizing high-dimensional data. The content emphasizes the need for dimension reduction techniques to extract meaningful information from large datasets.

Uploaded by

amit biswas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views140 pages

Multivariate

The document provides an overview of multivariate analysis, highlighting its importance in understanding complex data involving multiple variables. It discusses various statistical methods, including covariance, correlation, and canonical correlation, as well as the challenges of visualizing high-dimensional data. The content emphasizes the need for dimension reduction techniques to extract meaningful information from large datasets.

Uploaded by

amit biswas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Multivariate Analyses

Amit K Biswas

Indian Statistical Institute

@ PGDSMA 2023 − 24

Stat Methods II

January 20, 2024

Amit (ISI, Chennai) Multivariate January 20, 2024 1 / 70


Introduction

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 2 / 70


Introduction

Univariate → Multivariate : changes

While dealing with a single variable X we denote by xi the i th


observation of the variable and have it’s mean µ & variance
σ2,

Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70


Introduction

Univariate → Multivariate : changes

While dealing with a single variable X we denote by xi the i th


observation of the variable and have it’s mean µ & variance
σ 2 , but as soon as we have a second variable X2 , we have i th
observation xi = (xi1 , xi2 ) ∈ R2

Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70


Introduction

Univariate → Multivariate : changes

While dealing with a single variable X we denote by xi the i th


observation of the variable and have it’s mean µ & variance
σ 2 , but as soon as we have a second variable X2 , we have i th
observation xi = (xi1 , xi2 ) ∈ R2
 
µ1
mean µ =
µ2

σ12
 
variance Σ =
σ22
Of course the covariance will also claim it’s rightful place.
And we will rather talk about Variance-Covariance rather
than just Variance.

Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70


Introduction

Univariate → Multivariate : changes


As we travel to the p− variate space, we will deal with
 2 
σ1 · · · σ1p
 .. .. 
 . . 
Σ(p×p) =  .
 
.. 
 . . . 
σp1 · · · σp2

Amit (ISI, Chennai) Multivariate January 20, 2024 4 / 70


Introduction

Univariate → Multivariate : changes


As we travel to the p− variate space, we will deal with
 2 
σ1 · · · σ1p
 .. .. 
 . . 
Σ(p×p) =  .
 
.. 
 . . . 
σp1 · · · σp2

The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a


p-dimensional random variable X ∈ Rp and the population
mean is denoted by
 
µ1
µ =  ... 
 

µp

Amit (ISI, Chennai) Multivariate January 20, 2024 4 / 70


Introduction

Univariate → Multivariate : changes

With more than two variables, we can not plot anymore on


our two dimensional devices, with more than three variables
visualisation becomes a challenge.

Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70


Introduction

Univariate → Multivariate : changes

With more than two variables, we can not plot anymore on


our two dimensional devices, with more than three variables
visualisation becomes a challenge.

As dimension increases, the focus of analysis shifts to


reduction of dimensions.

Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70


Introduction

Univariate → Multivariate : changes

With more than two variables, we can not plot anymore on


our two dimensional devices, with more than three variables
visualisation becomes a challenge.

As dimension increases, the focus of analysis shifts to


reduction of dimensions.

Though this can be done in many manners, choosing the


appropriate reduction method is of paramount importance.

Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70


Introduction

Univariate → Multivariate : changes

With more than two variables, we can not plot anymore on


our two dimensional devices, with more than three variables
visualisation becomes a challenge.

As dimension increases, the focus of analysis shifts to


reduction of dimensions.

Though this can be done in many manners, choosing the


appropriate reduction method is of paramount importance.

Chosen method of dimension reduction should be able to


achieve the goal with least loss of information.

Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70


Introduction

Bivariate Normal Density

Amit (ISI, Chennai) Multivariate January 20, 2024 6 / 70


Introduction

Three variable data plot

Amit (ISI, Chennai) Multivariate January 20, 2024 7 / 70


Introduction

Multivariate Analysis

Statistical Methods is a journey from Data to Information.

Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70


Introduction

Multivariate Analysis

Statistical Methods is a journey from Data to Information.

Though it is easy to collect data, it’s far harder to collect


information.

Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70


Introduction

Multivariate Analysis

Statistical Methods is a journey from Data to Information.

Though it is easy to collect data, it’s far harder to collect


information.

Today’s technology make it very easy to collect large amounts


of data; multivariate methods are needed to determine
whether such massive amount of data actually contain
information.

Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70


Introduction

Introduction

Multivariate data occur in all branches of science. Almost all


data collected by today’s researchers can be classified as
multivariate data. For example a doctor trying to identify
characteristics of patients for whom a particular drug is more
effective than in others.

Amit (ISI, Chennai) Multivariate January 20, 2024 9 / 70


Introduction

Introduction

Multivariate data occur in all branches of science. Almost all


data collected by today’s researchers can be classified as
multivariate data. For example a doctor trying to identify
characteristics of patients for whom a particular drug is more
effective than in others.

A manager may like to identify areas in which workers need a


retraining among a large group of workers based on several
performance characteristics over the recent past.

Amit (ISI, Chennai) Multivariate January 20, 2024 9 / 70


Introduction

MVA can...

Multivariate methods can help determine whether there is


information in data, and they can also help to summarize
that information when it exists.

Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70


Introduction

MVA can...

Multivariate methods can help determine whether there is


information in data, and they can also help to summarize
that information when it exists.

Multivariate analysis is the best way to summarize a data


tables with many variables by creating a few new variables
containing most of the information.

Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70


Introduction

MVA can...

Multivariate methods can help determine whether there is


information in data, and they can also help to summarize
that information when it exists.

Multivariate analysis is the best way to summarize a data


tables with many variables by creating a few new variables
containing most of the information.

These new variables are then used for problem solving and
display, i.e., classification, relationships, control charts, and
more.

Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70


Introduction

MVA can...

Multivariate methods can do the following:

1 Data reduction or structural simplification.


2 Sorting and grouping.
3 Investigation of the dependence among variables.
4 Prediction.
5 Hypothesis construction and testing.

Amit (ISI, Chennai) Multivariate January 20, 2024 11 / 70


Introduction

Data reduction or simplification


• Using data on several variables related to cancer patient responses to
radio- therapy, a simple measure of patient response to radiotherapy
was constructed. (See Exercise 1.15.)
• nack records from many nations were used to develop an index of
performance for both male and female athletes. (See [8] and [22].)
• Multispectral image data collected by a high-altitude scanner were
reduced to a form that could be viewed as images (pictures) of a
shoreline in two dimensions. (See [23].)
• Data on several variables relating to yield and protein content were
used to create an index to select parents of subsequent generations of
improved bean plants. (See [13].)
• A matrix of tactic similarities was developed from aggregate data
derived from professional mediators. From this matrix the number of
dimensions by which professional mediators judge the tactics they use
in resolving disputes was determined. (See [21].)
Amit (ISI, Chennai) Multivariate January 20, 2024 12 / 70
Introduction

Multivariate thinking

Body of thought processes that illuminate the


interrelatedness between and within sets of variables.

Amit (ISI, Chennai) Multivariate January 20, 2024 13 / 70


Introduction

Multivariate thinking

Body of thought processes that illuminate the


interrelatedness between and within sets of variables.

The essence of multivariate thinking is to expose the inherent


structure and meaning revealed within these sets of variables
through application and interpretation of various statistical
methods

Amit (ISI, Chennai) Multivariate January 20, 2024 13 / 70


Introduction

Why the multivariate approach?

Big idea — multiple response outcomes

Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70


Introduction

Why the multivariate approach?

Big idea — multiple response outcomes

With univariate analyses we have just one dependent


variable of interest

Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70


Introduction

Why the multivariate approach?

Big idea — multiple response outcomes

With univariate analyses we have just one dependent


variable of interest

Although any analysis of data involving more than one


variable could be seen as multivariate, we typically
reserve the term for multiple dependent variables

Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70


Introduction

Why the multivariate approach?

Big idea — multiple response outcomes

With univariate analyses we have just one dependent


variable of interest

Although any analysis of data involving more than one


variable could be seen as multivariate, we typically
reserve the term for multiple dependent variables

So MV analysis is an extension of UV ones, or conversely,


many of the UV analyses are special cases of MV ones

Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70


Introduction

Projection of MV data

Projection of data is reduction of dimensionality, model in


latent variables

Amit (ISI, Chennai) Multivariate January 20, 2024 15 / 70


Introduction

Projection of MV data

Projection of data is reduction of dimensionality, model in


latent variables

Often the primary objective of Multivariate Analysis is to


summarize large amount of data by means of relatively few
parameters.
Amit (ISI, Chennai) Multivariate January 20, 2024 15 / 70
Introduction

Theme of MV analysis

The underlying theme behind many multivariate techniques


is simplification.

Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70


Introduction

Theme of MV analysis

The underlying theme behind many multivariate techniques


is simplification.

Multivariate techniques are often concerned with finding


relationships among

Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70


Introduction

Theme of MV analysis

The underlying theme behind many multivariate techniques


is simplification.

Multivariate techniques are often concerned with finding


relationships among

The response variables

Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70


Introduction

Theme of MV analysis

The underlying theme behind many multivariate techniques


is simplification.

Multivariate techniques are often concerned with finding


relationships among

The response variables


The experimental units

Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70


Introduction

Theme of MV analysis

The underlying theme behind many multivariate techniques


is simplification.

Multivariate techniques are often concerned with finding


relationships among

The response variables


The experimental units
Both response variables and experimental units

Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70


Introduction Examples

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 17 / 70


Introduction Examples

Textile Example
Example
A textile shop manager is studying the sales of ‘classic blue’ pullovers
over 10 different periods. He observes the number of pullovers sold
(X1 ), variation in price (X2 , in EUR), the advertisement costs in local
newspapers (X3 , in EUR) and the presence of a sales assistant (X4 , in
hours per period). Over the periods, he observes the following data
matrix:

Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70


Introduction Examples

Textile Example
Example
A textile shop manager is studying the sales of ‘classic blue’ pullovers
over 10 different periods. He observes the number of pullovers sold
(X1 ), variation in price (X2 , in EUR), the advertisement costs in local
newspapers (X3 , in EUR) and the presence of a sales assistant (X4 , in
hours per period). Over the periods, he observes the following data
matrix:
X1 X2 X3 X4
 
 
230 125 200 109

 


 

181 99 55 107

 


 

165 97 105 98

 


 

150 115 85 71

 

 
X(10×4) = 97 120 0 82
192 100 150 103

 


 

181 80 85 111

 


 

189 90 120 93

 


 

172 95 110 86

 


 

170 125 130 78
Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70
Introduction Examples

Textile Example

X1 X2 X3 X4
 
 
230 125 200 109

 


 

181 99 55 107

 


 

165 97 105 98

 


 

150 115 85 71

 

 
X(10×4) = 97 120 0 82
192 100 150 103

 


 

181 80 85 111

 


 

189 90 120 93

 


 

172 95 110 86

 


 

170 125 130 78
Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70
Covariance & Correlation

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 19 / 70


Covariance & Correlation

Covariance & Correlation

Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70


Covariance & Correlation

Covariance & Correlation

This convinces that the price must have a large influence on


the number of pullovers sold. So the following scatter-plot of
X2 vs. X1 . A rough impression is that the cloud is somewhat
downward-sloping. A computation of the empirical covariance
yields
10
1X
sX1 X2 = (X1i − X¯1 )(X2i − X¯2 ) = −80.02
9
i=1

a negative value as expected.

Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70


Covariance & Correlation

Covariance & Correlation

This convinces that the price must have a large influence on


the number of pullovers sold. So the following scatter-plot of
X2 vs. X1 . A rough impression is that the cloud is somewhat
downward-sloping. A computation of the empirical covariance
yields
10
1X
sX1 X2 = (X1i − X¯1 )(X2i − X¯2 ) = −80.02
9
i=1

a negative value as expected.

Note: The covariance function is scale dependent. Thus, if the prices


in this example were in Japanese Yen (JPY), we would obtain a
different answer. A measure of (linear) dependence independent of the
scale is the correlation.

Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70


Covariance & Correlation

Covariance & Correlation

Amit (ISI, Chennai) Multivariate January 20, 2024 21 / 70


Covariance & Correlation

Covariance & Correlation


Covariance is a measure of dependency between random
variables. Given two (random) variables X and Y the
(theoretical) covariance is defined by:
σXY = Cov (X , Y ) = E (XY ) − E (X ).E (Y )
If X and Y are independent of each other, the covariance
Cov (X , Y ) is necessarily equal to zero. The converse is not
true. The covariance of X with itself is the variance:

σXX = Var (X ) = Cov (X , X )


If the variable X is p-dimensional multivariate, e.g.,
 
X1
X =  ... 
 

Xp
Amit (ISI, Chennai) Multivariate January 20, 2024 21 / 70
Covariance & Correlation

Covariance & Correlation


The theoretical covariances among all the elements are put
into matrix form, i.e., the covariance matrix:
 
σX1 X1 · · · σX1 Xp
Σ =  ... .. ..
 
. . 
σXp X1 · · · σXp Xp
Sample versions of these quantities are of course
n
1X
sXY = (x1 − x̄)(y1 − ȳ )
n
i=1
1 1
For small n ≤ 30 factor n need be replaced by n−1 to correct
against bias.

Amit (ISI, Chennai) Multivariate January 20, 2024 22 / 70


Covariance & Correlation

Covariance & Correlation


The theoretical covariances among all the elements are put
into matrix form, i.e., the covariance matrix:
 
σX1 X1 · · · σX1 Xp
Σ =  ... .. ..
 
. . 
σXp X1 · · · σXp Xp
Sample versions of these quantities are of course
n
1X
sXY = (x1 − x̄)(y1 − ȳ )
n
i=1
1 1
For small n ≤ 30 factor need be replaced by n−1
n to correct
against [Link] a scatter-plot of two variables the covariances
measure how close the scatter is to a line, it should already
be understood here that in this sense covariance measures
only linear dependence.
Amit (ISI, Chennai) Multivariate January 20, 2024 22 / 70
Covariance & Correlation

Covariance & Correlation


The correlation between two variables X and Y is defined
from the covariance as the following

Cov (X , Y )
ρXY = p
Var (X )Var (Y )
The advantage of the correlation is that, it is independent of
the scale, i.e., changing the variables’ scale of measurement
does not change the value of the correlation. Therefore, the
correlation is more useful as a measure of association between
two random variables than the covariance.

The empirical version is

s(X , Y )
rXY = √
sXX sYY

Amit (ISI, Chennai) Multivariate January 20, 2024 23 / 70


Covariance & Correlation

Covariance & Correlation

The correlation is in absolute value always less than 1. It is


zero if the covariance is zero and vice-versa. For
p-dimensional vectors (X1 , . . . , Xp )T we have the theoretical
correlation matrix
 
ρX1 X1 · · · ρX1 Xp
ρ =  ... .. .. 

. . 
ρXp X1 · · · ρXp Xp

Amit (ISI, Chennai) Multivariate January 20, 2024 24 / 70


Covariance & Correlation

Covariance & Correlation


Theorem
If X and Y are independent, then ρ(X , Y ) = Cov (X , Y ) = 0

A In general, the converse is not true, as the following


example shows.
Example
Consider a standard normally-distributed random variable X and a
random variable Y = X 2 , which is surely not independent of X . Here
we have

Cov (X , Y ) = E (XY ) − E (X ).E (Y ) = E (X 3 ) = 0


(because E (X ) = 0 and E (X 2 ) = 1). Therefore ρ(X , Y ) = 0, as well.
This example also shows that correlations and covariances measure
only linear dependence. The quadratic dependence of Y = X 2 on X is
not reflected by these measures of dependence.
Amit (ISI, Chennai) Multivariate January 20, 2024 25 / 70
Covariance & Correlation

Covariance & Correlation

Remark
For two normal random variables, the converse of Theorem is true:
zero covariance for two normally-distributed random variables implies
independence.

Theorem enables us to check for independence between the


components of a bivariate normal random variable. That is,
we can use the correlation and test whether it is zero. The
distribution of rXY for an arbitrary (X , Y ) is unfortunately
complicated. The distribution of rXY will be more accessible if
(X , Y ) are jointly normal. If we transform the correlation by
Fisher’s Z -transformation, Skip TOH
 
1 1 + rXY
W = log
2 1 − rXY
Amit (ISI, Chennai) Multivariate January 20, 2024 26 / 70
Covariance & Correlation

Covariance & Correlation

We obtain a variable that has a more accessible distribution.


Under the hypothesis that ρ = 0, W has an asymptotic normal
distribution. Approximations of the expectation and variance
of W are given by the following:
 
1 1 + ρXY
E (W ) ≈ log
2 1 − ρXY
1
Var (W ) ≈
(n − 3)
The distribution is given in Theorem 2.2

Amit (ISI, Chennai) Multivariate January 20, 2024 27 / 70


Covariance & Correlation

Covariance & Correlation

Theorem
2.2

W − E (W ) D
Z= p −→ N(0, 1)
Var (W )
D
The symbol ”−→” denotes convergence in distribution.

Amit (ISI, Chennai) Multivariate January 20, 2024 28 / 70


Covariance & Correlation

Covariance & Correlation

Theorem
2.2

W − E (W ) D
Z= p −→ N(0, 1)
Var (W )
D
The symbol ”−→” denotes convergence in distribution.

Theorem allows us to test different hypotheses on correlation.


We can fix the level of significance α (the probability of
rejecting a true hypothesis) and reject the hypothesis if the
difference between the hypothetical value and the calculated
value of Z is greater than the corresponding critical value of
the normal distribution. The following example illustrates the
procedure.
Amit (ISI, Chennai) Multivariate January 20, 2024 28 / 70
Covariance & Correlation Test of correlation

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 29 / 70


Covariance & Correlation Test of correlation

Covariance & Correlation


Example
Let us consider again the pullovers data set from example. Consider
the correlation between the presence of the sales assistants (X4) vs. the
number of sold pullovers (X1 ) see Figure. Here we compute the
correlation as rX1 X4 = 0.633

The Z -transform of this value is


 
1 1 + rX1 X4
w = log = 0.746
2 1 − rX1 X4
The sample size is n√= 10, so for the hypothesis ρX1 X4 = 0, the statistic
to consider is: z = 7(0.746 − 0)
which is just statistically significant at the 5% level (i.e., 1.974 is just a
little larger than 1.96).

Amit (ISI, Chennai) Multivariate January 20, 2024 30 / 70


Covariance & Correlation Test of correlation

Covariance & Correlation


Example
Let us consider again the pullovers data set from example. Consider
the correlation between the presence of the sales assistants (X4) vs. the
number of sold pullovers (X1 ) see Figure. Here we compute the
correlation as rX1 X4 = 0.633

The Z -transform of this value is


 
1 1 + rX1 X4
w = log = 0.746
2 1 − rX1 X4
The sample size is n√= 10, so for the hypothesis ρX1 X4 = 0, the statistic
to consider is: z = 7(0.746 − 0)
which is just statistically significant at the 5% level (i.e., 1.974 is just a
little larger than 1.96).

Amit (ISI, Chennai) Multivariate January 20, 2024 30 / 70


Covariance & Correlation Test of correlation

Covariance & Correlation

Remark
The normalizing and variance stabilizing properties of W are
asymptotic. In addition the use of W in small samples (for n ≤ 25) is
improved by Hotelling’s transform

3W + tanh(W ) 1
W∗ = W − with Var (W ∗ ) =
4(n − 1) n−1

The transformed variable W ∗ is asymptotically distributed as a normal


distribution.

Amit (ISI, Chennai) Multivariate January 20, 2024 31 / 70


Covariance & Correlation Test of correlation

Covariance & Correlation

Remark
Note that the Fisher’s Z -transform is the inverse of the hyperbolic
tangent function: W = tanh−1 (rXY ); equivalently
2W
rXY = tanh(W ) = ee 2W −1
+1
.

Remark
Under the assumptions of normality of X and Y , we may test their
independence (ρXY = 0) using the exact t-distribution of the statistic
s
n − 2 ρXY =0
T = rXY 2
∼ tn−2
1 − rXY

Setting the probability of the first error type to α, we reject the null
hypothesis ρXY = 0 if |T | ≥ t1−α/2;n−2

Amit (ISI, Chennai) Multivariate January 20, 2024 32 / 70


Covariance & Correlation Need for plotting

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 33 / 70


Covariance & Correlation Need for plotting

Why plot?

Plots are important

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?

Plots are important but frequently neglected.

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?

Plots are important but frequently neglected. Lets look at


the following example of a small data set.

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?

Plots are important but frequently neglected. Lets look at


the following example of a small data set.
x1 3 4 2 6 8 2 5
x2 5 5.5 4 7 10 5 7.5

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?

Plots are important but frequently neglected. Lets look at


the following example of a small data set.

x1 5 4 6 2 2 8 3
x2 5 5.5 4 7 10 5 7.5

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?

So the difference is only on the bivariate plot, NOT on the


marginal plots.

Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70


Covariance & Correlation Need for plotting

Why plot?
Pullovers Data subjected to MiniTab

Amit (ISI, Chennai) Multivariate January 20, 2024 35 / 70


Canonical Correlation

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 36 / 70


Canonical Correlation From Multiple Regression

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 37 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

Canonical Correlation Analysis is a generalisation of the


Multiple Regression problem.

Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

Canonical Correlation Analysis is a generalisation of the


Multiple Regression problem.

The Multiple Correlation Coefficient can also be


interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.

Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

Canonical Correlation Analysis is a generalisation of the


Multiple Regression problem.

The Multiple Correlation Coefficient can also be


interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.

The response vector x is partitioned into two parts so


that:

Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

Canonical Correlation Analysis is a generalisation of the


Multiple Regression problem.

The Multiple Correlation Coefficient can also be


interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.

The response vector x is partitioned into two parts so


that:
     
x1 µ1 Σ11 Σ12
∼N ,
x2 µ2 Σ21 Σ12

Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70


Canonical Correlation From Multiple Regression

Introduction

If the q variables in x1 are measurements that are hard to


obtain or expensive to measure or not measurable but the
(p − q) variables are easy to measure, then we may attempt to
predict the difficult to measure x1 variables through the easy
to measure x2 variables.

Amit (ISI, Chennai) Multivariate January 20, 2024 39 / 70


Canonical Correlation From Multiple Regression

Introduction

If the q variables in x1 are measurements that are hard to


obtain or expensive to measure or not measurable but the
(p − q) variables are easy to measure, then we may attempt to
predict the difficult to measure x1 variables through the easy
to measure x2 variables.

The interrelationship that exist between the two sets, may


almost be described by the correlation between a few linear
combinations of the responses within each set of variables.

Amit (ISI, Chennai) Multivariate January 20, 2024 39 / 70


Canonical Correlation From Multiple Regression

Introduction

We’ve dealt with multiple regression, a case where we


used multiple independent variables to predict a single
dependent variable.

Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70


Canonical Correlation From Multiple Regression

Introduction

We’ve dealt with multiple regression, a case where we


used multiple independent variables to predict a single
dependent variable.

We came up with a linear combination of the predictors


that would result in the most variance accounted for in
the dependent variable (maximize R 2 )

Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70


Canonical Correlation From Multiple Regression

Introduction

We’ve dealt with multiple regression, a case where we


used multiple independent variables to predict a single
dependent variable.

We came up with a linear combination of the predictors


that would result in the most variance accounted for in
the dependent variable (maximize R 2 )

We will also see PC regression, in which linear


combinations of the predictors which account for all the
variance in the predictors can be used in lieu of the
original variables

Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70


Canonical Correlation From Multiple Regression

Introduction

Reason?

Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70


Canonical Correlation From Multiple Regression

Introduction

Reason?

Dimension reduction: select only the best components

Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70


Canonical Correlation From Multiple Regression

Introduction

Reason?

Dimension reduction: select only the best components

Independence of predictors: solves collinearity issue

Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

The problem posed now is what to do if we had multiple


dependent variables

Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

The problem posed now is what to do if we had multiple


dependent variables

How could we come up with a way to understand the


relationship between a set of predictors and DVs?

Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

The problem posed now is what to do if we had multiple


dependent variables

How could we come up with a way to understand the


relationship between a set of predictors and DVs?

Canonical Correlation measures the relationship between


two sets of variables

Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70


Canonical Correlation From Multiple Regression

Multiple Dependent Variable

The problem posed now is what to do if we had multiple


dependent variables

How could we come up with a way to understand the


relationship between a set of predictors and DVs?

Canonical Correlation measures the relationship between


two sets of variables

Example: looking to see if different personality traits


correlate with performance scores for a variety of tasks

Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70


Canonical Correlation From Multiple Regression

Canonical Correlation

CC extends bivariate correlation, allowing 2+ continuous IVs (on


left) with 2+ continuous DVs (variables on right).
Focus is on correlations & weights
Main question: How are the best linear combinations of predictors
related to the best linear combinations of the DVs?

Amit (ISI, Chennai) Multivariate January 20, 2024 43 / 70


Canonical Correlation From Multiple Regression

The Analysis

In CC, there are several layers of analysis:


Correlation between pairs of canonical variates
Loadings between IVs & their canonical variates (on left)
Loadings between DVs & their canonical variates (on right)
Adequacy
Communalities
Other
Redundancy between IVs & can. variates on other (right) side
Redundancy between DVs & can. variates on other (left) side
As we will discuss later, these are problematic

Amit (ISI, Chennai) Multivariate January 20, 2024 44 / 70


Canonical Correlation From Multiple Regression

When to use

As exploratory tool to see if two sets of variables are related


If significant overall shared variance, several layers of analysis can
explore variables & variates involved
In a modelling type approach where you have a theoretical reason
for considering the variables as sets, and that one predicts another
See if one set of two or more variables relate longitudinally across
two time points
Variables at t1 are predictors; same variables at t2 are DVs
See layers of analysis for tentative causal evidence

Amit (ISI, Chennai) Multivariate January 20, 2024 45 / 70


Canonical Correlation From Multiple Regression

Comparison with Multiple Regression

With canonical correlation we will again (like in MR and PCA) be


creating linear combinations for our sets of variables
In MR, the multiple correlation coefficient was between the
predicted values created by the linear combination of predictors
(i.e. the composite) and the dependent variable
R 2 = square of multiple correlation

Amit (ISI, Chennai) Multivariate January 20, 2024 46 / 70


Canonical Correlation From Multiple Regression

Comparison with Multiple Regression

Creating linear composites of our respective variable sets (X , Y ):


Creating some single variable that represents the X s and another
single variable that represents the Y s.
Given a linear combination of X variables:
F = f1 X1 + f2 X2 + · · · + fp Xp
and a linear combination of Y variables:
G = g1 Y1 + g2 Y2 + · · · + gq Yq
The first canonical correlation is:
The maximum correlation coefficient between F and G , for all F
and G
The correlation we are interested in here will be between the linear
combinations (variates) created for both sets of variables
However now there will be the possibility for several ways in which
to combine the variables that may provide a useful interpretation
of the relationship

Amit (ISI, Chennai) Multivariate January 20, 2024 47 / 70


Canonical Correlation From Multiple Regression

Canonical Correlation : Pictorially

Amit (ISI, Chennai) Multivariate January 20, 2024 48 / 70


Canonical Correlation From Multiple Regression

Terminologies

Variables
Those measurements recorded in the dataset
Canonical Variates
The linear combination of the sets of variables (predictor and DV)
One variate for each set of variables
Canonical correlation
The correlation between variates
Pairs of Canonical Variates
As mentioned, we may have more than one pair of linear
combinations that provide an interesting measure of the relationship

Amit (ISI, Chennai) Multivariate January 20, 2024 49 / 70


Canonical Correlation From Multiple Regression

More about Can Cor

Canonical Correlation is one of the most general multivariate


forms – multiple regression, discriminate function analysis and
MANOVA are all special cases of it
As the basic output of regression and ANOVA are equivalent, so too
is the case here between canonical correlation and regression
Example: CanCorr squared between predictor set and a lone DV =
R2 from multiple regression
Only one solution
The number of canonical variate pairs you can have is equal to the
number of variables in the smaller set
When you have many variables on both sides of the equation you
end up with many canonical correlations.
Arranged in descending order, in most cases the first one or two
will be the ones of interest.

Amit (ISI, Chennai) Multivariate January 20, 2024 50 / 70


Canonical Correlation From Multiple Regression

Care need be taken

Number of canonical variate pairs


How many prove to be statistically/practically significant?
Interpreting the canonical variates
Where is the meaning in the combinations of variables created?
Importance of canonical variates
How strong is the correlation between the canonical variates?
What is the nature of the variate’s relation to the individual
variables in its own set? The other set?
Canonical variate scores
If one did directly measure the variate, what would the subjects’
scores be?

Amit (ISI, Chennai) Multivariate January 20, 2024 51 / 70


Canonical Correlation From Multiple Regression

Limitations

The procedure maximizes the correlation between the linear


combination of variables, however the combination may not make
much sense theoretically
Nonlinearity poses the same problem as it does in simple
correlation i.e. if there is a nonlinear relationship between the sets
of variables the technique is not designed to pick up on that
Cancorr is very sensitive to data involved i.e. influential cases and
which variables are chosen or left out of analysis can have
dramatic effects on the results just like they do elsewhere
Correlation, as we know, does not automatically imply causality
It is a necessary but not sufficient condition for causality
Being a simple correlation in the end, cancorr is often thought of
as a descriptive procedure

Amit (ISI, Chennai) Multivariate January 20, 2024 52 / 70


Canonical Correlation From Multiple Regression

Conditions & Assumptions

Sample size
Normality
Linearity
Homoscedasticity
Multicollinearity/Singularity
Outliers
Assumptions may better be checked before embarking on a big leap.

Amit (ISI, Chennai) Multivariate January 20, 2024 53 / 70


Canonical Correlation From Multiple Regression

Data and Computations


Canonical Correlation uses the correlations from the raw data and uses
this as input data
 
R11 R12
R21 R12
The Canonical Correlation Matrix R is obtained by
−1 −1
R = RYY RYX RXX RXY

To get the canonical correlations, you get the eigenvalues of R and take
the square root
p
rci2 = λi ⇒ rci = λi
Consequently the eigenvector corresponding to each eigenvalue is
transformed into the coefficients that specify the linear combination
that will make up a canonical variate
Amit (ISI, Chennai) Multivariate January 20, 2024 54 / 70
Canonical Correlation From Multiple Regression

Data and Computations

Amit (ISI, Chennai) Multivariate January 20, 2024 55 / 70


Marginal & Conditional density functions

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 56 / 70


Marginal & Conditional density functions

Distribution & Density Functions

Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70


Marginal & Conditional density functions

Distribution & Density Functions


Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by

F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )

Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70


Marginal & Conditional density functions

Distribution & Density Functions


Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by

F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )

For continuous X , there exists a nonnegative probability


density function (pdf ) f , such that
Rx
F (x) = −∞ f (u)du

Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70


Marginal & Conditional density functions

Distribution & Density Functions


Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by

F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )

For continuous X , there exists a nonnegative probability


density function (pdf ) f , such that
Rx
F (x) = −∞ f (u)du
R∞
Note that −∞ f (u)du = 1
Most of the
R x integrals appearing below are multidimensional. For
instance, −∞ f (u)du means
Z xp Z x1
··· f (u1 , . . . , up )du1 · · · dup
−∞ −∞

Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70


Marginal & Conditional density functions

Marginal & Conditional pdf

Note also that the cdf F is differentiable with

δ p F (x)
f (x) =
δx1 · · · δxp

Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70


Marginal & Conditional density functions

Marginal & Conditional pdf

Note also that the cdf F is differentiable with

δ p F (x)
f (x) =
δx1 · · · δxp

For discrete X , the values of this random variable are


concentrated on a countable or finite set of points {cj }j∈J

Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70


Marginal & Conditional density functions

Marginal & Conditional pdf

Note also that the cdf F is differentiable with

δ p F (x)
f (x) =
δx1 · · · δxp

For discrete X , the values of this random variable are


concentrated on a countable or finite set of points {cj }j∈J

The probability of eventsPof the form {X ∈ D} can then be


computed as P(X ∈ D) = {j:cj ∈D} P(X = cj )

If we partition X as X T = (X1T , X2T ) with X1 ∈ R k and X2 ∈ R p−k ,


then the function FX1 (x1 ) = P(X1 ≤ x1 ) = F (x11 . . . x1k , ∞, . . . , ∞) is
called the marginal cdf.

F = F (x) is called the joint cdf.


Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70
Marginal & Conditional density functions

Marginal & Conditional pdf

For continuous X the marginal pdf can be computed from the


joint density by integrating out the variable not of interest.
Z ∞
fX1 (x1 ) = f (x1 , x2 )dx2
−∞

The conditional pdf of X2 given X1 = x1 is given as


f (x2 |x1 ) = ffX(x1(x,x12))
1

Amit (ISI, Chennai) Multivariate January 20, 2024 59 / 70


Marginal & Conditional density functions

Marginal & Conditional pdf


Consider the pdf
1
+ 32 x2
 
2 x1 0 ≤ x1 , x2 ≤ 1,
f (x1 , x2 ) =
0 otherwise

f (x1 , x2 ) is a density since


1 1
x12 x22
Z  
1 3 1 3
f (x1 , x2 )dx1 x2 = + = + =1
2 2 0 2 2 0 4 4
The marginal distributions are
Z Z 1
1 3 1 3
fX1 (x1 ) = f (x1 , x2 )dx2 = x1 + x2 dx2 = x1 +
0 2 2 2 4
Z Z 1  
1 3 3 1
fX2 (x2 ) = f (x1 , x2 )dx1 = x1 + x2 dx1 = x2 +
0 2 2 2 4
Amit (ISI, Chennai) Multivariate January 20, 2024 60 / 70
Marginal & Conditional density functions

Marginal & Conditional pdf

The conditional densities are therefore


1 3 1 3
2 x1 + 2 x2 2 x1 + 2 x2
f (x2 |x1 ) = 1 3
and f (x2 |x1 ) = 3 1
2 x1 + 4 2 x2 + 4
Note that these conditional pdf ’s are nonlinear in x1 and x2
although the joint pdf has a simple (linear) structure.

Amit (ISI, Chennai) Multivariate January 20, 2024 61 / 70


Matrix Notations

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 62 / 70


Matrix Notations

Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
 matrix. 
x11 · · · x1p
 .. .. 
 . . 
X(n×p) =  .

.. 

 . . . 
x1n · · · xnp

Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70


Matrix Notations

Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
 matrix. 
x11 · · · x1p
 .. .. 
 . . 
X(n×p) =  .

.. 

 .. . 
x1n · · · xnp
The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a
p-dimensional
  random variable X ∈ Rp .
x¯1
 .. 
X̄ =  .  = n−1 X T 1n .
x¯p

Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70


Matrix Notations

Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
 matrix. 
x11 · · · x1p
 .. .. 
 . . 
X(n×p) =  .

.. 

 .. . 
x1n · · · xnp
The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a
p-dimensional
  random variable X ∈ Rp .
x¯1
 .. 
X̄ =  .  = n−1 X T 1n .
x¯p
The sample varinace-covariance matrix is given by:

S = n−1 X T X − x̄ x̄ T = n−1 (X T X − n−1 X T 1n 1n X )


Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70
MultiNormal Distribution

Contents

1 Introduction
Examples

2 Covariance & Correlation


Test of correlation
Need for plotting

3 Canonical Correlation
From Multiple Regression

4 Marginal & Conditional density functions

5 Matrix Notations

6 MultiNormal Distribution

Amit (ISI, Chennai) Multivariate January 20, 2024 64 / 70


MultiNormal Distribution

Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .

Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70


MultiNormal Distribution

Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .

Definition
Multivariate normal :

x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn

Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70


MultiNormal Distribution

Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .

Definition
Multivariate normal :

x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn

However the Multivariate Normal distribution could also be


looked upon as a generalization of univariate normal
distribution to p ≥ 2 dimensions.

Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70


MultiNormal Distribution

Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .

Definition
Multivariate normal :

x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn

However the Multivariate Normal distribution could also be


looked upon as a generalization of univariate normal
distribution to p ≥ 2 dimensions.

1 x−µ 2
f (x) = √ e( σ ) −∞<x <∞
2πσ 2
Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70
MultiNormal Distribution

Definition, properties
The term
 2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ

Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70


MultiNormal Distribution

Definition, properties
The term
 2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ

in the exponent of the univariate normal density function


measures the distance of x from µ in standard deviation units.

Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70


MultiNormal Distribution

Definition, properties
The term
 2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ

in the exponent of the univariate normal density function


measures the distance of x from µ in standard deviation units.

This could be generalized for a p × 1 vector x of observations


on several variables as

Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70


MultiNormal Distribution

Definition, properties
The term
 2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ

in the exponent of the univariate normal density function


measures the distance of x from µ in standard deviation units.

This could be generalized for a p × 1 vector x of observations


on several variables as

(x − µ)′ Σ−1 (x − µ)
The p × 1 vector µ represents the expected value of random
vector X, and the p × p matrix Σ is its variance-covariance
matrix, symmetric and positive definite, in order that the
above represents the generalized distance from x to µ
Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70
MultiNormal Distribution

Definition, properties

In the multivariate normal density we replace the univariate


distance by the generalized mutivariate distance.

Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70


MultiNormal Distribution

Definition, properties

In the multivariate normal density we replace the univariate


distance by the generalized mutivariate distance.

The constant also needs a replacement by a more general


form and this turns out to be (2π)−p/2 |Σ|1/2 .

Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70


MultiNormal Distribution

Definition, properties

In the multivariate normal density we replace the univariate


distance by the generalized mutivariate distance.

The constant also needs a replacement by a more general


form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :

Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70


MultiNormal Distribution

Definition, properties

In the multivariate normal density we replace the univariate


distance by the generalized mutivariate distance.

The constant also needs a replacement by a more general


form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :
1 1 ′ −1 (x−µ)
f (x) = e − 2 (x−µ) Σ
(2π)−p/2 |Σ|1/2
where −∞ < xi < ∞, i = 1, 2, . . . , p and this will be denoted by
Np (µ, Σ).

Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70


MultiNormal Distribution

Definition, properties

In the multivariate normal density we replace the univariate


distance by the generalized mutivariate distance.

The constant also needs a replacement by a more general


form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :
1 1 ′ −1 (x−µ)
f (x) = e − 2 (x−µ) Σ
(2π)−p/2 |Σ|1/2
where −∞ < xi < ∞, i = 1, 2, . . . , p and this will be denoted by
Np (µ, Σ).

Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix


 
σ11 σ12
Σ=
σ12 σ22

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix


 
σ11 σ12
Σ=
σ12 σ22

 
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix


 
σ11 σ12
Σ=
σ12 σ22

 
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11

Correlation coefficient ρ12 help us write


√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix


 
σ11 σ12
Σ=
σ12 σ22

 
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11

Correlation coefficient ρ12 help us write


√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12

So the squared distance (x − µ)′ Σ−1 (x − µ) simplifies as follows :

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .

The inverse of covariance matrix


 
σ11 σ12
Σ=
σ12 σ22

 
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11

Correlation coefficient ρ12 help us write


√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12

So the squared distance (x − µ)′ Σ−1 (x − µ) simplifies as follows :

Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70


MultiNormal Distribution

Example : Bivariate

 √ 
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2

Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70


MultiNormal Distribution

Example : Bivariate

 √ 
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2


σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )

Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70


MultiNormal Distribution

Example : Bivariate

 √ 
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2


σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )

" 2  2   #
1 x1 − µ1 x2 − µ2 x1 − µ1 x2 − µ2
= √ + √ − 2ρ12 √ √
1 − ρ212 σ11 σ22 σ11 σ22

Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70


MultiNormal Distribution

Example : Bivariate

 √ 
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2


σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )

" 2  2   #
1 x1 − µ1 x2 − µ2 x1 − µ1 x2 − µ2
= √ + √ − 2ρ12 √ √
1 − ρ212 σ11 σ22 σ11 σ22

Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70


MultiNormal Distribution

Bivariate Normal Density

Amit (ISI, Chennai) Multivariate January 20, 2024 70 / 70

You might also like