Multivariate Analyses
Amit K Biswas
Indian Statistical Institute
@ PGDSMA 2023 − 24
Stat Methods II
January 20, 2024
Amit (ISI, Chennai) Multivariate January 20, 2024 1 / 70
Introduction
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 2 / 70
Introduction
Univariate → Multivariate : changes
While dealing with a single variable X we denote by xi the i th
observation of the variable and have it’s mean µ & variance
σ2,
Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70
Introduction
Univariate → Multivariate : changes
While dealing with a single variable X we denote by xi the i th
observation of the variable and have it’s mean µ & variance
σ 2 , but as soon as we have a second variable X2 , we have i th
observation xi = (xi1 , xi2 ) ∈ R2
Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70
Introduction
Univariate → Multivariate : changes
While dealing with a single variable X we denote by xi the i th
observation of the variable and have it’s mean µ & variance
σ 2 , but as soon as we have a second variable X2 , we have i th
observation xi = (xi1 , xi2 ) ∈ R2
µ1
mean µ =
µ2
σ12
variance Σ =
σ22
Of course the covariance will also claim it’s rightful place.
And we will rather talk about Variance-Covariance rather
than just Variance.
Amit (ISI, Chennai) Multivariate January 20, 2024 3 / 70
Introduction
Univariate → Multivariate : changes
As we travel to the p− variate space, we will deal with
2
σ1 · · · σ1p
.. ..
. .
Σ(p×p) = .
..
. . .
σp1 · · · σp2
Amit (ISI, Chennai) Multivariate January 20, 2024 4 / 70
Introduction
Univariate → Multivariate : changes
As we travel to the p− variate space, we will deal with
2
σ1 · · · σ1p
.. ..
. .
Σ(p×p) = .
..
. . .
σp1 · · · σp2
The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a
p-dimensional random variable X ∈ Rp and the population
mean is denoted by
µ1
µ = ...
µp
Amit (ISI, Chennai) Multivariate January 20, 2024 4 / 70
Introduction
Univariate → Multivariate : changes
With more than two variables, we can not plot anymore on
our two dimensional devices, with more than three variables
visualisation becomes a challenge.
Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70
Introduction
Univariate → Multivariate : changes
With more than two variables, we can not plot anymore on
our two dimensional devices, with more than three variables
visualisation becomes a challenge.
As dimension increases, the focus of analysis shifts to
reduction of dimensions.
Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70
Introduction
Univariate → Multivariate : changes
With more than two variables, we can not plot anymore on
our two dimensional devices, with more than three variables
visualisation becomes a challenge.
As dimension increases, the focus of analysis shifts to
reduction of dimensions.
Though this can be done in many manners, choosing the
appropriate reduction method is of paramount importance.
Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70
Introduction
Univariate → Multivariate : changes
With more than two variables, we can not plot anymore on
our two dimensional devices, with more than three variables
visualisation becomes a challenge.
As dimension increases, the focus of analysis shifts to
reduction of dimensions.
Though this can be done in many manners, choosing the
appropriate reduction method is of paramount importance.
Chosen method of dimension reduction should be able to
achieve the goal with least loss of information.
Amit (ISI, Chennai) Multivariate January 20, 2024 5 / 70
Introduction
Bivariate Normal Density
Amit (ISI, Chennai) Multivariate January 20, 2024 6 / 70
Introduction
Three variable data plot
Amit (ISI, Chennai) Multivariate January 20, 2024 7 / 70
Introduction
Multivariate Analysis
Statistical Methods is a journey from Data to Information.
Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70
Introduction
Multivariate Analysis
Statistical Methods is a journey from Data to Information.
Though it is easy to collect data, it’s far harder to collect
information.
Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70
Introduction
Multivariate Analysis
Statistical Methods is a journey from Data to Information.
Though it is easy to collect data, it’s far harder to collect
information.
Today’s technology make it very easy to collect large amounts
of data; multivariate methods are needed to determine
whether such massive amount of data actually contain
information.
Amit (ISI, Chennai) Multivariate January 20, 2024 8 / 70
Introduction
Introduction
Multivariate data occur in all branches of science. Almost all
data collected by today’s researchers can be classified as
multivariate data. For example a doctor trying to identify
characteristics of patients for whom a particular drug is more
effective than in others.
Amit (ISI, Chennai) Multivariate January 20, 2024 9 / 70
Introduction
Introduction
Multivariate data occur in all branches of science. Almost all
data collected by today’s researchers can be classified as
multivariate data. For example a doctor trying to identify
characteristics of patients for whom a particular drug is more
effective than in others.
A manager may like to identify areas in which workers need a
retraining among a large group of workers based on several
performance characteristics over the recent past.
Amit (ISI, Chennai) Multivariate January 20, 2024 9 / 70
Introduction
MVA can...
Multivariate methods can help determine whether there is
information in data, and they can also help to summarize
that information when it exists.
Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70
Introduction
MVA can...
Multivariate methods can help determine whether there is
information in data, and they can also help to summarize
that information when it exists.
Multivariate analysis is the best way to summarize a data
tables with many variables by creating a few new variables
containing most of the information.
Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70
Introduction
MVA can...
Multivariate methods can help determine whether there is
information in data, and they can also help to summarize
that information when it exists.
Multivariate analysis is the best way to summarize a data
tables with many variables by creating a few new variables
containing most of the information.
These new variables are then used for problem solving and
display, i.e., classification, relationships, control charts, and
more.
Amit (ISI, Chennai) Multivariate January 20, 2024 10 / 70
Introduction
MVA can...
Multivariate methods can do the following:
1 Data reduction or structural simplification.
2 Sorting and grouping.
3 Investigation of the dependence among variables.
4 Prediction.
5 Hypothesis construction and testing.
Amit (ISI, Chennai) Multivariate January 20, 2024 11 / 70
Introduction
Data reduction or simplification
• Using data on several variables related to cancer patient responses to
radio- therapy, a simple measure of patient response to radiotherapy
was constructed. (See Exercise 1.15.)
• nack records from many nations were used to develop an index of
performance for both male and female athletes. (See [8] and [22].)
• Multispectral image data collected by a high-altitude scanner were
reduced to a form that could be viewed as images (pictures) of a
shoreline in two dimensions. (See [23].)
• Data on several variables relating to yield and protein content were
used to create an index to select parents of subsequent generations of
improved bean plants. (See [13].)
• A matrix of tactic similarities was developed from aggregate data
derived from professional mediators. From this matrix the number of
dimensions by which professional mediators judge the tactics they use
in resolving disputes was determined. (See [21].)
Amit (ISI, Chennai) Multivariate January 20, 2024 12 / 70
Introduction
Multivariate thinking
Body of thought processes that illuminate the
interrelatedness between and within sets of variables.
Amit (ISI, Chennai) Multivariate January 20, 2024 13 / 70
Introduction
Multivariate thinking
Body of thought processes that illuminate the
interrelatedness between and within sets of variables.
The essence of multivariate thinking is to expose the inherent
structure and meaning revealed within these sets of variables
through application and interpretation of various statistical
methods
Amit (ISI, Chennai) Multivariate January 20, 2024 13 / 70
Introduction
Why the multivariate approach?
Big idea — multiple response outcomes
Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70
Introduction
Why the multivariate approach?
Big idea — multiple response outcomes
With univariate analyses we have just one dependent
variable of interest
Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70
Introduction
Why the multivariate approach?
Big idea — multiple response outcomes
With univariate analyses we have just one dependent
variable of interest
Although any analysis of data involving more than one
variable could be seen as multivariate, we typically
reserve the term for multiple dependent variables
Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70
Introduction
Why the multivariate approach?
Big idea — multiple response outcomes
With univariate analyses we have just one dependent
variable of interest
Although any analysis of data involving more than one
variable could be seen as multivariate, we typically
reserve the term for multiple dependent variables
So MV analysis is an extension of UV ones, or conversely,
many of the UV analyses are special cases of MV ones
Amit (ISI, Chennai) Multivariate January 20, 2024 14 / 70
Introduction
Projection of MV data
Projection of data is reduction of dimensionality, model in
latent variables
Amit (ISI, Chennai) Multivariate January 20, 2024 15 / 70
Introduction
Projection of MV data
Projection of data is reduction of dimensionality, model in
latent variables
Often the primary objective of Multivariate Analysis is to
summarize large amount of data by means of relatively few
parameters.
Amit (ISI, Chennai) Multivariate January 20, 2024 15 / 70
Introduction
Theme of MV analysis
The underlying theme behind many multivariate techniques
is simplification.
Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70
Introduction
Theme of MV analysis
The underlying theme behind many multivariate techniques
is simplification.
Multivariate techniques are often concerned with finding
relationships among
Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70
Introduction
Theme of MV analysis
The underlying theme behind many multivariate techniques
is simplification.
Multivariate techniques are often concerned with finding
relationships among
The response variables
Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70
Introduction
Theme of MV analysis
The underlying theme behind many multivariate techniques
is simplification.
Multivariate techniques are often concerned with finding
relationships among
The response variables
The experimental units
Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70
Introduction
Theme of MV analysis
The underlying theme behind many multivariate techniques
is simplification.
Multivariate techniques are often concerned with finding
relationships among
The response variables
The experimental units
Both response variables and experimental units
Amit (ISI, Chennai) Multivariate January 20, 2024 16 / 70
Introduction Examples
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 17 / 70
Introduction Examples
Textile Example
Example
A textile shop manager is studying the sales of ‘classic blue’ pullovers
over 10 different periods. He observes the number of pullovers sold
(X1 ), variation in price (X2 , in EUR), the advertisement costs in local
newspapers (X3 , in EUR) and the presence of a sales assistant (X4 , in
hours per period). Over the periods, he observes the following data
matrix:
Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70
Introduction Examples
Textile Example
Example
A textile shop manager is studying the sales of ‘classic blue’ pullovers
over 10 different periods. He observes the number of pullovers sold
(X1 ), variation in price (X2 , in EUR), the advertisement costs in local
newspapers (X3 , in EUR) and the presence of a sales assistant (X4 , in
hours per period). Over the periods, he observes the following data
matrix:
X1 X2 X3 X4
230 125 200 109
181 99 55 107
165 97 105 98
150 115 85 71
X(10×4) = 97 120 0 82
192 100 150 103
181 80 85 111
189 90 120 93
172 95 110 86
170 125 130 78
Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70
Introduction Examples
Textile Example
X1 X2 X3 X4
230 125 200 109
181 99 55 107
165 97 105 98
150 115 85 71
X(10×4) = 97 120 0 82
192 100 150 103
181 80 85 111
189 90 120 93
172 95 110 86
170 125 130 78
Amit (ISI, Chennai) Multivariate January 20, 2024 18 / 70
Covariance & Correlation
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 19 / 70
Covariance & Correlation
Covariance & Correlation
Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70
Covariance & Correlation
Covariance & Correlation
This convinces that the price must have a large influence on
the number of pullovers sold. So the following scatter-plot of
X2 vs. X1 . A rough impression is that the cloud is somewhat
downward-sloping. A computation of the empirical covariance
yields
10
1X
sX1 X2 = (X1i − X¯1 )(X2i − X¯2 ) = −80.02
9
i=1
a negative value as expected.
Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70
Covariance & Correlation
Covariance & Correlation
This convinces that the price must have a large influence on
the number of pullovers sold. So the following scatter-plot of
X2 vs. X1 . A rough impression is that the cloud is somewhat
downward-sloping. A computation of the empirical covariance
yields
10
1X
sX1 X2 = (X1i − X¯1 )(X2i − X¯2 ) = −80.02
9
i=1
a negative value as expected.
Note: The covariance function is scale dependent. Thus, if the prices
in this example were in Japanese Yen (JPY), we would obtain a
different answer. A measure of (linear) dependence independent of the
scale is the correlation.
Amit (ISI, Chennai) Multivariate January 20, 2024 20 / 70
Covariance & Correlation
Covariance & Correlation
Amit (ISI, Chennai) Multivariate January 20, 2024 21 / 70
Covariance & Correlation
Covariance & Correlation
Covariance is a measure of dependency between random
variables. Given two (random) variables X and Y the
(theoretical) covariance is defined by:
σXY = Cov (X , Y ) = E (XY ) − E (X ).E (Y )
If X and Y are independent of each other, the covariance
Cov (X , Y ) is necessarily equal to zero. The converse is not
true. The covariance of X with itself is the variance:
σXX = Var (X ) = Cov (X , X )
If the variable X is p-dimensional multivariate, e.g.,
X1
X = ...
Xp
Amit (ISI, Chennai) Multivariate January 20, 2024 21 / 70
Covariance & Correlation
Covariance & Correlation
The theoretical covariances among all the elements are put
into matrix form, i.e., the covariance matrix:
σX1 X1 · · · σX1 Xp
Σ = ... .. ..
. .
σXp X1 · · · σXp Xp
Sample versions of these quantities are of course
n
1X
sXY = (x1 − x̄)(y1 − ȳ )
n
i=1
1 1
For small n ≤ 30 factor n need be replaced by n−1 to correct
against bias.
Amit (ISI, Chennai) Multivariate January 20, 2024 22 / 70
Covariance & Correlation
Covariance & Correlation
The theoretical covariances among all the elements are put
into matrix form, i.e., the covariance matrix:
σX1 X1 · · · σX1 Xp
Σ = ... .. ..
. .
σXp X1 · · · σXp Xp
Sample versions of these quantities are of course
n
1X
sXY = (x1 − x̄)(y1 − ȳ )
n
i=1
1 1
For small n ≤ 30 factor need be replaced by n−1
n to correct
against [Link] a scatter-plot of two variables the covariances
measure how close the scatter is to a line, it should already
be understood here that in this sense covariance measures
only linear dependence.
Amit (ISI, Chennai) Multivariate January 20, 2024 22 / 70
Covariance & Correlation
Covariance & Correlation
The correlation between two variables X and Y is defined
from the covariance as the following
Cov (X , Y )
ρXY = p
Var (X )Var (Y )
The advantage of the correlation is that, it is independent of
the scale, i.e., changing the variables’ scale of measurement
does not change the value of the correlation. Therefore, the
correlation is more useful as a measure of association between
two random variables than the covariance.
The empirical version is
s(X , Y )
rXY = √
sXX sYY
Amit (ISI, Chennai) Multivariate January 20, 2024 23 / 70
Covariance & Correlation
Covariance & Correlation
The correlation is in absolute value always less than 1. It is
zero if the covariance is zero and vice-versa. For
p-dimensional vectors (X1 , . . . , Xp )T we have the theoretical
correlation matrix
ρX1 X1 · · · ρX1 Xp
ρ = ... .. ..
. .
ρXp X1 · · · ρXp Xp
Amit (ISI, Chennai) Multivariate January 20, 2024 24 / 70
Covariance & Correlation
Covariance & Correlation
Theorem
If X and Y are independent, then ρ(X , Y ) = Cov (X , Y ) = 0
A In general, the converse is not true, as the following
example shows.
Example
Consider a standard normally-distributed random variable X and a
random variable Y = X 2 , which is surely not independent of X . Here
we have
Cov (X , Y ) = E (XY ) − E (X ).E (Y ) = E (X 3 ) = 0
(because E (X ) = 0 and E (X 2 ) = 1). Therefore ρ(X , Y ) = 0, as well.
This example also shows that correlations and covariances measure
only linear dependence. The quadratic dependence of Y = X 2 on X is
not reflected by these measures of dependence.
Amit (ISI, Chennai) Multivariate January 20, 2024 25 / 70
Covariance & Correlation
Covariance & Correlation
Remark
For two normal random variables, the converse of Theorem is true:
zero covariance for two normally-distributed random variables implies
independence.
Theorem enables us to check for independence between the
components of a bivariate normal random variable. That is,
we can use the correlation and test whether it is zero. The
distribution of rXY for an arbitrary (X , Y ) is unfortunately
complicated. The distribution of rXY will be more accessible if
(X , Y ) are jointly normal. If we transform the correlation by
Fisher’s Z -transformation, Skip TOH
1 1 + rXY
W = log
2 1 − rXY
Amit (ISI, Chennai) Multivariate January 20, 2024 26 / 70
Covariance & Correlation
Covariance & Correlation
We obtain a variable that has a more accessible distribution.
Under the hypothesis that ρ = 0, W has an asymptotic normal
distribution. Approximations of the expectation and variance
of W are given by the following:
1 1 + ρXY
E (W ) ≈ log
2 1 − ρXY
1
Var (W ) ≈
(n − 3)
The distribution is given in Theorem 2.2
Amit (ISI, Chennai) Multivariate January 20, 2024 27 / 70
Covariance & Correlation
Covariance & Correlation
Theorem
2.2
W − E (W ) D
Z= p −→ N(0, 1)
Var (W )
D
The symbol ”−→” denotes convergence in distribution.
Amit (ISI, Chennai) Multivariate January 20, 2024 28 / 70
Covariance & Correlation
Covariance & Correlation
Theorem
2.2
W − E (W ) D
Z= p −→ N(0, 1)
Var (W )
D
The symbol ”−→” denotes convergence in distribution.
Theorem allows us to test different hypotheses on correlation.
We can fix the level of significance α (the probability of
rejecting a true hypothesis) and reject the hypothesis if the
difference between the hypothetical value and the calculated
value of Z is greater than the corresponding critical value of
the normal distribution. The following example illustrates the
procedure.
Amit (ISI, Chennai) Multivariate January 20, 2024 28 / 70
Covariance & Correlation Test of correlation
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 29 / 70
Covariance & Correlation Test of correlation
Covariance & Correlation
Example
Let us consider again the pullovers data set from example. Consider
the correlation between the presence of the sales assistants (X4) vs. the
number of sold pullovers (X1 ) see Figure. Here we compute the
correlation as rX1 X4 = 0.633
The Z -transform of this value is
1 1 + rX1 X4
w = log = 0.746
2 1 − rX1 X4
The sample size is n√= 10, so for the hypothesis ρX1 X4 = 0, the statistic
to consider is: z = 7(0.746 − 0)
which is just statistically significant at the 5% level (i.e., 1.974 is just a
little larger than 1.96).
Amit (ISI, Chennai) Multivariate January 20, 2024 30 / 70
Covariance & Correlation Test of correlation
Covariance & Correlation
Example
Let us consider again the pullovers data set from example. Consider
the correlation between the presence of the sales assistants (X4) vs. the
number of sold pullovers (X1 ) see Figure. Here we compute the
correlation as rX1 X4 = 0.633
The Z -transform of this value is
1 1 + rX1 X4
w = log = 0.746
2 1 − rX1 X4
The sample size is n√= 10, so for the hypothesis ρX1 X4 = 0, the statistic
to consider is: z = 7(0.746 − 0)
which is just statistically significant at the 5% level (i.e., 1.974 is just a
little larger than 1.96).
Amit (ISI, Chennai) Multivariate January 20, 2024 30 / 70
Covariance & Correlation Test of correlation
Covariance & Correlation
Remark
The normalizing and variance stabilizing properties of W are
asymptotic. In addition the use of W in small samples (for n ≤ 25) is
improved by Hotelling’s transform
3W + tanh(W ) 1
W∗ = W − with Var (W ∗ ) =
4(n − 1) n−1
The transformed variable W ∗ is asymptotically distributed as a normal
distribution.
Amit (ISI, Chennai) Multivariate January 20, 2024 31 / 70
Covariance & Correlation Test of correlation
Covariance & Correlation
Remark
Note that the Fisher’s Z -transform is the inverse of the hyperbolic
tangent function: W = tanh−1 (rXY ); equivalently
2W
rXY = tanh(W ) = ee 2W −1
+1
.
Remark
Under the assumptions of normality of X and Y , we may test their
independence (ρXY = 0) using the exact t-distribution of the statistic
s
n − 2 ρXY =0
T = rXY 2
∼ tn−2
1 − rXY
Setting the probability of the first error type to α, we reject the null
hypothesis ρXY = 0 if |T | ≥ t1−α/2;n−2
Amit (ISI, Chennai) Multivariate January 20, 2024 32 / 70
Covariance & Correlation Need for plotting
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 33 / 70
Covariance & Correlation Need for plotting
Why plot?
Plots are important
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
Plots are important but frequently neglected.
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
Plots are important but frequently neglected. Lets look at
the following example of a small data set.
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
Plots are important but frequently neglected. Lets look at
the following example of a small data set.
x1 3 4 2 6 8 2 5
x2 5 5.5 4 7 10 5 7.5
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
Plots are important but frequently neglected. Lets look at
the following example of a small data set.
x1 5 4 6 2 2 8 3
x2 5 5.5 4 7 10 5 7.5
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
So the difference is only on the bivariate plot, NOT on the
marginal plots.
Amit (ISI, Chennai) Multivariate January 20, 2024 34 / 70
Covariance & Correlation Need for plotting
Why plot?
Pullovers Data subjected to MiniTab
Amit (ISI, Chennai) Multivariate January 20, 2024 35 / 70
Canonical Correlation
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 36 / 70
Canonical Correlation From Multiple Regression
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 37 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
Canonical Correlation Analysis is a generalisation of the
Multiple Regression problem.
Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
Canonical Correlation Analysis is a generalisation of the
Multiple Regression problem.
The Multiple Correlation Coefficient can also be
interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.
Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
Canonical Correlation Analysis is a generalisation of the
Multiple Regression problem.
The Multiple Correlation Coefficient can also be
interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.
The response vector x is partitioned into two parts so
that:
Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
Canonical Correlation Analysis is a generalisation of the
Multiple Regression problem.
The Multiple Correlation Coefficient can also be
interpreted as the measure of maximum correlation that
is attainable between the dependent variable and any
linear combination of the Independent Variables.
The response vector x is partitioned into two parts so
that:
x1 µ1 Σ11 Σ12
∼N ,
x2 µ2 Σ21 Σ12
Amit (ISI, Chennai) Multivariate January 20, 2024 38 / 70
Canonical Correlation From Multiple Regression
Introduction
If the q variables in x1 are measurements that are hard to
obtain or expensive to measure or not measurable but the
(p − q) variables are easy to measure, then we may attempt to
predict the difficult to measure x1 variables through the easy
to measure x2 variables.
Amit (ISI, Chennai) Multivariate January 20, 2024 39 / 70
Canonical Correlation From Multiple Regression
Introduction
If the q variables in x1 are measurements that are hard to
obtain or expensive to measure or not measurable but the
(p − q) variables are easy to measure, then we may attempt to
predict the difficult to measure x1 variables through the easy
to measure x2 variables.
The interrelationship that exist between the two sets, may
almost be described by the correlation between a few linear
combinations of the responses within each set of variables.
Amit (ISI, Chennai) Multivariate January 20, 2024 39 / 70
Canonical Correlation From Multiple Regression
Introduction
We’ve dealt with multiple regression, a case where we
used multiple independent variables to predict a single
dependent variable.
Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70
Canonical Correlation From Multiple Regression
Introduction
We’ve dealt with multiple regression, a case where we
used multiple independent variables to predict a single
dependent variable.
We came up with a linear combination of the predictors
that would result in the most variance accounted for in
the dependent variable (maximize R 2 )
Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70
Canonical Correlation From Multiple Regression
Introduction
We’ve dealt with multiple regression, a case where we
used multiple independent variables to predict a single
dependent variable.
We came up with a linear combination of the predictors
that would result in the most variance accounted for in
the dependent variable (maximize R 2 )
We will also see PC regression, in which linear
combinations of the predictors which account for all the
variance in the predictors can be used in lieu of the
original variables
Amit (ISI, Chennai) Multivariate January 20, 2024 40 / 70
Canonical Correlation From Multiple Regression
Introduction
Reason?
Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70
Canonical Correlation From Multiple Regression
Introduction
Reason?
Dimension reduction: select only the best components
Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70
Canonical Correlation From Multiple Regression
Introduction
Reason?
Dimension reduction: select only the best components
Independence of predictors: solves collinearity issue
Amit (ISI, Chennai) Multivariate January 20, 2024 41 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
The problem posed now is what to do if we had multiple
dependent variables
Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
The problem posed now is what to do if we had multiple
dependent variables
How could we come up with a way to understand the
relationship between a set of predictors and DVs?
Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
The problem posed now is what to do if we had multiple
dependent variables
How could we come up with a way to understand the
relationship between a set of predictors and DVs?
Canonical Correlation measures the relationship between
two sets of variables
Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70
Canonical Correlation From Multiple Regression
Multiple Dependent Variable
The problem posed now is what to do if we had multiple
dependent variables
How could we come up with a way to understand the
relationship between a set of predictors and DVs?
Canonical Correlation measures the relationship between
two sets of variables
Example: looking to see if different personality traits
correlate with performance scores for a variety of tasks
Amit (ISI, Chennai) Multivariate January 20, 2024 42 / 70
Canonical Correlation From Multiple Regression
Canonical Correlation
CC extends bivariate correlation, allowing 2+ continuous IVs (on
left) with 2+ continuous DVs (variables on right).
Focus is on correlations & weights
Main question: How are the best linear combinations of predictors
related to the best linear combinations of the DVs?
Amit (ISI, Chennai) Multivariate January 20, 2024 43 / 70
Canonical Correlation From Multiple Regression
The Analysis
In CC, there are several layers of analysis:
Correlation between pairs of canonical variates
Loadings between IVs & their canonical variates (on left)
Loadings between DVs & their canonical variates (on right)
Adequacy
Communalities
Other
Redundancy between IVs & can. variates on other (right) side
Redundancy between DVs & can. variates on other (left) side
As we will discuss later, these are problematic
Amit (ISI, Chennai) Multivariate January 20, 2024 44 / 70
Canonical Correlation From Multiple Regression
When to use
As exploratory tool to see if two sets of variables are related
If significant overall shared variance, several layers of analysis can
explore variables & variates involved
In a modelling type approach where you have a theoretical reason
for considering the variables as sets, and that one predicts another
See if one set of two or more variables relate longitudinally across
two time points
Variables at t1 are predictors; same variables at t2 are DVs
See layers of analysis for tentative causal evidence
Amit (ISI, Chennai) Multivariate January 20, 2024 45 / 70
Canonical Correlation From Multiple Regression
Comparison with Multiple Regression
With canonical correlation we will again (like in MR and PCA) be
creating linear combinations for our sets of variables
In MR, the multiple correlation coefficient was between the
predicted values created by the linear combination of predictors
(i.e. the composite) and the dependent variable
R 2 = square of multiple correlation
Amit (ISI, Chennai) Multivariate January 20, 2024 46 / 70
Canonical Correlation From Multiple Regression
Comparison with Multiple Regression
Creating linear composites of our respective variable sets (X , Y ):
Creating some single variable that represents the X s and another
single variable that represents the Y s.
Given a linear combination of X variables:
F = f1 X1 + f2 X2 + · · · + fp Xp
and a linear combination of Y variables:
G = g1 Y1 + g2 Y2 + · · · + gq Yq
The first canonical correlation is:
The maximum correlation coefficient between F and G , for all F
and G
The correlation we are interested in here will be between the linear
combinations (variates) created for both sets of variables
However now there will be the possibility for several ways in which
to combine the variables that may provide a useful interpretation
of the relationship
Amit (ISI, Chennai) Multivariate January 20, 2024 47 / 70
Canonical Correlation From Multiple Regression
Canonical Correlation : Pictorially
Amit (ISI, Chennai) Multivariate January 20, 2024 48 / 70
Canonical Correlation From Multiple Regression
Terminologies
Variables
Those measurements recorded in the dataset
Canonical Variates
The linear combination of the sets of variables (predictor and DV)
One variate for each set of variables
Canonical correlation
The correlation between variates
Pairs of Canonical Variates
As mentioned, we may have more than one pair of linear
combinations that provide an interesting measure of the relationship
Amit (ISI, Chennai) Multivariate January 20, 2024 49 / 70
Canonical Correlation From Multiple Regression
More about Can Cor
Canonical Correlation is one of the most general multivariate
forms – multiple regression, discriminate function analysis and
MANOVA are all special cases of it
As the basic output of regression and ANOVA are equivalent, so too
is the case here between canonical correlation and regression
Example: CanCorr squared between predictor set and a lone DV =
R2 from multiple regression
Only one solution
The number of canonical variate pairs you can have is equal to the
number of variables in the smaller set
When you have many variables on both sides of the equation you
end up with many canonical correlations.
Arranged in descending order, in most cases the first one or two
will be the ones of interest.
Amit (ISI, Chennai) Multivariate January 20, 2024 50 / 70
Canonical Correlation From Multiple Regression
Care need be taken
Number of canonical variate pairs
How many prove to be statistically/practically significant?
Interpreting the canonical variates
Where is the meaning in the combinations of variables created?
Importance of canonical variates
How strong is the correlation between the canonical variates?
What is the nature of the variate’s relation to the individual
variables in its own set? The other set?
Canonical variate scores
If one did directly measure the variate, what would the subjects’
scores be?
Amit (ISI, Chennai) Multivariate January 20, 2024 51 / 70
Canonical Correlation From Multiple Regression
Limitations
The procedure maximizes the correlation between the linear
combination of variables, however the combination may not make
much sense theoretically
Nonlinearity poses the same problem as it does in simple
correlation i.e. if there is a nonlinear relationship between the sets
of variables the technique is not designed to pick up on that
Cancorr is very sensitive to data involved i.e. influential cases and
which variables are chosen or left out of analysis can have
dramatic effects on the results just like they do elsewhere
Correlation, as we know, does not automatically imply causality
It is a necessary but not sufficient condition for causality
Being a simple correlation in the end, cancorr is often thought of
as a descriptive procedure
Amit (ISI, Chennai) Multivariate January 20, 2024 52 / 70
Canonical Correlation From Multiple Regression
Conditions & Assumptions
Sample size
Normality
Linearity
Homoscedasticity
Multicollinearity/Singularity
Outliers
Assumptions may better be checked before embarking on a big leap.
Amit (ISI, Chennai) Multivariate January 20, 2024 53 / 70
Canonical Correlation From Multiple Regression
Data and Computations
Canonical Correlation uses the correlations from the raw data and uses
this as input data
R11 R12
R21 R12
The Canonical Correlation Matrix R is obtained by
−1 −1
R = RYY RYX RXX RXY
To get the canonical correlations, you get the eigenvalues of R and take
the square root
p
rci2 = λi ⇒ rci = λi
Consequently the eigenvector corresponding to each eigenvalue is
transformed into the coefficients that specify the linear combination
that will make up a canonical variate
Amit (ISI, Chennai) Multivariate January 20, 2024 54 / 70
Canonical Correlation From Multiple Regression
Data and Computations
Amit (ISI, Chennai) Multivariate January 20, 2024 55 / 70
Marginal & Conditional density functions
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 56 / 70
Marginal & Conditional density functions
Distribution & Density Functions
Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70
Marginal & Conditional density functions
Distribution & Density Functions
Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by
F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )
Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70
Marginal & Conditional density functions
Distribution & Density Functions
Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by
F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )
For continuous X , there exists a nonnegative probability
density function (pdf ) f , such that
Rx
F (x) = −∞ f (u)du
Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70
Marginal & Conditional density functions
Distribution & Density Functions
Let X = (X1 , X2 , . . . , Xp )T be a random vector. The cumulative
distribution function (cdf ) of X is defined by
F (x) = P(X ≤ x) = P(X1 ≤ x1 , X2 ≤ x2 , . . . , Xp ≤ xp )
For continuous X , there exists a nonnegative probability
density function (pdf ) f , such that
Rx
F (x) = −∞ f (u)du
R∞
Note that −∞ f (u)du = 1
Most of the
R x integrals appearing below are multidimensional. For
instance, −∞ f (u)du means
Z xp Z x1
··· f (u1 , . . . , up )du1 · · · dup
−∞ −∞
Amit (ISI, Chennai) Multivariate January 20, 2024 57 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
Note also that the cdf F is differentiable with
δ p F (x)
f (x) =
δx1 · · · δxp
Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
Note also that the cdf F is differentiable with
δ p F (x)
f (x) =
δx1 · · · δxp
For discrete X , the values of this random variable are
concentrated on a countable or finite set of points {cj }j∈J
Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
Note also that the cdf F is differentiable with
δ p F (x)
f (x) =
δx1 · · · δxp
For discrete X , the values of this random variable are
concentrated on a countable or finite set of points {cj }j∈J
The probability of eventsPof the form {X ∈ D} can then be
computed as P(X ∈ D) = {j:cj ∈D} P(X = cj )
If we partition X as X T = (X1T , X2T ) with X1 ∈ R k and X2 ∈ R p−k ,
then the function FX1 (x1 ) = P(X1 ≤ x1 ) = F (x11 . . . x1k , ∞, . . . , ∞) is
called the marginal cdf.
F = F (x) is called the joint cdf.
Amit (ISI, Chennai) Multivariate January 20, 2024 58 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
For continuous X the marginal pdf can be computed from the
joint density by integrating out the variable not of interest.
Z ∞
fX1 (x1 ) = f (x1 , x2 )dx2
−∞
The conditional pdf of X2 given X1 = x1 is given as
f (x2 |x1 ) = ffX(x1(x,x12))
1
Amit (ISI, Chennai) Multivariate January 20, 2024 59 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
Consider the pdf
1
+ 32 x2
2 x1 0 ≤ x1 , x2 ≤ 1,
f (x1 , x2 ) =
0 otherwise
f (x1 , x2 ) is a density since
1 1
x12 x22
Z
1 3 1 3
f (x1 , x2 )dx1 x2 = + = + =1
2 2 0 2 2 0 4 4
The marginal distributions are
Z Z 1
1 3 1 3
fX1 (x1 ) = f (x1 , x2 )dx2 = x1 + x2 dx2 = x1 +
0 2 2 2 4
Z Z 1
1 3 3 1
fX2 (x2 ) = f (x1 , x2 )dx1 = x1 + x2 dx1 = x2 +
0 2 2 2 4
Amit (ISI, Chennai) Multivariate January 20, 2024 60 / 70
Marginal & Conditional density functions
Marginal & Conditional pdf
The conditional densities are therefore
1 3 1 3
2 x1 + 2 x2 2 x1 + 2 x2
f (x2 |x1 ) = 1 3
and f (x2 |x1 ) = 3 1
2 x1 + 4 2 x2 + 4
Note that these conditional pdf ’s are nonlinear in x1 and x2
although the joint pdf has a simple (linear) structure.
Amit (ISI, Chennai) Multivariate January 20, 2024 61 / 70
Matrix Notations
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 62 / 70
Matrix Notations
Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
matrix.
x11 · · · x1p
.. ..
. .
X(n×p) = .
..
. . .
x1n · · · xnp
Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70
Matrix Notations
Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
matrix.
x11 · · · x1p
.. ..
. .
X(n×p) = .
..
.. .
x1n · · · xnp
The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a
p-dimensional
random variable X ∈ Rp .
x¯1
..
X̄ = . = n−1 X T 1n .
x¯p
Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70
Matrix Notations
Summary Statistics
Let X ∈ Rp . Let X represent n realisations of X , hence X is an
(n × p) data
matrix.
x11 · · · x1p
.. ..
. .
X(n×p) = .
..
.. .
x1n · · · xnp
The rows xi = (xi1 , . . . , xip ) ∈ Rp denote the i-th observation of a
p-dimensional
random variable X ∈ Rp .
x¯1
..
X̄ = . = n−1 X T 1n .
x¯p
The sample varinace-covariance matrix is given by:
S = n−1 X T X − x̄ x̄ T = n−1 (X T X − n−1 X T 1n 1n X )
Amit (ISI, Chennai) Multivariate January 20, 2024 63 / 70
MultiNormal Distribution
Contents
1 Introduction
Examples
2 Covariance & Correlation
Test of correlation
Need for plotting
3 Canonical Correlation
From Multiple Regression
4 Marginal & Conditional density functions
5 Matrix Notations
6 MultiNormal Distribution
Amit (ISI, Chennai) Multivariate January 20, 2024 64 / 70
MultiNormal Distribution
Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .
Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70
MultiNormal Distribution
Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .
Definition
Multivariate normal :
x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn
Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70
MultiNormal Distribution
Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .
Definition
Multivariate normal :
x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn
However the Multivariate Normal distribution could also be
looked upon as a generalization of univariate normal
distribution to p ≥ 2 dimensions.
Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70
MultiNormal Distribution
Definition, properties
Let Σ = (σij ) ∈ Rn×n be symmetric, positive definite and µ ∈ Rn .
Definition
Multivariate normal :
x ∼ Nn (µ, Σ) iff t ′ x ∼ N(t ′ µ, t ′ Σt), ∀t ∈ Rn
However the Multivariate Normal distribution could also be
looked upon as a generalization of univariate normal
distribution to p ≥ 2 dimensions.
1 x−µ 2
f (x) = √ e( σ ) −∞<x <∞
2πσ 2
Amit (ISI, Chennai) Multivariate January 20, 2024 65 / 70
MultiNormal Distribution
Definition, properties
The term
2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ
Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70
MultiNormal Distribution
Definition, properties
The term
2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ
in the exponent of the univariate normal density function
measures the distance of x from µ in standard deviation units.
Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70
MultiNormal Distribution
Definition, properties
The term
2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ
in the exponent of the univariate normal density function
measures the distance of x from µ in standard deviation units.
This could be generalized for a p × 1 vector x of observations
on several variables as
Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70
MultiNormal Distribution
Definition, properties
The term
2
x −µ
= (x − µ)(σ 2 )−1 (x − µ)
σ
in the exponent of the univariate normal density function
measures the distance of x from µ in standard deviation units.
This could be generalized for a p × 1 vector x of observations
on several variables as
(x − µ)′ Σ−1 (x − µ)
The p × 1 vector µ represents the expected value of random
vector X, and the p × p matrix Σ is its variance-covariance
matrix, symmetric and positive definite, in order that the
above represents the generalized distance from x to µ
Amit (ISI, Chennai) Multivariate January 20, 2024 66 / 70
MultiNormal Distribution
Definition, properties
In the multivariate normal density we replace the univariate
distance by the generalized mutivariate distance.
Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70
MultiNormal Distribution
Definition, properties
In the multivariate normal density we replace the univariate
distance by the generalized mutivariate distance.
The constant also needs a replacement by a more general
form and this turns out to be (2π)−p/2 |Σ|1/2 .
Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70
MultiNormal Distribution
Definition, properties
In the multivariate normal density we replace the univariate
distance by the generalized mutivariate distance.
The constant also needs a replacement by a more general
form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :
Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70
MultiNormal Distribution
Definition, properties
In the multivariate normal density we replace the univariate
distance by the generalized mutivariate distance.
The constant also needs a replacement by a more general
form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :
1 1 ′ −1 (x−µ)
f (x) = e − 2 (x−µ) Σ
(2π)−p/2 |Σ|1/2
where −∞ < xi < ∞, i = 1, 2, . . . , p and this will be denoted by
Np (µ, Σ).
Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70
MultiNormal Distribution
Definition, properties
In the multivariate normal density we replace the univariate
distance by the generalized mutivariate distance.
The constant also needs a replacement by a more general
form and this turns out to be (2π)−p/2 |Σ|1/2 .
So the p-dimensional normal density for the random vector
X = [X1 , X2 , . . . , Xp ]′ is given by :
1 1 ′ −1 (x−µ)
f (x) = e − 2 (x−µ) Σ
(2π)−p/2 |Σ|1/2
where −∞ < xi < ∞, i = 1, 2, . . . , p and this will be denoted by
Np (µ, Σ).
Amit (ISI, Chennai) Multivariate January 20, 2024 67 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
σ11 σ12
Σ=
σ12 σ22
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
σ11 σ12
Σ=
σ12 σ22
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
σ11 σ12
Σ=
σ12 σ22
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11
Correlation coefficient ρ12 help us write
√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
σ11 σ12
Σ=
σ12 σ22
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11
Correlation coefficient ρ12 help us write
√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12
So the squared distance (x − µ)′ Σ−1 (x − µ) simplifies as follows :
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
Let p = 2. E (X1 ) = µ1 , E (X2 ) = µ2 , Var (X1 ) = σ11 , Var (X2 ) = σ22
√ √
and Corr (X1 , X2 ) = σ12/( σ11 σ22 ) = ρ12 .
The inverse of covariance matrix
σ11 σ12
Σ=
σ12 σ22
1 σ22 −σ12
Σ−1 = 2
σ11 σ22 − σ12 −σ12 σ11
Correlation coefficient ρ12 help us write
√ 2 = σ σ (1 − ρ2 )
σ12 = ρ12 σ11 σ22 ⇒ σ11 σ22 − σ12 11 22 12
So the squared distance (x − µ)′ Σ−1 (x − µ) simplifies as follows :
Amit (ISI, Chennai) Multivariate January 20, 2024 68 / 70
MultiNormal Distribution
Example : Bivariate
√
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2
Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70
MultiNormal Distribution
Example : Bivariate
√
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2
√
σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )
Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70
MultiNormal Distribution
Example : Bivariate
√
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2
√
σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )
" 2 2 #
1 x1 − µ1 x2 − µ2 x1 − µ1 x2 − µ2
= √ + √ − 2ρ12 √ √
1 − ρ212 σ11 σ22 σ11 σ22
Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70
MultiNormal Distribution
Example : Bivariate
√
1 σ22 −ρ12 σ11 σ22 x1 − µ 1
[x1 −µ1 , x2 −µ2 ] √
σ11 σ22 (1 − ρ212 ) −ρ12 σ11 σ22 σ11 x 2 − µ2
√
σ22 (x1 − µ1 )2 + σ11 (x2 − µ2 )2 − 2ρ12 σ11 σ22 (x1 − µ1 )(x2 − µ2 )
=
σ11 σ22 (1 − ρ212 )
" 2 2 #
1 x1 − µ1 x2 − µ2 x1 − µ1 x2 − µ2
= √ + √ − 2ρ12 √ √
1 − ρ212 σ11 σ22 σ11 σ22
Amit (ISI, Chennai) Multivariate January 20, 2024 69 / 70
MultiNormal Distribution
Bivariate Normal Density
Amit (ISI, Chennai) Multivariate January 20, 2024 70 / 70