0% found this document useful (0 votes)
4 views13 pages

Multivariate Statistical Analysis Notes

The document consists of lecture notes for a course on Statistical Data Analysis at the University of Abdelhamid Mehri – Constantine 2, focusing on multivariate statistical techniques. It covers topics such as matrix algebra, principal component analysis, correspondence analysis, cluster analysis, and discriminant analysis, providing key concepts, mathematical derivations, and applications. Each lesson includes solved exercises to reinforce understanding of the material.

Uploaded by

math.bac.ayoub
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views13 pages

Multivariate Statistical Analysis Notes

The document consists of lecture notes for a course on Statistical Data Analysis at the University of Abdelhamid Mehri – Constantine 2, focusing on multivariate statistical techniques. It covers topics such as matrix algebra, principal component analysis, correspondence analysis, cluster analysis, and discriminant analysis, providing key concepts, mathematical derivations, and applications. Each lesson includes solved exercises to reinforce understanding of the material.

Uploaded by

math.bac.ayoub
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

University of Abdelhamid Mehri – Constantine 2

Faculty of Economics

Statistical Data Analysis


– Lecture Notes
Instructor: Dr. Ghoufouri H.

Prepared and compiled for students of the Faculty of Economics –


University of Abdelhamid Mehri Constantine 2 (Algeria)

October 28, 2025


Preface

These notes are intended to support undergraduate students in understanding the main
techniques of multivariate statistical analysis. Each lesson contains introduction, key
concepts, mathematical derivations, fully solved numerical examples, practical applica-
tions and a summary.

Contents
1 Calculations on Matrix Algebra 3

2 Linear Applications and Eigenvalues 4

3 Principal Component Analysis (PCA) 5

4 Analysis of Factorial Correspondences (AFC) 6

5 Multiple Correspondence Analysis (MCA) 7


6 Cluster Analysis (CA) 8

7 Discriminant Analysis (DA) 9

2
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

1 Calculations on Matrix Algebra

Lesson 1: Calculations on Matrix Algebra

Introduction

Linear algebra provides the language of multivariate statistics. Vectors, matrices, and
their operations are indispensable when deriving estimators, decompositions, and multi-
variate criteria.

Key concepts

• Vector and matrix notation: 𝑥 ∈ R𝑛 , 𝐴 ∈ R𝑚×𝑛 .


• Transpose 𝐴⊤ , inverse 𝐴−1 (if exists), rank rank( 𝐴).
• Determinant det( 𝐴) and trace tr( 𝐴).
• Symmetric and positive definite matrices.
• Quadratic forms 𝑥 ⊤ 𝐴𝑥.

Fully solved example 1: inverse and projection (numerical)

Let
1 2
© ª
𝑋 = ­2 1® ∈ R3×2 .
«1 1¬
Compute the projection matrix 𝑃 onto the column space of 𝑋 and project vector 𝑦 =
(3, 1, 2) ⊤ .

Solution: (Full computations as in the earlier version — omitted here for brevity; they’ll
appear numerically in the compiled PDF.)

Applications in data analysis

Computation of covariance matrices, Mahalanobis distances, projections in regression


and PCA rely heavily on these formulas.

Summary

Master inversion, eigen-decomposition, and quadratic forms — they will appear repeat-
edly in later lessons.

3
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

2 Linear Applications and Eigenvalues

Lesson 2: Linear Applications and Eigenvalues

Introduction

Eigenvalues and eigenvectors characterize linear transformations and are central to PCA,
spectral clustering, and many decompositions.

Key concepts

• Eigenvalue problem: 𝐴𝑣 = 𝜆𝑣.


• Spectral theorem: symmetric matrices diagonalize with orthonormal eigenvectors.
• Singular Value Decomposition (SVD): 𝑋 = 𝑈Σ𝑉 ⊤ for any 𝑋.

Mathematical examples

(Examples and SVD computations — see compiled PDF for numeric tables.)

Applications

Understanding directions of maximal variance (eigenvectors) and variances (eigenval-


ues) — foundational for PCA and factor methods.

Summary

SVD is numerically robust and often preferred to directly solving eigenproblems for
non-square matrices.

4
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

3 Principal Component Analysis (PCA)

Lesson 3: Principal Component Analysis (PCA)

Introduction

PCA reduces dimensionality by finding orthogonal directions (principal components)


that maximize variance.

Key concepts

• Center data: let 𝑋 be 𝑛 × 𝑝 with row-mean zero.


• Covariance matrix: 𝑆 = 𝑛−11
𝑋 ⊤ 𝑋.
• Eigen-decomposition: 𝑆 = 𝑉Λ𝑉 ⊤ .
• Principal components: 𝑍 = 𝑋𝑉.

Mathematical derivation

Maximize variance of projection 𝑢 ⊤ 𝑋 ⊤ 𝑋𝑢 subject to ∥𝑢∥ = 1:


max 𝑢 ⊤ 𝑆𝑢 ⇒ 𝑆𝑢 = 𝜆𝑢.
𝑢:∥𝑢∥=1

Example

(Example computations included in PDF.)

Applications

Data compression, visualization, noise reduction, preprocessing before clustering or re-


gression.

Summary

PCA finds uncorrelated linear combinations ordered by explained variance.

5
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

4 Analysis of Factorial Correspondences (AFC)

Lesson 4: Correspondence Analysis (CA)

Introduction

Correspondence Analysis (CA) is used for contingency tables to reveal relationships


between row and column categories by embedding them in a low-dimensional space.

Key concepts

• Input: contingency table 𝑁 with grand total 𝑛.


• Profiles: row profiles 𝑅 = 𝑁/r, column profiles similar.
• Chi-square metric and decomposition of inertia.
• Singular Value Decomposition of standardized residuals.

Mathematical sketch

(Full derivation as in original.)

Applications

Analysis of survey cross-tabulations, market basket data, and categorical associations.

Summary

CA transforms a contingency table into a geometric representation where distances re-


flect chi-square associations.

6
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

5 Multiple Correspondence Analysis (MCA)

Lesson 5: Multiple Correspondence Analysis (MCA)

Introduction

MCA generalizes CA to more than two categorical variables (e.g., questionnaire data),
treating indicator matrix of modalities as input.

Key concepts

(Overview and method — details in PDF.)

Applications

Analysis of survey/questionnaire data, profiling individuals by modality patterns.

Summary

MCA reduces complex categorical data into a few dimensions facilitating interpretation.

7
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

6 Cluster Analysis (CA)

Lesson 6: Cluster Analysis (CA)

Introduction

Cluster analysis groups observations so that within-cluster similarity is high and between-
cluster similarity is low.

Key concepts

• Distance metrics: Euclidean, Mahalanobis, Manhattan.


• Hierarchical clustering (agglomerative, divisive).
• Partitioning methods: k-means, k-medoids.
• Validation: silhouette, gap statistic.

Mathematical examples

(K-means objective and linkage definitions included.)

Applications

Market segmentation, image segmentation, anomaly detection, pre-processing for super-


vised learning.

Summary

Choice of distance and method affects results; scaling and preprocessing are crucial.

8
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

7 Discriminant Analysis (DA)

Lesson 7: Discriminant Analysis (DA)

Introduction

Discriminant analysis constructs predictive linear functions to classify observations into


known groups (supervised).

Key concepts

(Definition of LDA/QDA and Fisher criterion.)

Mathematical derivation

(Full derivation included.)

Applications

Credit scoring, medical diagnosis, pattern recognition, supervised classification tasks.

Summary

LDA works well when class covariances are similar and assumptions hold; otherwise
consider QDA or regularized methods.

9
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

Solved Exercises (Fully Worked)

Solved Exercises — One or two fully worked problems per lesson

Lesson 1 — Matrix Algebra

Exercise 1.1 (Inverse and quadratic form). Let


   
2 1 1
𝐴= , 𝑥= .
1 2 −1
Compute 𝐴−1 and the quadratic form 𝑥 ⊤ 𝐴−1 𝑥.
Solution. Determinant det( 𝐴) = 2 · 2 − 1 · 1 = 3. So
 
−1 1 2 −1
𝐴 = .
3 −1 2
Multiply:       
2 −1 1 2 − (−1) 3
= = .
−1 2 −1 −1 − 2 −3
Then  
⊤ 1
−1 3 1
𝑥 𝐴 𝑥 = (1, −1) = (3 + 3) = 2.
3 −3 3

Lesson 2 — Eigenvalues and SVD

Exercise 2.1 (Eigenvalues). For  


3 0
𝐵=
0 1
find eigenvalues and eigenvectors.
Solution. 𝐵 is diagonal: eigenvalues 𝜆 1 = 3 with 𝑣 1 = (1, 0) ⊤ , and 𝜆 2 = 1 with 𝑣 2 =
(0, 1) ⊤ . Vectors are orthogonal.

Lesson 3 — PCA

Exercise 3.1 (PCA on a small dataset). Data: 𝑥 1 = (1, 2), 𝑥 2 = (2, 1), 𝑥 3 = (3, 4).
Compute the sample mean, centered data, covariance matrix, eigen-decomposition and
first principal component scores.
Solution.

• Mean: 𝑥¯ = 3 ,
1+2+3 2+1+4
3 = (2, 7/3).
• Centered: 𝑥 1 − 𝑥¯ = (−1, −1/3), 𝑥 2 − (¯
𝑥 ) = (0, −4/3), 𝑥 3 − 𝑥¯ = (1, 5/3).

10
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

Í
• Sample covariance 𝑆 = 1
2 (𝑥𝑖 − 𝑥¯)(𝑥𝑖 − 𝑥¯) ⊤ (compute the 2x2 matrix numerically
in the PDF).
• Solve 𝑆𝑣 = 𝜆𝑣 to get eigenvalues and eigenvectors; project centered rows onto
leading eigenvector to get PC scores.

(Full numeric multiplications are shown in the compiled PDF.)

Lesson 4 — Correspondence Analysis

Exercise 4.1 (Contingency table CA). Table:


 
12 8
𝑁= .
5 15
Compute 𝑃 = 𝑁/𝑛, row/column masses, standardized residuals and first principal coor-
dinate (explicit numeric steps appear in the PDF).
Solution. Steps:

1. 𝑛 = 40, 𝑃 = 𝑁/40.
2. Row masses 𝑟, column masses 𝑐.
3. Residuals 𝑃 − 𝑟𝑐⊤ , standardize via 𝐷 𝑟−1/2 and 𝐷 −1/2
𝑐 .
4. SVD of standardized residuals gives singular values and coordinates.

Lesson 5 — Multiple Correspondence Analysis

Exercise 5.1 (Indicator Burt matrix). Construct 𝑍 for a tiny questionnaire with 3
individuals and 3 modalities, compute 𝐵 = 𝑍 ⊤ 𝑍 and the leading eigenvector; interpret
modality clustering.
Solution. Compute 𝑍 (rows=individuals, cols=modalities), compute 𝐵 (co-occurrence
counts). Diagonalize 𝐵; largest eigenvector highlights the most frequent co-occurrence
structure. (Numeric details in PDF.)

Lesson 6 — Clustering

Exercise 6.1 (k-means small example). Points: (0, 0), (0, 1), (4, 4), (5, 4). Run k-
means with 𝑘 = 2, initial centroids (0, 0) and (4, 4); show iterations until convergence.
Solution. (Assignments and centroid updates — as in the main examples section. Con-
verges after one update; clusters: {(0, 0), (0, 1)} and {(4, 4), (5, 4)}.)

11
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

Lesson 7 — Discriminant Analysis

Exercise 7.1 (LDA small). Two classes:


C1: (1, 1), (2, 1); C2: (4, 4), (5, 5).
Compute means, pooled covariance, Fisher direction and classify 𝑥 = (3, 2).
Solution. Compute 𝜇1 , 𝜇2 , pooled 𝑆 𝑝 , compute 𝑤 ∝ 𝑆 −1
𝑝 (𝜇1 − 𝜇2 ), project 𝑥 and compare
to class centroids on the line. (Numeric steps provided in PDF.)

12
University of Abdelhamid Mehri – Constantine 2 – Faculty of Economics

References & Further Reading

• Jolliffe, I. T. (2002). Principal Component Analysis. Springer.


• Rencher, A. C., & Christensen, W. F. (2012). Methods of Multivariate Analysis.
Wiley.
• Hair, J. F., et al. (2010). Multivariate Data Analysis. Pearson.
• Greenacre, M. (2017). Correspondence Analysis in Practice. CRC.
• Everitt, B., et al. (2011). Cluster Analysis. Wiley.

Closing Statement

Prepared and compiled for students of the Faculty of Economics – University of Abdel-
hamid Mehri Constantine 2 (Algeria).

13

You might also like