0% found this document useful (0 votes)
4 views41 pages

Correlation and Regression Analysis Guide

Uploaded by

ireshaperera337
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views41 pages

Correlation and Regression Analysis Guide

Uploaded by

ireshaperera337
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Correlation and Regression

Prof Nihal Hennayake


Introduction
• Understand relationships between variables and use
them for prediction and decision-making in
management contexts.
Learning Objectives
• - Define correlation and regression analysis.
• - Understand how to calculate and interpret correlation.
• - Perform simple and multiple regression analyses.
• - Apply these methods in real-world management
scenarios.
What is Correlation?
• Measures strength and direction of a linear
relationship between two quantitative variables.
• Range: -1 to +1 (Positive, Negative, or No correlation).
Types of Correlation
• Positive: Advertising ↑ → Sales ↑
• Negative: Turnover ↑ → Job Satisfaction ↓
• Zero: Coffee consumption and employee height.
Example: Correlation in Management
• Relationship between employee training hours (X) and
productivity (Y).
• Example data:
Employer Employee Productivity
Training
Hours
A 5 50
B 8 65
C 10 70
D 12 80
E 15 90
n=5
∑𝑋=5+8+10+12+15=50
∑𝑌=50+65+70+80+90=355
Xˉ=n∑X​=550​=10.0

Yˉ=n∑Y​=5355​=71.0
(x) (y) (x-bar X) (y-bar Y) (x-bar X)(y-bar Y) x-bar X^2) (y-bar Y)^2

5 50 (5-10=-5) (50-71=-21) (-5)(-21)=105 (25) (441)

8 65 (8-10=-2) (65-71=-6) (-2)(-6)=12 (4) (36)

10 70 (10-10=0) (70-71=-1) (0).(1)=0 (0) (1)

12 80 (12-10=2) (80-71=9) (2)(9)=18 (4) (81)

(5)(19)=95 (25) (361)


15 90 (15-10=5) (90-71=19)
230 58 920
Calculating Correlation
• Pearson's

• Example result: r = 0.97 (Strong positive correlation).


Interpretation of Correlation
• r = 0.97 → Strong positive relationship.
• As training hours increase, productivity increases.
• Managerial insight: Invest more in employee training.
Regression Analysis
Year X Y
1 10 44
Scatter Diagram
2 9 40
3 11 42
4 12 46
5 11 48
6 12 52
7 13 54
8 13 58
9 14 56
10 15 60
Regression Analysis
Regression Line: Line of Best Fit

Regression Line: Minimizes the sum of the


squared vertical deviations (et) of each point
from the regression line.

Ordinary Least Squares (OLS) Method


Regression Analysis
Ordinary Least Squares (OLS)

Model: Yt a  bX t  et

ˆ
Yˆt aˆ  bX t

et Yt  Yˆt
Ordinary Least Squares (OLS)
Estimation Example
Time Xt Yt Xt  X Yt  Y ( X t  X )(Yt  Y ) ( X t  X )2
1 10 44 -2 -6 12 4
2 9 40 -3 -10 30 9
3 11 42 -1 -8 8 1
4 12 46 0 -4 0 0
5 11 48 -1 -2 2 1
6 12 52 0 2 0 0
7 13 54 1 4 4 1
8 13 58 1 8 8 1
9 14 56 2 6 12 4
10 15 60 3 10 30 9
120 500 106 30
n n n
106
n 10  X t 120
t 1
 Yt 500
t 1
 (X
t 1
t  X ) 2 30 bˆ 
30
3.533

n n n
X t 120 Yt 500 aˆ 50  (3.533)(12) 7.60
X   12 Y   50  (X t  X )(Yt  Y ) 106
t 1 n 10 t 1 n 10 t 1
Ordinary Least Squares (OLS)
Estimation Example
n
X t 120
n 10 X   12
t 1 n 10
n
n n Yt 500
X 120 Y 500 Y   50
t 1 n 10
t t
t 1 t 1

n
ˆ 106
 (X
t 1
t
2
 X ) 30 b
30
3.533

 (X
t 1
t  X )(Yt  Y ) 106 aˆ 50  (3.533)(12) 7.60
Tests of Significance

Standard Error of the Slope Estimate

sbˆ 
 (Yt  Y )
ˆ 2


e 2
t

( n  k ) ( X t  X ) 2
(n  k ) ( X t  X )2
Tests of Significance
Example Calculation
Time Xt Yt Yˆt et Yt  Yˆt et2 (Yt  Yˆt ) 2 ( X t  X )2
1 10 44 42.90 1.10 1.2100 4
2 9 40 39.37 0.63 0.3969 9
3 11 42 46.43 -4.43 19.6249 1
4 12 46 49.96 -3.96 15.6816 0
5 11 48 46.43 1.57 2.4649 1
6 12 52 49.96 2.04 4.1616 0
7 13 54 53.49 0.51 0.2601 1
8 13 58 53.49 4.51 20.3401 1
9 14 56 57.02 -1.02 1.0404 4
10 15 60 60.55 -0.55 0.3025 9
65.4830 30

n n n
 t
(Y  Yˆ ) 2
65.4830
  t t
e 
t 1
(Y2
t Yˆ ) 2

t 1
65.4830  t 30
( X
t 1
 X ) 2 sbˆ 
(n  k ) ( X t  X ) 2

(10  2)(30)
0.52
Tests of Significance
Example Calculation
n n

 t  t t 65.4830
e 2

t 1
 (Y  Yˆ ) 2

t 1
n

 t
( X
t 1
 X ) 2
30

sbˆ 
 (Yt  Y )
ˆ 2


65.4830
0.52
(n  k ) ( X t  X ) 2
(10  2)(30)
Tests of Significance
Calculation of the t Statistic

bˆ 3.53
t  6.79
sbˆ 0.52

Degrees of Freedom = (n-k) = (10-2) = 8


Critical Value at 5% level =2.306
Tests of Significance
Decomposition of Sum of Squares

Total Variation = Explained Variation + Unexplained Variation

 (Yt  Y )  (Y  Y )   (Yt  Yt )
2ˆ 2 ˆ 2
Tests of Significance
Decomposition of Sum of Squares
Tests of Significance

Coefficient of Determination

R2 
Explained Variation

 (Yˆ  Y ) 2

TotalVariation  t
(Y  Y ) 2

2 373.84
R  0.85
440.00
Tests of Significance

Coefficient of Correlation

r  R 2 with the sign of bˆ

 1 r 1

r  0.85 0.92
Multiple Regression Analysis

Y a  b1 X 1  b2 X 2    bk ' X k '
Model:

Adjusted Coefficient of Determination

2 2 (n  1)
R 1  (1  R )
(n  k )
Adjusted R-square should be used to measure "goodness
of fit" in a model that contains more than one
independent variable, since simple R-square systematically
overstates goodness of fit in model with more than
one independent variable.

Adjusted R2 is a modification of R2 that adjusts for the


number of explanatory terms in a model.
Adjusted R2 is used to compensate for the addition of
variables to the model. As more independent variables
are added to the regression model, unadjusted R2 will
generally increase but there will never be a
decrease. This will occur even when the additional
variables do little to help explain the dependent
variable. To compensate for this, adjusted R2 is corrected
for the number of independent variables in the
model. The result is an adjusted R2 can go up or down
depending on whether the addition of another variable
adds or does not add to the explanatory power of the
model. Adjusted R2 will always be lower than unadjusted.
Multiple Regression Analysis

Analysis of Variance and F Statistic

Explained Variation /(k  1)


F
Unexplained Variation /(n  k )

R 2 /(k  1)
F
(1  R 2 ) /( n  k )
Problems in Regression Analysis
Multicollinearity: Two or more explanatory
variables are highly correlated.
Heteroskedasticity: Variance of error term is not
independent of the Y variable.
Autocorrelation: Consecutive error terms are
correlated.
What is multicollinearity?
Multicollinearity exists whenever two or more of the
predictors in a regression model are moderately or highly
correlated.

why can't a researcher just collect his data in such a way


to ensure that the predictors aren't highly correlated?
(Then, multicollinearity wouldn't be a problem)

Unfortunately, researchers often can't control the


predictors. Obvious examples include a person's gender,
race, grade point average, IQ, and starting salary. For each
of these predictor examples, the researcher just observes
the values as they occur for the people in her random
sample.
Causes of Multicollinearity

• Improper use of dummy variables (e.g. failure to exclude


one category)

• Including a variable that is computed from other variables


in the equation (e.g. family income = husband’s income +
wife’s income, and the regression includes all 3 income
measures)

• In effect, including the same or almost the same variable


twice (height in feet and height in inches)
Consequences of multicollinearity
• Even extreme multicollinearity does not violate OLS
assumptions. OLS estimates are still unbiased and BLUE
(Best Linear Unbiased Estimators)

• Nevertheless, the greater the multicollinearity, the greater


the standard errors. When high multicollinearity is present,
confidence intervals for coefficients tend to be very wide
and t statistics tend to be very small. Coefficients will have
to be larger in order to be statistically significant, i.e. it will
be harder to reject the null when multicollinearity is
present.
Heteroskedasticity

What is heteroskedasticity?

Non-uniform variance or var(ui) = E(ui2) = σi2

Why do we have heteroskedasticity?

Variance may be positively related to independent variable


(for example there may be a larger variation in consumption
for people with high incomes than for people with low
incomes).
Consequences of heteroskedasticity
- The βOLS estimator remains unbiased but will not have the
minimum variance among all linear unbiased estimators
(because less precise high variance and more precise low-
variance observations are given equal weight)
- The σ2
OLS and var(βOLS) estimators are biased → t-statistics and
F-statistics are invalid
How to detect heteroskedasticity
- Plot, economic theory, or statistical tests
What can we do?
- If we can recognize a pattern in σi2, use Feasible Generalized
Least Squares (FGLS)
- If we cannot recognize a pattern in σi2, use robust
estimators
Autocorrelation refers to the correlation of a time series with
its own past and future values.
Autocorrelation is also sometimes called “lagged correlation”
or “serial correlation”, which refers to the correlation
between members of a series of numbers arranged in time.

A time series is a sequence of observations on a variable


over time. Macroeconomists generally work with time series
(e.g., quarterly observations on GDP and monthly
observations on the unemployment rate).

You might also like