Regression Analysis Basics and Applications
Regression Analysis Basics and Applications
© Faculty of Management
Regression
# of doctors
# of TVs
2
Covariance and
Correlation
© Faculty of Management
Variance of multiple dependent r.v.’s
E.g.: Stock prices – if economy is good, all stocks tend to move up.
4
Covariance and correlation
Variance must account for “complementarity” or “together-ness” of X1 and X2
If X1 and X2 are complements, both move in the same direction
If X1 and X2 are substitutes, they move in opposite directions
Need a measure of “together-ness” of 2 random variables.
Covariance:
Cov(X,Y) = E[(X-μX)(Y-μY)]= E[XY] – E[X]E[Y]= μXY - μXμY
Correlation:
covariance divided by standard deviations of the variables
Cov( X , Y )
( X ,Y ) =
X Y
5
Covariance
Sample covariance:
( xi − X )( yi − Y )
n
Cov( X , Y ) =
i =1 n −1
6
Covariance
7
Correlation
𝑥𝑖 − 𝑋ሜ 𝑦𝑖 − 𝑌ሜ
σ𝑛𝑖=1 ∗
𝑠𝑋 𝑠𝑌
𝑟𝑋,𝑌 =
𝑛−1
8
Example
© Faculty of Management
Correlation
10
Correlation ( or r)
Properties
i. -1 1, unitless, scale-free.
ii. = 0: No linear relationship. Same as Cov = 0
iii. As gets closer to 1 or (-1), the linear relationship gets stronger
• = +1: Perfectly linear positive relationship.
• = -1: Perfectly linear negative relationship.
iv. Slope of the line is not related to correlation coefficient.
(or r) = 0 (or r) = 1
120 800
700
100
600
80
500
60 400
300
40
200
20 100
0 0
0 20 40 60 80 100 120 0 20 40 60 80 100 120
11
Correlation ( or r)
Make a guess for the following cases
(i) = -1.0, (ii) = -0.7 (iii) = 0,
(iv) = +0.7, (v) = 1.0
250 A 4.5 B 800 C
4 700
200 3.5
600
3
150 500
2.5
400
2
100 300
1.5
200
50 1
0.5 100
0 0 0
0 20 40 60 80 100 120 0 20 40 60 80 100 120 0 20 40 60 80 100 120
0
0
D 20 40 60 80 100 120
120
E 60 F
-50 100 40
-100 80 20
-150 60 0
0 20 40 60 80 100 120
-200 40 -20
-250 -40
20
-300 -60
0
0 20 40 60 80 100 120 -80
-350
12
Correlation ( or r)
Make a guess for the following cases
(i) = -1.0, D (ii) = -0.7 F (iii) = 0, E
(iv) = +0.7, A (v) = 1.0 B C
250 A 4.5 B 800 C
4 700
200 3.5
600
3
150 500
2.5
400
2
100 300
1.5
200
50 1
0.5 100
0 0 0
0 20 40 60 80 100 120 0 20 40 60 80 100 120 0 20 40 60 80 100 120
0
0
D 20 40 60 80 100 120
120
E 60 F
-50 100 40
-100 80 20
-150 60 0
0 20 40 60 80 100 120
-200 40 -20
-250 -40
20
-300 -60
0
0 20 40 60 80 100 120 -80
-350
13
Linear Regression
© Faculty of Management
Linear regression
Basic model
Estimating coefficients
Using Excel
Interpreting results
Coefficient estimates: Confidence intervals and hypothesis
tests
Evaluating the overall model: ANOVA and R2
Standard error
Prediction and confidence Intervals
Validating the model
15
Linear regression
16
Regression: Location analysis for Starbuck’s
17
Using regression analysis
Hypothesized
Business Reality relationships
Need for location analysis Profit = 0 + 1XPop + 2XCollege…
18
Simple linear regression
Basic model
Estimating coefficients
Using Excel
Interpreting results
Coefficient estimates: Confidence intervals and hypothesis
tests
Evaluating the overall model: ANOVA and R2
Standard error
Prediction and confidence Intervals
Validating the model
19
Example: Montreal house price and square footage
20
Example: Montreal house price and square footage
Chart ➔ Scatter plot
Add a trend line (sample regression line)
Right click data points ➔ trend line / options / show equation
Datapoint Sq. Ft Price (in $100's)
Montreal House Price by Sq. Ft.
1 1185 2400
4 1504 2599
7000
5 2050 2890
Selling price ($000's)
6000
6 1757 3195
9 2400 3590
3000
10 2544 3900
2000
11 2201 3949
1000
12 2437 4290
13 2877 4499 0
0 1000 2000 3000 4000 5000 6000
14 2944 5190
Square footage
15 3721 6140
16 3190 6200
17 3992 7350
18 5015 7350
21
This is great! Why learn regression analysis?
Price and Sq. Ft
9000
Sample 1
8000
5000
4000
3000
Y= 370.27+ 1.4792X
2000
1000
0
0 1000 2000 3000 4000 5000 6000
Sq. Ft
22
Simple Linear
Regression
© Faculty of Management
Simple regression model: Population model yi = 0 + 1xi + i
Regression Assumptions:
1. is independent of x and follows N(0, ) for all x.
2. i do not exhibit autocorrelation (are independent of each other):
Cov(i, j)=0
24
Simple regression model
yˆ i estimates Y|x = 0 + 1 xi
ˆi = b0 + b1 xi
y
the average value of y at x i
How do we find b0 and b1?
ei: residual (difference between an actual y and a predicted value ŷ)
residual : e i = y i - ŷ i = yi - (b 0 + b1x i )
25
Simple regression: Least square criterion
Intercept = b0
xi x
26
Estimating Regression
Coefficients
© Faculty of Management
Simple linear regression: Estimating coefficients
min ei = 2
i i =
(y −ŷ ) 2
i 0 1 i
(y − (b + b x )) 2
σ(𝑥 − 𝑥)(𝑦
lj − 𝑦)
lj 𝐶𝑜𝑣(𝑥, 𝑦) 𝑟𝑋,𝑌 𝑠𝑋 𝑠𝑌 𝑠𝑌
𝑏1 = = = = 𝑟𝑋,𝑌
σ(𝑥 − 𝑥)lj 2 𝑠𝑥2 𝑠𝑥2 𝑠𝑋
y ŷ = b 0 + b1x
y
ei Slope = b1
Intercept =
b0
xi x x
28
Using Excel: Montreal house price example
Want to explain the house price (Y) with the sq. ft. (X)
Pop. model: yi=β0+β1xi+εi Estimated model: yˆ i = b0 + b1 xi
Tools / Data Analysis / Regression
Range of Y
(include label)
Range of X
(include label)
output
range
output options
29
Using Excel: Montreal house example Excel output
SUMMARY OUTPUT
Regression Statistics
Multiple R 0.9623 3. Regression summary statistics
R Square 0.9261
Adjusted R Square 0.9214
Standard Error 455.2725
Observations 18 2. Overall significance of a model
ANOVA
df SS MS F Significance F
Regression 1 41530449.03 41530449 200.3659 1.81687E-10
Residual 16 3316369.417 207273.1
Total 17 44846818.44
x x x
The expression “the coefficient of ___ is significant” refers to this default
hypothesis test: You can conclude the coefficient is different from zero (at
the specified α)
31
Evaluating Regression
Model
© Faculty of Management
Evaluating the overall regression model
Pop. regression model yi = 0 +1xi +I
Our fit line: ŷi = b0 +b1xi
How well does our model explain the variation in y (from the mean y )?
xi x
33
Evaluating the overall regression model
n
i
SST (Total Sum of Squares)
variation of the y values around their mean
( y − y ) 2
i =1
=
SSR (Regression Sum of Squares) n
portion of the total sum of squares explained
by the estimated model
i
( ˆ
y
i =1
− y ) 2
+
SSE (Sum of Squared Errors) n
the sum of squares of the errors (between the
model prediction and the actual observed
i i
( ˆ
y
i =1
− y ) 2
data)
35
Evaluating the overall model: ANOVA Table
Regression Statistics
SSR = ( yˆ i − y ) 2
SSE = ( yˆ i − yi ) 2
Multiple R 0.9623
R Square 0.9261
Adjusted R Square 0.9214
Standard Error 455.2725 SST = ( yi − y ) 2
Observations 18
ANOVA
df SS MS F Significance F
Regression 1 41530449.03 41530449.03 200.366 1.81687E-10
Residual 16 3316369.417 207273.0886
Total 17 44846818.44
36
Evaluating the overall model: Coefficient of determination, R2
Regression Statistics 𝑆𝑆𝑅 𝑆𝑆𝑇 − 𝑆𝑆𝐸 original − unexplained
Multiple R 0.9623 𝑅2 = = =
R Square 0.9261
𝑆𝑆𝑇 𝑆𝑆𝑇 original
Adjusted R Square 0.9214
Standard Error 455.2725 The portion of the total variation explained
Observations 18
by the regression model.
ANOVA
df SS MS F Significance F
Regression 1 41530449.03 41530449.03 200.366 1.81687E-10
Residual 16 3316369.417 207273.0886
Total 17 44846818.44
37
Evaluating the overall model: How do I determine a bad model?
Regression Statistics
Multiple R 0.9623 1. If you cannot reject H0 that all independent
R Square 0.9261 variable parameters i are zero, or
Adjusted R Square 0.9214 2. R2 is too small, or
Standard Error 455.2725
Observations 18
3. if sb is so large that CI for i includes zero
ANOVA
df SS MS F Significance F
Regression 1 41530449.03 41530449.03 200.366 1.81687E-10
Residual 16 3316369.417 207273.0886
Total 17 44846818.44
38
Evaluating the overall model: Estimating σ
Regression Statistics Recall regression model yi= 0 +1xi +i
Multiple R 0.9623
R Square 0.9261 where i~N(0, ).
Adjusted R Square 0.9214
Standard Error 455.2725 This is the estimate of σ: the standard deviation of
Observations 18 the errors around the regression line. It is called
the “standard error of the estimate”
ANOVA
df SS MS F Significance F
Regression 1 41530449.03 41530449.03 200.366 1.81687E-10
Residual 16 3316369.417 207273.0886
Total 17 44846818.44
𝑆𝑆𝐸
𝑠𝜀 = 𝑀𝑆𝐸 =
𝑛−𝑝−1
is an unbiased estimator for
𝑆𝑆𝐸 3,316,369.417
𝑠𝜀 = = = 455.2725
𝑛−𝑝−1 16
39
Confidence and
Prediction Intervals
© Faculty of Management
Estimation for mean value (y|x) & prediction for an individual value (yi)
41
100(1-)% CI for the avg. y value at x (y|x)
Confidence Interval for the average y value for given x
(estimating y|xp=0+1x using ŷ=b0 + b1x )
42
MTL house example: CI for the avg. y value at x (y|x)
Construct a 95% CI for the avg. price of the 2600 sq. ft. houses.
Coefficients Standard Error t Stat P-value Regression Statistics
Intercept 325.403 293.0507 1.1104 0.2832 Multiple R 0.9623
R Square 0.9261
Sq. Ft 1.555 0.1099 14.1551 0.0000 Adjusted R Square 0.9214
Standard Error 455.2725
Observations 18
lj 2
1 (𝑥 − 𝑥) 1
𝑦𝑥 ± 𝑡𝛼/2 𝑠𝜀 1+ + 𝑦𝑥 ± 𝑡𝛼/2 𝑠𝜀 1 +
𝑛 (𝑛 − 1)𝑠𝑥2 𝑛
For a given x, the prediction interval is (wider or narrower) than the CI.
44
MTL house example: Prediction interval for a specific y value
Construct a 95% Prediction Interval for the price of a 2600 sq. ft. house.
Coefficients Standard Error t Stat P-value Regression Statistics
Multiple R 0.9623
Intercept 325.403 293.0507 1.1104 0.2832
R Square 0.9261
Sq. Ft 1.555 0.1099 14.1551 0.0000 Adjusted R Square 0.9214
Standard Error 455.2725
Observations 18
45
Confidence and prediction interval: Exact vs. approximated
4368.4 appr. CI
47