0% found this document useful (0 votes)
3 views45 pages

Understanding Multiple Regression Dynamics

The document discusses the complexities of multiple regression analysis, emphasizing the importance of understanding the interaction between variables and the significance of R-squared and adjusted R-squared values. It highlights the challenges of including additional variables in a model and the potential impact of outliers on regression results. Key lessons include the need to evaluate the contribution of each variable and the implications of influential points in econometric modeling.

Uploaded by

parthbatra2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views45 pages

Understanding Multiple Regression Dynamics

The document discusses the complexities of multiple regression analysis, emphasizing the importance of understanding the interaction between variables and the significance of R-squared and adjusted R-squared values. It highlights the challenges of including additional variables in a model and the potential impact of outliers on regression results. Key lessons include the need to evaluate the contribution of each variable and the implications of influential points in econometric modeling.

Uploaded by

parthbatra2006
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

…Interesting things to know

“If you torture the data long enough,


Nature will confess”

“There are two things you are better


off not watching in the making:
sausages and econometric
— Edward E. Leamer, 1983.

estimates.”
…Interesting things to know

“If you torture the data long enough,


Nature will confess”

“There are two things you are better


off not watching in the making:
sausages and econometric
— Edward E. Leamer, 1983.

estimates.”
Challenge No.1 for the
day…
Let’s get deeper into Multiple
Regression and try to understand its
dynamics!
Assume that you are
working with Simple
Wish … I Regression Model…
could
understand Y    X 1  
mystery
behind
Multiple
Regression!
• What would happen to our
results if we introduce another
variable X2 into our model?

• If we run a model only with X1


and only with X2 separately,
can we say that total explained
variation will be the sum of
their R2?
Let’s do an experiment…
We run the following model!

Y    X 1  

We also run the following


model!

Y    X 2  
Then, finally we run the
following Multiple
Regression model!

Y   1X 1   2 X 2  
SUMMARY OUTPUT

Regression Statistics Look at the


Multiple R 0.6671
R Square 0.4450 Values of
Adjusted R Square
Standard Error
0.4198
4.1698
R2 !
Observations 24 Y    X 1  
ANOVA
df SS MS F Significance F
Regression 1 306.7323 306.7323 17.6408 0.0004
Residual 22 382.5277 17.3876
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 39.3477 3.7067 10.6154 0.0000 31.6605 47.0348
X1 2.8278 0.6733 4.2001 0.0004 1.4315 4.2241

SUMMARY OUTPUT

Regression Statistics
Multiple R 0.8586
R Square
Adjusted R Square
0.7371
0.7252 Y    X 2  
Standard Error 2.8698
Observations 24

ANOVA
df SS MS F Significance F
Regression 1 508.0688 508.0688 61.6891 0.0000
Residual 22 181.1912 8.2360
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 44.0478 1.4540 30.2944 0.0000 41.0324 47.0632
X2 0.4188 0.0533 7.8542 0.0000 0.3082 0.5294
SUMMARY OUTPUT
Regression Statistics
Look at the
Multiple R 0.9098 Values of
R Square 0.8277 R2 !
Adjusted R Square 0.8113
Standard Error 2.3778
Observations 24
ANOVA
df SS MS F Significance F
Regression 2 570.5268 285.2634 50.4537 0.0000
Residual 21 118.7332 5.6540
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 38.2508 2.1198 18.0448 0.0000 33.8425 42.6591
X1 1.4430 0.4342 3.3237 0.0032 0.5401 2.3459
X2 0.3412 0.0500 6.8306 0.0000 0.2374 0.4451

Y   1X 1   2 X 2   X1 Coefficient = 1.4430 ;


X2 Coefficient = 0.3412
Wish … I
could
understand
mystery
behind
Multiple
Regression!

The Mystery lies in the fact that X1 may be


interacting with X2 and …remember that X2
explains that part of Y which is not explained
by X1.
Let’s understand Multiple
Regression through Venn
Diagrams …

Demystifying the
Y
Mystery!

X1
This Venn Diagram suggests the
following regression model …
Y    X 1  
Y
SUMMARY OUTPUT

X1
Regression Statistics
Multiple R 0.6671 What is the motivation
R Square 0.4450 to introduce another
Adjusted R Square 0.4198
Standard Error 4.1698 variable in the model?
Observations 24

ANOVA
df SS MS F Significance F
Regression 1 306.7323 306.7323 17.6408 0.0004
Residual 22 382.5277 17.3876
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 39.3477 3.7067 10.6154 0.0000 31.6605 47.0348
X1 2.8278 0.6733 4.2001 0.0004 1.4315 4.2241
What X2 is supposed to do in
the regression model?

Demystifying the
Y
Mystery!

X1 X2
X2 tries to explain that part of Y which is not explained by X1 and
through that which is NOT there already in X1.
It means that X2 will explain the residuals of Y which are
left after regressing with X1.

Y Will the whole X2 explain


the residuals of Y ? NO.

X1 X2
Only that part of X2 which

is not contained in X1!!!


Challenge No.2 for the
day…
I have been
told that
higher the R2,
better it is.
Hence, to get the higher value of R2 I should
add more and more variables into my
model and get a better model!!!!!!!!
I do not
think so,
do you?
R2 vs Adjusted R2…

• R2 is defined as a ratio of Explained


Variation to Total Variation and is given as

R2 = ESS/TSS = 1 – RSS/TSS

• Adjusted R2 is defined as -

2 RSS /(n - k - 1) RSS (n - 1)


Adjusted R =1 - =1 -
TSS /(n - 1) TSS (n - k - 1)
I am getting
negative Adjust R2.
But, I don’t
understand how can
a squared quantity
be a negative
number?
2 RSS /(n - k - 1) RSS (n - 1)
Adjusted R =1 - =1 -
TSS /(n - 1) TSS (n - k - 1)
ANOTHER CHALLENGE OF
THE DAY…
This boy is
struggling with a
problem!!!!!
INFLUENTIAL POINTS!!!!

They are Outliers!!!!


Coefficientsa

Standardi
zed
Unstandardized Coefficien
Coefficients ts
Model B Std. Error Beta t Sig.
1 (Constant) 9.458 13.184 .717 .477
WEIGHT OF THE CAR 44.368 13.192 .501 3.363 .002
ENGINE SIZE 19.400 6.689 .483 2.900 .006
WHEEL BASE -9.395 17.093 -.091 -.550 .585
a. Dependent Variable: KMs PER LITTER

Histogram
Dependent Variable: KMs PER LITTER
16

14

12
What do
What do think 10 say
what such 8 about it?
points are 6
Frequency

called? 4
Std. Dev = .97
2 Mean = 0.00
0 N = 50.00
-1.50 -.50 .50 1.50 2.50 3.50 4.50 5.50
-1.00 0.00 1.00 2.00 3.00 4.00 5.00

Regression Standardized Residual


Influential Points…
• These are individual observations that exert
undue influence on the coefficients of a
Regression model.

• An observation that is substantially different


from all other observations can make a large
difference in the results of Regression.
If observations substantially change
our results, do not we like to know
about them and investigate them
further?
Outliers…

• The Influential Points may be called OUTLIERS.

• An outlier may indicate a data entry error or


other problem.
The boy removed
the outlier!!!!!
Coefficientsa

Standardi
zed
Unstandardized Coefficien
Coefficients ts
Model B Std. Error Beta t Sig.
1 (Constant) 26.128 7.597 3.439 .001
WEIGHT OF THE CAR 64.213 7.677 .808 8.365 .000
ENGINE SIZE 17.857 3.764 .492 4.744 .000
WHEEL BASE -33.367 9.903 -.353 -3.369 .002
a. Dependent Variable: KMs PER LITTER

Histogram

8
Dependent Variable: KMs PER LITTER Better
6
Results?
4
Frequency

2
Std. Dev = .97
Mean = 0.00
0 N = 49.00

Regression Standardized Residual


First thing for the DAY!!!!!!…
What you will say about such points?
Something Technical…

What these
points are
called?
Next Challenge …

??????????
??????????
nifty_ret
???????
LET’S
UNDERSTAND
CERTAIN
IMPORTANT AND
INTERESTING
POINTS ABOUT
ECONOMETRICS
MODELS!!!
Lesson#1
IMPORTANT THINGS TO REMEMBER
R2 and Adjusted R2 …
• Adjusted R2 is more important when we are
dealing with Multiple Regression.
• Adjusted R2 penalizes R2 for loosing degrees of
freedom when additional independent variables
are added into the model.
• Adjusted R2 may be NEGATIVE!
IMPORTANT THINGS TO REMEMBER What is the impact of an independent
variable introduced in the Model on
Adjusted R2 …
• Assume that t-Statistic is less than 1 of an
independent variable introduced in a model, then
dropping it from the model, will increase Adjusted R2.
• If it has t-Statistic more than 1, then Adjusted R2 will
decrease if the corresponding independent variable
is dropped.
Lesson#2
IMPORTANT THINGS TO REMEMBER
Let’s try to understand how
much R2 increases by adding
one MORE VARIABLE into the
model!
• Let’s understand it by looking at the Regression
Output for one variable case and two variable
case!
IMPORTANT THINGS TO REMEMBER
Regression Table – One Variable
SUMMARY OUTPUT
• Please note that Total Sum of Square is
Regression Statistics 689.26 which is totally unexplained
Multiple R 0.6671 when there is no variable in the model.
• Out of it, X1 explains 306.7323 which is
R Square 0.4450 about 44.50%.
Adjusted R
Square 0.4198

Standard Error 4.1698


Observations 24

ANOVA
df SS MS F Significance F
Regression 1 306.7323 306.7323 17.6408 0.0004
Residual 22 382.5277 17.3876

Total 23 689.2600

Coefficient Standard
s Error t Stat P-value Lower 95% Upper 95%
Intercept 39.3477 3.7067 10.6154 0.0000 31.6605 47.0348
X1 2.8278 0.6733 4.2001 0.0004 1.4315 4.2241
IMPORTANT THINGS TO REMEMBER Regression Table – Two Variables
SUMMARY OUTPUT

Regression Statistics • Please note that Total Sum of


Multiple R 0.9098 Square is 689.26 is still same!
R Square 0.8277 • Out of it, X2 explains
Adjusted R Square 0.8113 (570.5268-
Standard Error 2.3778 306.7323=263.7944) that is
Observations 24 about 38.27%.

ANOVA
Significance
df SS MS F F
Regression 2 570.5268 285.2634 50.4537 0.0000
Residual 21 118.7332 5.6540
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 38.2508 2.1198 18.0448 0.0000 33.8425 42.6591
X1 1.4430 0.4342 3.3237 0.0032 0.5401 2.3459
X2 0.3412 0.0500 6.8306 0.0000 0.2374 0.4451
Don’t forget the
relation between F-
Statistic and t-
Statistic.
It will also help in understanding how do
we get the t-Statistic of X2 Variable?
IMPORTANT THINGS TO REMEMBER
F-Statistic is the Square of t-
Statistic!!!

• Note that in Regression, F is the ratio of Mean


Sum of Square due to Regression to Mean Sum of
Square due to residuals.
IMPORTANT THINGS TO REMEMBER Revisiting the Problem…
SUMMARY OUTPUT

Regression Statistics • Please note that X2 explains


Multiple R 0.9098 (570.5268-
R Square 0.8277 306.7323=263.7944) which is
Adjusted R Square 0.8113 also its MS.
Standard Error 2.3778
• If we divide 263.7994 by 5.6540,
Observations 24 we get 46.65656 which is nothing
but Square of t-Statistic of X2.
ANOVA
Significance
df SS MS F F
Regression 2 570.5268 285.2634 50.4537 0.0000
Residual 21 118.7332 5.6540
Total 23 689.2600

Coefficients Standard Error t Stat P-value Lower 95% Upper 95%


Intercept 38.2508 2.1198 18.0448 0.0000 33.8425 42.6591
X1 1.4430 0.4342 3.3237 0.0032 0.5401 2.3459
X2 0.3412 0.0500 6.8306 0.0000 0.2374 0.4451
After learning all
these finer points,
let’s move to another
Challenge of the Day…
W h a t

You might also like