0% found this document useful (0 votes)
9 views11 pages

Extensions MLR

The document discusses the application of regression through origin in multiple linear regression (MLR) using the Wage1 dataset, highlighting the differences in results when including or excluding an intercept. It also addresses challenges such as biased OLS estimators and the impact of scaling variables on regression outcomes. Additionally, it explains standardization of variables and its implications for interpreting regression coefficients.

Uploaded by

Smiksha Karnavat
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views11 pages

Extensions MLR

The document discusses the application of regression through origin in multiple linear regression (MLR) using the Wage1 dataset, highlighting the differences in results when including or excluding an intercept. It also addresses challenges such as biased OLS estimators and the impact of scaling variables on regression outcomes. Additionally, it explains standardization of variables and its implications for interpreting regression coefficients.

Uploaded by

Smiksha Karnavat
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Regression through Origin in MLR

OLS estimators when there is no intercept in a MLR


Estimating a multiple regression model
through origin: An example
• The Wage1 dataset is used to estimate wage

True model: 𝑤𝑎𝑔𝑒 = 𝛽0 + 𝛽1 𝑒𝑑𝑢𝑐𝑎𝑡𝑖𝑜𝑛 + 𝛽2 𝑒𝑥𝑝𝑒𝑟𝑖𝑒𝑛𝑐𝑒 + 𝜇 −−−− −(1)

Regression through origin: 𝑤𝑎𝑔𝑒=


෧ 𝛽 ෪1 𝑒𝑑𝑢 + 𝛽
෪2 𝑒𝑥𝑝 −−−−−− −(2)
STATA Result 1: Regression Results STATA Result 2: Regression Results
. reg wage education experience
. reg wage education experience, noconstant
Source SS df MS Number of obs = 526
F(2, 523) = 75.99 Source SS df MS Number of obs = 526
Model 1612.2545 2 806.127251 Prob > F = 0.0000 F(2, 524) = 896.32
Residual 5548.15979 523 10.6083361 R-squared = 0.2252 Model 19690.6003 2 9845.30014 Prob > F = 0.0000
Adj R-squared = 0.2222 Residual 5755.69207 524 10.9841452 R-squared = 0.7738
Total 7160.41429 525 13.6388844 Root MSE = 3.257 Adj R-squared = 0.7729
Total 25446.2924 526 48.3769817 Root MSE = 3.3142

wage Coefficient Std. err. t P>|t| [95% conf. interval]


wage Coefficient Std. err. t P>|t| [95% conf. interval]
education .6442721 .0538061 11.97 0.000 .5385695 .7499747
experience .0700954 .0109776 6.39 0.000 .0485297 .0916611 education .4170463 .0162767 25.62 0.000 .3850707 .449022
_cons -3.390539 .7665661 -4.42 0.000 -4.896466 -1.884613 experience .0454382 .0096228 4.72 0.000 .0265341 .0643422

Source: Author’s estimation using Wage1 dataset in STATA, refer Do file for commands
Challenges of regression through origin

▪ Numerical properties of OLS do not apply to regression through origin

▪ Problem of negative R-squared value

▪ If the intercept 𝛽0 in the population model of the regression through origin is


not zero, then the OLS estimators of the slope parameters will be biased

▪ The variances of the OLS slope estimators are larger when estimating an
intercept when 𝛽0 is truly zero.
Scaling of variables
Changing the scale of the dependent variable
• The Wage1 dataset is used to estimate wage
ෟ =𝛽
𝑤𝑎𝑔𝑒 ෢0 + 𝛽
෢1 𝑒𝑑𝑢cation + 𝛽
෢2 𝑒𝑥𝑝erience−−(1) wage_cents=(
ෟ β෢0 *100) +(β
෢1 ∗ 100)) education+(β
෢2 *100) experience --(2)
STATA Result 1: Regression Results STATA Result 2: Regression Results
. reg wage education experience . gen wage_cents = wage*100

. reg wage_cents education experience


Source SS df MS Number of obs = 526
F(2, 523) = 75.99 Source SS df MS Number of obs = 526
Model 1612.2545 2 806.127251 Prob > F = 0.0000 F(2, 523) = 75.99
Residual 5548.15979 523 10.6083361 R-squared = 0.2252 Model 16122545.2 2 8061272.59 Prob > F = 0.0000
Adj R-squared = 0.2222 Residual 55481597.8 523 106083.361 R-squared = 0.2252
Total 7160.41429 525 13.6388844 Root MSE = 3.257 Adj R-squared = 0.2222
Total 71604143 525 136388.844 Root MSE = 325.7

wage Coefficient Std. err. t P>|t| [95% conf. interval]


wage_cents Coefficient Std. err. t P>|t| [95% conf. interval]

education .6442721 .0538061 11.97 0.000 .5385695 .7499747 education 64.42721 5.380607 11.97 0.000 53.85695 74.99747
experience .0700954 .0109776 6.39 0.000 .0485297 .0916611 experience 7.00954 1.097764 6.39 0.000 4.852972 9.166107
_cons -3.390539 .7665661 -4.42 0.000 -4.896466 -1.884613 _cons -339.054 76.65661 -4.42 0.000 -489.6466 -188.4613

. predict wagehat_cents
. predict wagehat_dollars
(option xb assumed; fitted values)
(option xb assumed; fitted values)
. sum wagehat_cents
. sum wagehat_dollars
Variable Obs Mean Std. dev. Min Max
Variable Obs Mean Std. dev. Min Max
wagehat_ce~s 526 589.6103 175.2416 -184.8441 1023.912
wagehat_do~s 526 5.896103 1.752416 -1.848441 10.23912

Source: Author’s estimation using Wage1 dataset in STATA, refer Do file for commands

wage_cents=wage*100
Changing the scale of the dependent
variable(contd)
• In our example we converted the wage in dollars to wage in cents

• Mean of 𝑤𝑎𝑔𝑒 ෟ
ෟ =5.89 dollars and 𝑤𝑎𝑔𝑒_𝑐𝑒𝑛𝑡𝑠 =589 cents

• Thus, if the dependent variable is divided or multiplied by some nonzero constant, c,


then the OLS slope coefficients as well as the intercept is divided or multiplied by c
Changing the scale of an independent
variable
• The Wage1 dataset is used to estimate wage
𝑤𝑎𝑔𝑒
ෟ =𝛽 ෢0 + 𝛽
෢1 𝑒𝑑𝑢𝑐𝑎𝑡𝑖𝑜𝑛 + 𝛽
෢2 𝑒𝑥𝑝𝑒𝑟𝑖𝑒𝑛𝑐𝑒 ------(1) ෢1
𝛽
𝑤𝑎𝑔𝑒
ෟ =𝛽 ෢0 +( ) 𝑒𝑑𝑢𝑐𝑎𝑡𝑖𝑜𝑛 _months+ 𝛽
෢2 experience−−−−−(3)
12
STATA Result 1: Regression Results STATA Result 3: Regression Results
. reg wage education experience
. gen education_months= education*12
Source SS df MS Number of obs = 526
F(2, 523) = 75.99 . reg wage education_months experience
Model 1612.2545 2 806.127251 Prob > F = 0.0000
Residual 5548.15979 523 10.6083361 R-squared = 0.2252 Source SS df MS Number of obs = 526
F(2, 523) = 75.99
Adj R-squared = 0.2222
Model 1612.2545 2 806.127251 Prob > F = 0.0000
Total 7160.41429 525 13.6388844 Root MSE = 3.257 Residual 5548.15979 523 10.6083361 R-squared = 0.2252
Adj R-squared = 0.2222
Total 7160.41429 525 13.6388844 Root MSE = 3.257
wage Coefficient Std. err. t P>|t| [95% conf. interval]

education .6442721 .0538061 11.97 0.000 .5385695 .7499747 wage Coefficient Std. err. t P>|t| [95% conf. interval]
experience .0700954 .0109776 6.39 0.000 .0485297 .0916611
education_months .0536893 .0044838 11.97 0.000 .0448808 .0624979
_cons -3.390539 .7665661 -4.42 0.000 -4.896466 -1.884613 experience .0700954 .0109776 6.39 0.000 .0485297 .0916611
_cons -3.390539 .7665661 -4.42 0.000 -4.896466 -1.884613

. predict wagehat_edu_year
(option xb assumed; fitted values) . predict wagehat_edu_months
(option xb assumed; fitted values)
. sum wagehat_edu_year
. sum wagehat_edu_months

Variable Obs Mean Std. dev. Min Max Variable Obs Mean Std. dev. Min Max

wagehat_ed~r 526 5.896103 1.752416 -1.848441 10.23912 wagehat_ed~s 526 5.896103 1.752416 -1.848441 10.23912

Source: Author’s estimation using Wage1 dataset in STATA, refer Do file for commands
Changing the scale of an independent
variable(contd)
• In our example we converted education in years to education in months

• Thus ,if the independent variable is divided or multiplied by some nonzero


constant, c, then the OLS slope coefficient is multiplied or divided by c,
respectively.

• Changing the units of measurement of only the independent variable does not
affect the intercept

• If we transform the variables to logs, then the change in the scale does not make
any change to the betas of the transformed variable
Standardized Coefficients
(beta coefficients)
Standardizing a variable

▪ To standardize a variable:

❑ Subtract the mean value of the variable

❑ Divide by the standard deviation of the variable

❑ For Example: std_ 𝑤𝑎𝑔𝑒 = (𝑤𝑎𝑔𝑒 - 𝑤𝑎𝑔𝑒) / standard deviation(𝑤𝑎𝑔𝑒)


Standardizing a variable(contd)
▪ The mean value of a standardized variable is always zero and its standard deviation value is
always 1

▪ Standardized variables/Unit-free numbers because the numerator and denominator are


measured in the same unit of measurement and while standardizing the units cancel

▪ Since the mean of the standardized units is always zero , the intercept automatically takes
the value zero

▪ The slope coefficient of a standardized regression is interpreted as : If the standardized


regressor increases by one standard deviation unit, on average, the standardized
dependent variable increases by β𝑗 units, keeping other standardized variable constant.

▪ A regression on standardized variables puts the explanatory variables on equal


footing

You might also like