Simple Linear Regression Overview
Simple Linear Regression Overview
• Y >
dependent variable ( explained variable , LHS variable )
• ✗ >
independent variable ( explanatory variable .
Control variable, RHS variable )
.
Example :
Y; =
GO By ✗ 1- i t Ui > (>
testing out whether the independent variable affects the dependent variable
↳ student ✗ Y
test score
•
Simple linear
regression model only cares about 1 X 's affect on Y
.
•
Bo :
intercept
•
B ,
: the slope > main interest ( the value is unknown ) > estimate it with sample
.
other
way to write simple linear
regression model
Y; =
do
-
+ 2, Xi + Ui or Yi =
So
-
+
8, ✗ i + Ui
- -
•
Symbols are just representing the model 's parameters
-
The term U
↳ in data collection ,
they're either unmeasurable variables they're not included in the
regression specification ( for )
I
some or some reason
controlled
they're included
> not mean ✗
not in
Example :
Yield i =
Both fertilizer ; t Ui
( measured in
>
nectar
↳ measured in
kilogram
• U could be > weather >
quality of the land
Pests / disease
I
(
t
"h^d°9T → should be included in the model cause it could affect ✗ s Y → biased !
, farm ,
,, , ay, , , ,y
>
irrigation system
-
• how to recover
Lpo Bn ), from the data ? fulfill the key assumptions
• how can we do it ?
Ordinary Least Squares 101s ) estimator
•
Deriving the OLS estimates
Key assumption
-
f- [ ulx ] :O
✗
( >
key assumption involves an unobservable variable > it cannot be tested statistically
under the assumption :
+ E[ ulx ]
from
E- [ Both ✗ tulx ] :-[ [ Bolx ] EEPIXH ]
+
f- [ YIX ] Bo +
Pix >
f- [ Bok ] t Elp , ✗ I ✗ ] to
=
↳ B, E [ ✗ I ✗ ] =p , ✗
Bo 1- Bix =
Simple Linear Regression
.
Example
wage equation
Correlated with Education > Zero mean conditional assumption FAILS
Wagei =
Bo +
B Education i
,
+ UM /
ay, ,,yy
y \ uncorrelated with Education > Zero mean conditional assumption SUCCEED
• Zero mean conditional
assumptions implication
the Zero mean conditional assumption requires that the average level of ability is the same years of schooling
•
regardless of no correlation
>
•
if we think that average ability is
higher for individual with higher years of schooling ,
the assumption is false
• on the average , ability s education are correlated , hence the assumption is FALSE
• ELYIX ] is a linear function of ✗ ; for any given value of X, the distribution of Y is centered about ECYIX ]
fist is
> ✗
✗,
.
linear regression model
Y =
Bo + Bix + U
with
}
Bo :
intercept population
estimated using
also known slopes
B' :
coefficients of their ✗ > as
•
fitted values of Y are denoted as Y^ Yi: =
pi pixi
+ > also known as predicted value of Y
-
How to estimate the model : data
•
needs data
}
n ,
( ✗ %)
2, > second observation > value of the dependent variable
✓ Of the i th observation -
(✗ n
, Yn ) > n - th observation
↳ when we say
that we estimate the model > pi pi is what we are estimating Elylx ) Both ✗
-
-
(
,
Yy - - - -
- - - - - - - -
-9.3in
- - -
Y : the true
Ñ Bi
value
> we choose .
to minimize the sum of squared residuals ;
>
Ui
ji : :-. : : : : : .
-
§ - - -
% }Ñ :
-
his
;&
I
1
minimize
I :↳j :
i
?
,
¥5 ÷ }v
,
y, - - - - - - - - -
: i
predicted / fitted valve
xj :*
→
✗
✗i ✗
z ,
?⃝
Simple Linear Regression
.
OLS estimators
•
if the optimization problem is solved , the OLS estimators are obtained as below
→ covlxiy)
&? ( ✗ i 5) ( Yi F)
-
Bi
-
Bi I pig
"
=
and = -
8 ! Ix ,
;
-
x-p -> varlx )
depends B
on
? Cov ( Ky )
covariance
=
+1
l×)-
-
- var +
the > always
•
X and Y negatively correlated > the slope is
negative
.
Example
CEO salary s return on equity ( ROE )
salary =
Bo +
Be ROE + U
dependent value /
Predicted value
y -
> fitted
residual
O O O - >
independent
Salary =
963.191 t 10.501 ✗ ROE
( > if ROE = 0 ,
the predicted annual salary is 963.191 thousands of dollars ( $963,191 =
963.191 ✗
)
$1,000
Interpreting coefficients
• assumes that the Zero conditional mean assumption is satisfied
salary =
963.191 t 18.501 ✗ ROE
• Goodness of Fit
.
total sum of square ISST )
SSE > variation of the variable that can be explained by the data
(
?
variation of the
g- µ
> SST data
{ ( Yi by
( Y) variation of the variable that cannot be explained the
dependent variable Ssr >
SST
-
=
i ,
SSE ÷ É (Yi -
F)
i = ,
-
residual sum of Square ( SSR)
n
n '
Yi )
'
SSR :{ ( Yi - =
{ lñ ;)
i = 1 i =p
. SST =
SSE + SSR
?⃝
Simple Linear Regression
Square ( R )
"
model the
" '
-
'
SSE SSR
R =
sg ,
=
y -
SST
↳ o -
pi -
I
• R' = I >
perfect pi near zero >
regressor ✗ is not good at
predicting Y
'
• R does not depend on the units of measurement > measurement free
example
•
CEO salary S ROE
satay =
963.191 + 18.501 ROE
> based on book
h 209 R2 =
g. Oyzzf of the variation of the data can be explained by the variation of the data
=
- > i. 32%
•
Wages education
526 R2 0.163
n
16.31 Of the variations explained by the model
=
be
=
can
- >
.
• Functional Forms
-
Incorporating nonlinearities
•
•
example
% wage = 1100 ✗
Be ) educ
n= 256 R2 =
0.186
model
-
Log log -
up
+
=
,
: constant elasticity > 11 . increase in ✗ leads to a
pit increase in y ( percentage change )
-
Linear -
log model
Y =
pot B , 1091×1 tu
Bo expected
: value of y given ✗ = 1 because
log 111=0
Y =
Bot Bix + U L , E [ ui I ✗i ] :O
ow random sampling > to get the representative sample of population •↳o homoskedasticity
{
"
( ✗i
explanatory
- I )
'
> o
variables are not all the same
heteroskedasticity
to be a constant
i = I
Simple Linear Regression
-
OLS estimators are unbiased
E- IBI ) Bo =
and Elpi ) =p ,
unbiased Field
• On average ,
they will be equal to the values that characterize the true relationship between y and ✗ in the population
\ > multiple random sampling
\ > we can't do multiple random sampling
an unbiased ness is a
feature of the sampling distribution of Bi and Bo >
says nothing of the estimate that are obtained from a given sample
unbiased ness
an
generally fails if any of our four assumptions fail
Varluilxi ) Ñ ✗i
for
=
,
the width of the distribution is similar all ✗
t >
constant for all value of ✗ ( > does not dependent on the value of the explanatory variable
fly ) ay
'
= T
)
>
varluilxi
,
F
EIYIX ] Bot Pix
=
> ✗
✗,
Homoskedasticity
•. predictive power of regression line
E ( Yilxi ) Bo = +
Bi Xi differs across ✗i
↳ the error term exhibits homoskedasti city or the error term is homoskedastic
-
Variances of the OLS estimators
'
i r in '
&? ✗i
( pi )
-
pi ) ,
=
var =
=
and var ( =
g. = ,
lxi -
IT sstx
&! (✗i - I )
'
standard deviation of Pi is sd ( pi ) =
FAE THE =
valve)
[Link] I \ > Smaller standard error > more precise the regression coeffients , larger variance
,
ggy,
µµgg high Se ( pi ) lower precision
>