0% found this document useful (0 votes)
8 views5 pages

Simple Linear Regression Overview

Simple linear regression models how a dependent variable (Y) varies with changes in an independent variable (X). It studies the population regression function (PRE) where Y=Bo + B1X + U. The model has parameters including the intercept (Bo), the slope (B1), and the error term (U). The slope measures how Y changes as X changes and is the main parameter of interest. Simple linear regression only considers how one independent variable affects the dependent variable.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views5 pages

Simple Linear Regression Overview

Simple linear regression models how a dependent variable (Y) varies with changes in an independent variable (X). It studies the population regression function (PRE) where Y=Bo + B1X + U. The model has parameters including the intercept (Bo), the slope (B1), and the error term (U). The slope measures how Y changes as X changes and is the main parameter of interest. Simple linear regression only considers how one independent variable affects the dependent variable.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Simple Linear Regression

• A Simple Linear Regression Model


Population Regression Function ( PRE)
×
1; =
Bo +
B, Xi + Ui

the function studies on how Y varies with


changes in ✗

let i= 1,2 , N indicates the


entity
• . . .
. >
Could be students workers households
, , , firms etc
,

• Y >
dependent variable ( explained variable , LHS variable )

• ✗ >
independent variable ( explanatory variable .
Control variable, RHS variable )
.

Example :

> parental income


y student 's test if parental differs
what will happen to score income

Y; =
GO By ✗ 1- i t Ui > (>
testing out whether the independent variable affects the dependent variable

↳ student ✗ Y
test score


Simple linear
regression model only cares about 1 X 's affect on Y
.

parameters of the model


Bo :
intercept

B ,
: the slope > main interest ( the value is unknown ) > estimate it with sample
.
other
way to write simple linear
regression model

Y; =
do
-
+ 2, Xi + Ui or Yi =
So
-
+
8, ✗ i + Ui
- -


Symbols are just representing the model 's parameters

• the most used symbols are


B

-
The term U

U is defined as error term or disturbance

u represents unobservable variables that controlled Y


not in this but
they also affects
• are model ,

↳ in data collection ,
they're either unmeasurable variables they're not included in the
regression specification ( for )
I
some or some reason

controlled
they're included
> not mean ✗
not in

Example :

model relating farming yield s fertilizer use

Yield i =
Both fertilizer ; t Ui

( measured in
>
nectar
↳ measured in
kilogram
• U could be > weather >
quality of the land

Pests / disease
I
(
t
"h^d°9T → should be included in the model cause it could affect ✗ s Y → biased !

, farm ,
,, , ay, , , ,y
>
irrigation system
-

Estimating the parameters


The goal of a simple linear
regression is to estimate the parameters (Bo , p, ) which are unknown

• how to recover
Lpo Bn ), from the data ? fulfill the key assumptions

• how can we do it ?
Ordinary Least Squares 101s ) estimator


Deriving the OLS estimates

Key assumption
-

involves the conditional expectation of a given ×


,
Elul × )

Zero conditional mean assumption

f- [ ulx ] :O

implies that u and ✗ are uncorrelated > exp : productivity i =


Bo +
p, sleep i + ui

( >
key assumption involves an unobservable variable > it cannot be tested statistically
under the assumption :
+ E[ ulx ]
from
E- [ Both ✗ tulx ] :-[ [ Bolx ] EEPIXH ]
+

f- [ YIX ] Bo +
Pix >

f- [ Bok ] t Elp , ✗ I ✗ ] to
=

↳ B, E [ ✗ I ✗ ] =p , ✗
Bo 1- Bix =
Simple Linear Regression
.

Example
wage equation
Correlated with Education > Zero mean conditional assumption FAILS
Wagei =
Bo +
B Education i
,
+ UM /
ay, ,,yy
y \ uncorrelated with Education > Zero mean conditional assumption SUCCEED
• Zero mean conditional
assumptions implication

f- [ ul Education ] = 0 > U and Education are uncorrelated

↳ assumed that u is the same as innate ability in the wage equation

the Zero mean conditional assumption requires that the average level of ability is the same years of schooling

regardless of no correlation
>


if we think that average ability is
higher for individual with higher years of schooling ,
the assumption is false

• on the average , ability s education are correlated , hence the assumption is FALSE

↳ to verify whether the assumption is TRUE or FALSE is


by identifying the variables in Ui that is important in the determination of the dependent variable ( Y )

don't think both are not correlated > assumption is FALSE



if you

if you think both are not correlated >
assumption is TRUE

Assumption in linear regression


model

• ELYIX ] is a linear function of ✗ ; for any given value of X, the distribution of Y is centered about ECYIX ]

fist is

EIYIX ] potpix = interpretation : the distribution of Y is expected to

be centered around the predicted value of Y

> ✗
✗,

.
linear regression model

Basic model ( PRE )

Y =
Bo + Bix + U

with
}
Bo :
intercept population
estimated using
also known slopes
B' :
coefficients of their ✗ > as

• estimated Ps are denoted as


pi pi,
> estimated using sample


fitted values of Y are denoted as Y^ Yi: =
pi pixi
+ > also known as predicted value of Y

• residuals are denoted as u^ uii


: = Yi Yi -

-
How to estimate the model : data

needs data

↳ population (costly) >


sample (randomly selected from the population ) >
representative of the population
> sample are not randomly selected > BIASED ! > estimate will be biased as well

• a random sample of n observations > each entity will have a pair of (✗ , y)

(✗ Yn ) > first observation

}
n ,

( ✗ %)
2, > second observation > value of the dependent variable

✓ Of the i th observation -

( × , ,y , ) > third observation {( ✗ i. 5;) :i=i .


. . .
.
n
}
[ value
i. >
of the explanatory variable
Of i-th Observation

(✗ n
, Yn ) > n - th observation

. How to estimate the model :


OLS

• to estimate Bo Pi , we use the approach of OLS Y

↳ when we say
that we estimate the model > pi pi is what we are estimating Elylx ) Both ✗
-
-

(
,

Yy - - - -
- - - - - - - -

-9.3in
- - -

Y : the true
Ñ Bi
value
> we choose .
to minimize the sum of squared residuals ;
>
Ui
ji : :-. : : : : : .
-

§ - - -
% }Ñ :
-

his
;&
I
1

minimize
I :↳j :
i

> calculus / optimization approach


pi pi
,
,

?
,

¥5 ÷ }v
,
y, - - - - - - - - -

: i
predicted / fitted valve
xj :*


✗i ✗
z ,
?⃝
Simple Linear Regression
.
OLS estimators

if the optimization problem is solved , the OLS estimators are obtained as below

→ covlxiy)
&? ( ✗ i 5) ( Yi F)
-

Bi
-

Bi I pig
"
=
and = -

8 ! Ix ,
;
-
x-p -> varlx )

I and I are the corresponding sample averages


• the slope estimator is basically the ratio of Cov IX. y ) and the var 1×1

depends B
on
? Cov ( Ky )
covariance
=

+1
l×)-
-

- var +
the > always

↳ the relationship between ✗ and Y will determine the slope of Bi


• X and Y positively correlated > the slope is positive


X and Y negatively correlated > the slope is
negative

.
Example
CEO salary s return on equity ( ROE )

• Y : annual salary ( in thousands of dollars ) s ✗ :


average ROE ( in percentage )

salary =
Bo +
Be ROE + U

dependent value /
Predicted value

y -
> fitted
residual
O O O - >

independent

the OLS estimates >


Bo = 963.191 and pi = 10.501

Salary =
963.191 t 10.501 ✗ ROE

( > if ROE = 0 ,
the predicted annual salary is 963.191 thousands of dollars ( $963,191 =
963.191 ✗
)
$1,000

Interpreting coefficients
• assumes that the Zero conditional mean assumption is satisfied

↳ other factors in U are held fixed >


change in u is Zero ( u = 0 )

✗ has a linear effect on Y

Y =P, ✗ if u = 0 - > ceteris paribvs

↳ key feature of applied economics

• assumes that the Zero conditional mean assumption is satisfied

salary =
963.191 t 18.501 ✗ ROE

↳ if ROE increases by one percentage point the salary , would increase


by $18,501 > 18.501 ✗ $1000

• Goodness of Fit
.
total sum of square ISST )
SSE > variation of the variable that can be explained by the data

(
?
variation of the
g- µ
> SST data
{ ( Yi by
( Y) variation of the variable that cannot be explained the
dependent variable Ssr >

SST
-

=
i ,

explained sum of squares ( SSE )


'

SSE ÷ É (Yi -
F)
i = ,

-
residual sum of Square ( SSR)
n
n '
Yi )
'

SSR :{ ( Yi - =
{ lñ ;)
i = 1 i =p

. SST =
SSE + SSR
?⃝
Simple Linear Regression
Square ( R )
"
model the
" '
-

measure of the goodness of fit of a


regression is R

'
SSE SSR
R =
sg ,
=
y -

SST

• pi measures the predictive power of our model

↳ o -
pi -
I

• R' = I >
perfect pi near zero >
regressor ✗ is not good at
predicting Y

pi near one >


regressor
✗ is good at predicting Y

'
• R does not depend on the units of measurement > measurement free

example

CEO salary S ROE

satay =
963.191 + 18.501 ROE
> based on book

h 209 R2 =
g. Oyzzf of the variation of the data can be explained by the variation of the data
=

- > i. 32%


Wages education

Wa^ge : -0.90 t 0.54 educ

526 R2 0.163
n
16.31 Of the variations explained by the model
=
be
=
can
- >
.

R squared / > mean that the regression has a causal interpretation


A high
-

> a good indicator for prediction or


forecasting

• Functional Forms
-

Incorporating nonlinearities


example

the percentage increase in is the given one more year of education


wage same ,

↳ the model that gives ( approximately) a constant percentage effect is

log / wage ) Potpieduct =


U

1091 ) denotes the natural logarithm


\ > if u = 0

% wage = 1100 ✗
Be ) educ

suppose that log ( wage) is the dependent variable , the relationship is as


followed
1 year of schooling would
→ 100% ✗ - > increase wage by d→ return
Iogtwage) = 0.584 + 0.003 educ + U
of
to another
education
year

n= 256 R2 =
0.186

model
-

Log log -

log ly) Both log A)

up
+
=

,
: constant elasticity > 11 . increase in ✗ leads to a
pit increase in y ( percentage change )

-
Linear -

log model

Y =
pot B , 1091×1 tu

Bn : I 't increase in ✗ leads to Be ✗ 0.01 Units increase in y

Bo expected
: value of y given ✗ = 1 because
log 111=0

• Unbiased ness of OLS


. Standard assumptions for simple linear regression model

•wo linear in parameters oq• zero conditional mean assumption

Y =
Bot Bix + U L , E [ ui I ✗i ] :O

ow random sampling > to get the representative sample of population •↳o homoskedasticity

•wo sample variation in explanatory variable ↳ var luilxi ) =


02 > variance of the unobservable variables of Ui on ✗i is

↳ the values of the

{
"

( ✗i
explanatory

- I )
'

> o
variables are not all the same

( > not constant =


going

heteroskedasticity
to be a constant

i = I
Simple Linear Regression
-
OLS estimators are unbiased

ma theorem (unbiased ness )

E- IBI ) Bo =
and Elpi ) =p ,

unbiased Field

Interpretation of unbiased ness


• the estimated coefficients may be smaller or larger
↳ depending on the sample that is the result of a random draw

• On average ,
they will be equal to the values that characterize the true relationship between y and ✗ in the population
\ > multiple random sampling
\ > we can't do multiple random sampling

an unbiased ness is a
feature of the sampling distribution of Bi and Bo >
says nothing of the estimate that are obtained from a given sample

unbiased ness
an
generally fails if any of our four assumptions fail

ooo Variance of 015 Estimators


-
Variances of 015 Estimators
• estimator that has the smallest variance is called
efficient estimator

↳ simpler under the homoskedasticity assumption

Varluilxi ) Ñ ✗i
for
=
,
the width of the distribution is similar all ✗

t >

constant for all value of ✗ ( > does not dependent on the value of the explanatory variable

fly ) ay
'
= T
)

>
varluilxi
,

F
EIYIX ] Bot Pix
=

> ✗
✗,

Homoskedasticity
•. predictive power of regression line

E ( Yilxi ) Bo = +
Bi Xi differs across ✗i

↳ similar across Xi → homo

different across Xi → hetero

When lulx ) costant → Predictive power of Elylx ) does


•. var not
vary with ✗

↳ the error term exhibits homoskedasti city or the error term is homoskedastic

-
Variances of the OLS estimators
'
i r in '
&? ✗i
( pi )
-

pi ) ,
=

var =
=
and var ( =

g. = ,
lxi -

IT sstx
&! (✗i - I )
'

standard deviation of Pi is sd ( pi ) =
FAE THE =

↳ higher sampling variability of the estimated regression coefficients >


larger the variability of the unobserved factors
I > lower sampling variability of the estimated regression coefficients >
higher the variation in the explanatory variable →
preferable ( more condifidence in the estimated

valve)

variance is unknown → the data is used to estimate 82


↳ the unbiased estimator for J '
is sample variance of the residuals £2 → Elf )=Ñ '

Calculating of SE for regression coefficients


• standard error of the regression > I =
IF
• standard error of B1 var ( pi ) and/or Sel pi ) measure how precisely the
regression coefficients are estimated

[Link] I \ > Smaller standard error > more precise the regression coeffients , larger variance
,

ggy,
µµgg high Se ( pi ) lower precision
>

You might also like