0% found this document useful (0 votes)
24 views82 pages

Simple Linear Regression Overview

The document outlines the principles of Simple Linear Regression, including scatter plots, covariance, and the coefficient of correlation as tools to analyze the relationship between two numerical variables. It discusses how to visualize and measure linear associations, make predictions, and the importance of understanding correlation in the context of investment diversification. Examples are provided to illustrate these concepts using data from students' height and weight, as well as stock returns.

Uploaded by

matthewlee0312
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views82 pages

Simple Linear Regression Overview

The document outlines the principles of Simple Linear Regression, including scatter plots, covariance, and the coefficient of correlation as tools to analyze the relationship between two numerical variables. It discusses how to visualize and measure linear associations, make predictions, and the importance of understanding correlation in the context of investment diversification. Examples are provided to illustrate these concepts using data from students' height and weight, as well as stock returns.

Uploaded by

matthewlee0312
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CB2200 Business Statistics

Topic 8
Simple Linear Regression
Reference
Levine, D.M., Kathryn, A.S. and David, F.S. Business Statistics: A First Course, Pearson
Education Ltd, Chapter 2 & 3 & 12

1
Outline

 Scatter Plot
 Covariance and the Coefficient of Correlation
 Simple Linear Regression
 Least Squares Estimation
 Predictions in Regression Analysis
 Coefficient of Determination
 Inferences about the Slope
 Applications of Linear Regression

2
Association Between Two Numerical
Variables
 To visualize the relationship between two numerical
variables
 Using scatter plot
 To measure the degree of linear association
 Using coefficient of correlation
 To forecast one variable for given values of the other
 Using regression model
 Examples
 Apartment price vs. Gross floor area
 Weekly sales for chain stores vs. Number of customers

3
Association Between Two Numerical
Variables Cont’d

 We will look at two variables measuring different


characteristics of some population of individuals
 Usually consist of paired sample data corresponding to
pairs of observations on the two variables for members
of a sample taken from the population
 If two variables are related, then the nature of the
relationship may be indicated by plotting paired samples
of observations for both variables on a scatter plot

4
Association Between Two Numerical
Variables – Example Cont’d

 Consider the following data for variables from a sample


of 10 students
 𝑋 = Height (in cm.)
 𝑌 = Weight (in kg.)
𝑋 𝑌
170 75
185 80
165 75
140 50
180 70
150 60
200 100
160 65
175 80
170 70 5
Association Between Two Numerical
Variables – Example Cont’d

 There is a clear tendency for small values of to be


associated with small values of , and large with large

 The dots on the scatter plot lie “close to” a straight line
with a positive slope
 We say that these two variables, height and weight, have
a positive linear association

6
Association Between Two Numerical
Variables – Example Cont’d

 Consider the scatter plot between Height ( ) and IQ ( )


for the same 10 students
𝑋 𝒁
170 120
185 130
165 140
140 135
180 100
150 115
200 130
160 125
175 145
170 110

 The diagram indicates no obvious relationship between


and , as you might well expected, since there is no
7
known relationship between height and IQ
Association Between Two Numerical
Variables – Example Cont’d

 If the dots on the scatter plot lie “close to” a straight line
with negative slope, we say that the variables exhibit a
negative linear association

8
Covariance

 How do we measure the degree of linear association


between two variables and ?
 The answer to this question is the covariance
 A quantity that measures the linear association

 Population covariance

 Sample covariance

 An estimator of 𝜎 based on 𝑛 pairs of sample values


9
Covariance
Cont’d

 The cross product term will be


positive in quadrants I and III, and negative in quadrants
II and IV
 With positive
linear association,
there is a 𝑋
Quadrant II Quadrant I
tendency for the
dots to lie
predominantly in 𝑌
quadrants I and III
Quadrant III Quadrant IV

10
Covariance
Cont’d

 On the other hand, with negative linear association,


there is a tendency for the dots to lie predominantly in
quadrants II and IV

11
Covariance
Cont’d

 If there is no or very weak linear association, then there


is a tendency for the dots to scatter across all four
quadrants

12
Covariance
Cont’d

 The covariance only measures linear association


 A covariance of zero does not necessarily imply that
and have no association because they may be related
in a non‐linear way

13
Covariance – Example
Cont’d

 Consider the sample data regarding Height ( ) and


Weight ( )
𝑋 𝑌 𝑿 𝑿 𝒀 𝒀 𝑿 𝑿 𝒀 𝒀
170 75 0.5 2.5 1.25
185 80 15.5 7.5 116.25
165 75 ‐4.5 2.5 ‐11.25
140 50 ‐29.5 ‐22.5 663.75
180 70 10.5 ‐2.5 ‐26.25
150 60 ‐19.5 ‐12.5 243.75
200 100 30.5 27.5 838.75
160 65 ‐9.5 ‐7.5 71.25
175 80 5.5 7.5 41.25
170 70 0.5 ‐2.5 ‐1.25

𝑿 169.5 𝒀 72.5 𝑺𝑿𝒀 215.28


14
Covariance – Example
Cont’d

 Let’s convert the height of the students from cm. to m.


𝑋′ 𝒀 𝑿′ 𝑿′ 𝒀 𝒀 𝑿′ 𝑿′ 𝒀 𝒀
1.7 75 0.005 2.5 0.0125
1.85 80 0.155 7.5 1.1625
1.65 75 ‐0.045 2.5 ‐0.1125
1.4 50 ‐0.295 ‐22.5 6.6375
1.8 70 0.105 ‐2.5 ‐0.2625
1.5 60 ‐0.195 ‐12.5 2.4375
2 100 0.305 27.5 8.3875
1.6 65 ‐0.095 ‐7.5 0.7125
1.75 80 0.055 7.5 0.4125
1.7 70 0.005 ‐2.5 ‐0.0125

𝑿′ 1.695 𝒀 72.5 𝑺𝑿 𝒀 2.1528

 The sample covariance is reduced by a factor of 100 15


Covariance
Cont’d

 One problem with the covariance is that it is dependent


on the units used to measure and
 Its value does not indicate the strength of the linear relationship
of the two variables
 Its value cannot be directly compared for different variables

16
Coefficient of Correlation

 The coefficient of correlation measures the relative


strength of a linear association between two variables
that is not affected by the variables’ units of measure
 It adjusts the covariance by the standard deviations of 𝑋 and 𝑌
so that the resulting measure is unit‐free
 It is a “standardized score” of the covariance

17
Coefficient of Correlation
Cont’d

 Population coefficient of correlation


pronounced rho

∑ ∑

 Sample coefficient of correlation


∑ ∑

 An estimator of 𝜌
 The sign of ( ) is the same as that of ( )
 As the denominator of 𝜌 is always non‐negative
18
Coefficient of Correlation – Example

 Consider the sample data regarding Height ( ) and


Weight ( ) again
𝑋 𝑌 𝟐 𝟐
𝑿 𝑿 𝒀 𝒀 𝑿 𝑿 𝒀 𝒀 𝑿 𝑿 𝒀 𝒀
170 75 0.5 2.5 1.25 0.25 6.25
185 80 15.5 7.5 116.25 240.25 56.25
165 75 ‐4.5 2.5 ‐11.25 20.25 6.25
140 50 ‐29.5 ‐22.5 663.75 870.25 506.25
180 70 10.5 ‐2.5 ‐26.25 110.25 6.25
150 60 ‐19.5 ‐12.5 243.75 380.25 156.25
200 100 30.5 27.5 838.75 930.25 756.25
160 65 ‐9.5 ‐7.5 71.25 90.25 56.25
175 80 5.5 7.5 41.25 30.25 56.25
170 70 0.5 ‐2.5 ‐1.25 0.25 6.25
𝑿 169.5 𝒀 72.5 𝑺𝑿𝒀 215.28 𝑺𝑿 17.232 𝑺𝒀 13.385
 19
Coefficient of Correlation – Example
Cont’d

 What if the height is measured in m.?


𝑋′ 𝑌 𝟐 𝟐
𝑿′ 𝑿′ 𝒀 𝒀 𝑿′ 𝑿′ 𝒀 𝒀 𝑿′ 𝑿′ 𝒀 𝒀
1.7 75 0.005 2.5 0.0125 0.000025 6.25
1.85 80 0.155 7.5 1.1625 0.024025 56.25
1.65 75 ‐0.045 2.5 ‐0.1125 0.002025 6.25
1.4 50 ‐0.295 ‐22.5 6.6375 0.087025 506.25
1.8 70 0.105 ‐2.5 ‐0.2625 0.011025 6.25
1.5 60 ‐0.195 ‐12.5 2.4375 0.038025 156.25
2 100 0.305 27.5 8.3875 0.093025 756.25
1.6 65 ‐0.095 ‐7.5 0.7125 0.009025 56.25
1.75 80 0.055 7.5 0.4125 0.003025 56.25
1.7 70 0.005 ‐2.5 ‐0.0125 0.000025 6.25
𝑿′ 1.695 𝒀 72.5 𝑺𝑿 𝒀 2.1528 𝑺𝑿 0.1723 𝑺𝒀 13.385

 The sample correlation remains unchanged although the sample
20
covariance has been reduced by a factor of 100
Coefficient of Correlation
Cont’d

 It can be shown it is always the case that


and
 Three special values of and are of interest
 When 𝜌 0 (𝑟 0), 𝑋 and 𝑌 are not linearly related, and
we say that 𝑋 and 𝑌 are uncorrelated in the population (sample)
 When all population (sample) values of 𝑋 and 𝑌 lie exactly on a
straight line having a positive slope, then 𝜌 1 (𝑟 1)
 When all population (sample) values of 𝑋 and 𝑌 lie exactly on a
straight line having a negative slope, then 𝜌 1 (𝑟 1)
 If the population (sample) values of and lie close to a
straight line, then ( ) will be close to 1 or ‐1
21
Coefficient of Correlation
Cont’d

 Here are some diagrams illustrating different values of

22
Coefficient of Correlation
Cont’d

 Correlation alone cannot prove that there is a


causation effect
 Causation effect means that the change in the value of
one variable caused the change in the other variable

23
Diversifying Your Investments

 One basic theory of investing is diversification


 The idea is that you want to have a basket of stocks that do not
all “move in the same direction’
 If one investment goes down, you don’t want a second
investment in your portfolio that is also likely to go down
 One hallmark of a good portfolio is a low correlation
between investments

24
Diversifying Your Investments
Cont’d

 The following data represent the annual rates of return


for various stocks
Year Cisco Systems Walt Disney General Electric Exxon Mobil TECO Energy Dell
1999 1.310 ‐0.015 0.574 0.151 ‐0.303 ‐0.319
2000 ‐0.286 ‐0.004 ‐0.055 0.127 0.849 ‐0.661
2001 ‐0.527 ‐0.277 ‐0.151 ‐0.066 ‐0.150 0.553
2002 ‐0.277 ‐0.203 ‐0.377 ‐0.089 ‐0.369 ‐0.031
2003 0.850 0.444 0.308 0.206 0.004 0.254
2004 ‐0.203 0.202 0.207 0.281 0.128 0.234
2005 0.029 ‐0.129 ‐0.014 0.118 0.170 ‐0.288
2006 0.434 0.443 0.093 0.391 0.051 ‐0.164
2007 0.044 ‐0.043 0.126 0.243 0.058 ‐0.033
2008 ‐0.396 ‐0.306 ‐0.593 ‐0.193 ‐0.355 ‐0.580
2009 0.459 0.417 ‐0.102 ‐0.171 0.249 0.393
2010 ‐0.185 0.155 0.053 0.023 0.044 ‐0.323

Source: Yohoo!Finance
25
Calculating Coefficient of Correlation
in Excel
 Find “Data Analysis” in the “Data” menu bar

 Choose “Correlation” at
Data Cells
“Data Analysis” browser

Output Cell
26
Diversifying Your Investments
Cont’d
Cisco Systems Walt Disney General Electric Exxon Mobil TECO Energy Dell

Cisco Systems 1
Walt Disney 0.5512 1
General Electric 0.7461 0.5110 1
Exxon Mobil 0.3625 0.4701 0.7024 1
TECO Energy ‐0.1211 0.3432 0.1477 0.2828 1
Dell 0.0630 0.2906 0.1448 ‐0.0445 ‐0.1768 1

 If you only wish to invest in two stocks


 Which two would you select if your goal is to have low
correlation between the two investments?

 Which two would you select if your goal is to have one stock go
up when the other goes down?
27
Linear Regression Model

 Suppose that a scatter plot or the coefficient of


correlation indicates linear association between two
variables, then it is quite easy
 To fit a straight line to the scatter plot, and
 Use the fitted straight line to forecast values of one variable
(indicated as 𝑌 variable or dependent variable) given values of
the other (indicated as 𝑋 variable or independent variable)
 In other words, given that the variable 𝑋 takes a specific value, we
expect a response in the variable 𝑌
 This can be thought of as a dependency of 𝑌 on 𝑋

28
Linear Regression Model
Cont’d

 Our concern is with the value taken by the variable ,


when the variable takes a specific value
 The variable could take many different values for a
specific value
 For example, we may be interested in the value of retail sales
per household in a year in which disposable income per
household is $12,000. At that income, the retail sales value per
household in the population could be $5,800, or $5,900, or
$6,000, etc. It is not reasonable to think of just a single possible
retail sales level resulting from a particular value for disposable
income

29
Linear Regression Model
Cont’d

 It is more realistic to consider a distribution of possible


values resulting from each possible value
 A crucial characteristic of this distribution is the
population mean, or the expected value, of when
takes a specific value
 For example, we can ask what would be the average (population
mean) retail sales per household in which disposable income per
household was $12,000

30
Linear Regression Model
Cont’d

 In general, we will denote the expected value of the


variable , when the variable takes the specific value
of by

 Our assumption of linearity is the assumption that this


conditional expectation depends linearly on
 This implies that

where the fixed numbers and determine a specific


straight line
 The true values of 𝛽 and 𝛽 are unknown to us
31
Linear Regression Model
Cont’d

 We hypothesize that the conditional expected value of


𝑖 depends linearly on 𝑖
 Such hypothesis will not hold exactly in the real world
 In addition, we do not actually observe the expected value of 𝑌𝑖
for the 𝑋𝑖
 Denote the discrepancy between the observed 𝑖 and its
conditional expected value by such that

32
Linear Regression Model
Cont’d

 The population (or true) regression line is defined as

 The response of 𝑖 to a particular value 𝑖 will be the


sum of two parts
 An expectation 𝛽 𝛽 𝑋 reflecting their systematic
relationship
 A discrepancy 𝜀 from the expectation, often called the error
term
 Since the population regression line involves on one
independent variable ( 𝑖 ), the line is sometimes called
simple linear regression model
33
Least Squares Estimation
13
12
11

 If the random error term = 0 for 10


9
8
7

all , it implies Y

y
6
5
4

exactly
3
2
1
0
0 1 2 3 4 5 6 7

X
x

 If the random error term , = 1, …, 13


12

, are not all equal to 0, then the 11


10
9

observed pairs ( 𝑖 , 𝑖 ), = 1, …, , Y6
8
7

y
cannot be drawn on a straight line 5
4
3

 It is possible to find a straight line that 2


1
0

will fit the set as accurately as possible, 0 1 2 3

Xx
4 5 6 7

i.e. the fitting errors should be


minimized
34
Least Squares Estimation
Cont’d

 Consider two lines A and B, both are fitted to the same


set of ( 𝑖 , 𝑖 ) pairs
13
12
A
11
10
9
8
7
Y
y

6
5
4
3
2
1 B
0
0 1 2 3 4 5 6 7
Xx

 Which line seems to fit the set of points better? Why?

35
Least Squares Estimation
Cont’d

 Consider another two possible lines A and C, both are


fitted to the same set of ( 𝑖 , 𝑖 ) pairs
13
12
A
11
10
9
8
7
y

Y 6
5
4
3 C
2
1
0
0 1 2 3 4 5 6 7

Xx

 Which line seems to fit the set of points better? Why?

36
Least Squares Estimation
Cont’d

 For a given , when a line pass through the (X,Y)


Error = 0
point ( , ) exactly, we say there is no error
 When a line does not pass through the
point, we say there is an error
(X,Y)
 The amount of error is represented by the Error = Y ‐ 𝑌

distance between the actual value ( ) and 𝑋, 𝑌)


the fitted (or predicted) value given by
the straight line for the same
 That is, =
 This error is also called residual in regression
analysis, and denoted as 𝑒
37
Least Squares Estimation
Cont’d

 We must consider the entire set of ( 𝑖 , 𝑖 ), = 1, …, , for


determining the goodness of fit
 Consider an observed set of ( 𝑖 , 𝑖 ), = 1, …, , suppose
there exists a straight line

such that it minimizes the sum of squared errors (SSE)

Min. SSE =
 Least‐squares criterion is about finding such for 𝑏0 and 𝑏1
 The resulting line is often called the least‐squares regression line

38
Least Squares Estimation
Cont’d

 It is possible to show using calculus that the least‐


squares form of and can be determined as

and ∑

 𝑏 and 𝑏 are the least squares estimates for 𝛽 and 𝛽


respectively

39
Least Squares Estimation
Cont’d

 The estimated is also related to the sample coefficient


of correlation as follows

 Since and are non‐negative, will have the same


sign as

40
Least Squares Estimation – Example

 The following table gives data collected last year for


seven employees of a company
 = Number of years of service
 = Number of days taken off work
𝟐 𝟐
𝑋 𝑌 𝑿 𝑿 𝑿 𝑿 𝒀 𝒀 𝒀 𝒀 𝑿 𝑿 𝒀 𝒀
2 8 ‐3 9 1 1 ‐3
5 7 0 0 0 0 0
7 5 2 4 ‐2 4 ‐4
3 12 ‐2 4 5 25 ‐10
8 3 3 9 ‐4 16 ‐12
3 9 ‐2 4 2 4 ‐4
7 5 2 4 ‐2 4 ‐4

𝑿 5 𝒀 7 ∑ 𝟎 ∑ 𝟑𝟒 ∑ 𝟎 ∑ 𝟓𝟒 ∑ 𝟑𝟕 41
Least Squares Estimation – Example
Cont’d

 Therefore

or

 The least‐squares regression line is

where 𝑌 = predicted or fitted value of 𝑌 for


a given value of 𝑋
 𝑌 𝑌 for all sample values if and only if
𝑟 1
42
Predictions in Regression Analysis
– Example Cont’d

 Suppose we want to predict the number of days off work


this year for employees with 0, 5, 6, 8 and 14 years of
service
 All we have to do is to substitute these given values into
the estimated regression equation
 For 𝑋 = 0, 𝑌 12.45 1.09 0 12.45 days off work
 For 𝑋 = 5, 𝑌 12.45 1.09 5 7 days off work
 For 𝑋 = 6, 𝑌 12.45 1.09 6 5.91 days off work
 For 𝑋 = 8, 𝑌 12.45 1.09 8 3.73 days off work
 For 𝑋 = 14, 𝑌 12.45 1.09 14 2.81 days off work
What???
43
Interpreting the Estimated
Coefficients – Example Cont’d

 Interpreting : From the prediction of for = 0, we


see that = 12.45 is the predicted number of days off
for an employee with 0 years of service
 We should not take this interpretation seriously as this probably
would never happen
 The level 𝑋 = 0 is beyond the range of data studied
 Linearity assumption seems reasonable in the range of 2 and 8
years of service as shown by the data, it would be dangerous to
extrapolate our conclusions far outside that range

44
Interpreting the Estimated
Coefficients – Example Cont’d

 Interpreting : Subtracting the prediction for = 5 (i.e.


= 7) from the prediction for = 6 (i.e. = 5.91) gives
= ‐1.09, thus is the change in the estimated number
of days off for an additional year’s service
 We are estimating that each 1 year increase in service leads, on
average, to a decrease of 1.09 days off work

45
Interpreting the Estimated
Coefficients – Example Cont’d

 The regression line gives a non‐sense prediction of ‐2.81


days off work for = 14 years of service, because
 The relationship between 𝑋 and 𝑌 is approximately linear over
the range covered by the sample, but the regression line cannot
be extended indefinitely without cutting the 𝑋‐axis
 Once we go beyond the sample range, the relationship may
cease to be approximately linear
 We should only predict within the range of observed 𝑋 values

46
Developing Regression Model
in Excel
 Find “Data Analysis” in the “Data” menu bar

 Choose “Regression” at “Data Analysis” browser

47
Developing Regression Model
in Excel Cont’d

Data Cells for 𝒀 and 𝑿 variables


 Data

Output Cell

48
Developing Regression Model
in Excel Cont’d

 Output

|𝒓𝑿𝒀 | Multiple R measures the correlation


between 𝑌 and 𝑌

SSE

𝒃𝟎
𝒃𝟏

49
Coefficient of Determination

 By comparing the actual against predicted values, we


obtain the errors (
 When 𝑋 5, 𝑌 7, 𝑌 7, 𝑒 0
 When 𝑋 8, 𝑌 3, 𝑌 3.73, 𝑒 0.73
 It over‐estimates the number of days off work
 This does not mean our model is bad as the regression
line can never make a precise prediction without errors
unless the linear association is perfect

50
Coefficient of Determination
Cont’d

 The least‐squares regression line minimizes the sum of


squared errors, . In theory, no
other straight line will give a smaller value of SSE for the
same set of data
 In general, the smaller the amount of SSE, the better the
data fit to a straight line
 However, SSE is scale dependent, it can be made as large
or as small by adjusting the scale of

51
Coefficient of Determination
Cont’d

 A better way to measure the goodness of fit for a least‐


squares regression line is to compare its SSE value to that
of another regression line based on the same set of

 A natural second line to be compared with is ,
that is, estimating the mean value of without using
 The corresponding SSE is

SST
 SST is called the total variation in 𝑌 or the total sum of squares

52
Coefficient of Determination
Cont’d

 The goal is to determine by how much the SSE is smaller


than SST
 Or, the amount of improvement in using the regression line and
the independent variable 𝑋 rather than just the sample mean to
predict 𝑌
 This measure is provided through a statistic called the
coefficient of determination ( )

 𝑅 is unit‐free with value in between 0 and 1 inclusive


 The higher the 𝑅 , the better the fitting (the stronger linear
association between 𝑋 and 𝑌)
 However, it does not mean that 𝑋 causes 𝑌 53
Coefficient of Determination –
Example Cont’d

 Thus, in our example on number of days taken off work


𝑋 𝑌 𝒀 𝑒 𝒆𝟐
2 8 10.27 ‐2.27 5.1529
SSE = 13.7354
5 7 7 0 0
SST = 54
.
7 5 4.82 0.18 0.0324
3 12 9.18 2.82 7.9524
8 3 3.73 ‐0.73 0.5329
3 9 9.18 ‐0.18 0.0324
7 5 4.82 0.18 0.0324

𝑿 5 𝒀 7 ∑ 𝟎 ∑ 𝟏𝟑. 𝟕𝟑𝟓𝟒

54
Coefficient of Determination –
Example Cont’d

 Commonly, the coefficient of determination is


interpreted as
 74.56% of the sample variability in 𝑌 is explained by its linear
dependency on 𝑋
 Or, alternatively, by taking the linear dependence on 𝑋 into
account, the SSE is reduced by 74.56%

55
Coefficient of Determination
Cont’d

 In a regression model containing only one variable,

 Hence, in our example, the sample correlation coefficient


between and is
 We know 𝑟 has a negative sign because 𝑏 is negative
 𝑟 would have a positive sign if 𝑏 was positive

56
Inferences about the Slope
 At times, tests concerning are of interest, particularly
one of the forms: H0: vs H1:
 If , there is no linear relationship between and
 The means of the probability distribution of 𝑌 are all equal,
namely 𝐸 𝑌 𝑋 𝑥 𝛽 0𝑥 𝛽 for all levels of 𝑋
 A change in 𝑋 does not induce any change in 𝑌
 Similar to those discussed in Topics 6 & 7, we need to
consider the sampling distribution of , the least squares
point estimate of , in order to perform the inferences
on

57
Inferences about the Slope
Cont’d

 The population regression line is defined as

 It is very common to assume that the error terms are


independent and normally distributed with mean 0 and
variance , = 1, …,
 This assumption can be relaxed, but it will make the inference on
the slope parameter (and others) more complicated
 Under this assumption, the dependent variables 𝑖 are
also independent and normally distributed with mean
and variance , = 1, …,
 We are treating 𝑋𝑖 as known constants
58
Inferences about the Slope
Cont’d

 Sampling distribution of 1
 Since the 𝑌𝑖 are normal, the estimator 𝑏1 is also normal. It can be
shown that 𝑏1 has mean and variance

𝐸 𝑏 𝛽 𝜎 ∑

 The variance 𝜎 can be estimated by 𝑆 as


⁄ ∑ ⁄
𝑆 ∑ ∑ ∑

 𝑆 measures the variability in the slope of regression lines arise


from different possible samples
 𝑆 is called the mean squared error (MSE) of the regression model.
It measures the variance of the errors around the regression line. It
is an unbiased estimator of 𝜎
59
Inferences about the Slope
Cont’d

 Confidence intervals for the population regression slope


 Since 𝑏1 is normally distributed, when 𝜎 is estimated by 𝑆 ,
the statistic
~ 𝑡 with 𝑛‐2 degrees of freedom
 If the error term 𝜀 are normally distribution as assumed, a
100(1‐𝛼)% confidence interval for the population regression
slope 𝛽 is given by
𝑏 𝑡 ⁄ , 𝑆 ,𝑏 𝑡 ⁄ , 𝑆
where 𝑡 ⁄ , is the value corresponding to an upper‐tail
probability of  / 2 from the 𝑡 distribution at degrees of freedom
𝑛 2
60
Inferences about the Slope
Cont’d

 The confidence interval for the population regression


slope is interpreted as
 The 100(1‐𝛼)% confidence interval for the expected change in 𝑌
resulting from one‐unit increase in 𝑋 is between 𝑏
𝑡 ⁄ , 𝑆 ,𝑏 𝑡 ⁄ , 𝑆

61
Inferences about the Slope –
Exercise Cont’d

 Refer to the example on number of days taken off work,


given , and
 A 95% CI for is
95% CI for
⁄ ,

62
Inferences about the Slope
Cont’d

 Hypothesis testing for


 For hypotheses 𝐻0: 𝛽 0 and 𝐻1: 𝛽 0 , the 𝑡 test statistic is
𝑡
 Critical value approach
 At 𝛼 significance level, reject 𝐻0 if 𝑡 𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒 or 𝑡
𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒 where the critical values are obtained from the 𝑡
distribution table at 𝑛 – 2 degrees of freedom
 𝑝‐value approach
 𝑝‐value = 𝑃 𝑡 𝑡 𝑃 𝑡 𝑡
 Reject 𝐻0 if 𝑝‐value 𝛼

 The same 𝑡 can also be used for testing the hypotheses


𝐻0 : 𝛽 0 vs 𝐻1 : 𝛽 0 , or 𝐻0 : 𝛽 0 and 𝐻1 : 𝛽 0
63
Inferences about the Slope –
Exercise Cont’d

 In the example on number of days taken off work , test


at 5% level of significance, is years of service linearly
influencing the number of days taken off work?
𝐻 : Given 𝑏 1.09 and 𝑆 0.2842,
𝐻: 𝑏
𝑡
At 𝛼 0.05 𝑆
𝑛 7 𝑑𝑓 5
Critical Value =
At 𝛼 0. 05,
Reject 𝐻 if

64
Developing Regression Model
in Excel
 Data

Level of confidence

65
Developing Regression Model
in Excel Cont’d

 Output

|𝒓𝑿𝒀 |
𝑹𝟐

𝑺𝒆
𝒏

SSE
SST

𝒃𝟎
𝒃𝟏 𝑺𝒃𝟏

𝒕 for 𝒑‐value 95% CI for 𝜷𝟏 90% CI for 𝜷𝟏


𝜷𝟏 𝟎 for 𝜷𝟏 𝟎 66
Calculating Correlation and
Regression Coefficients in Calculator
(For Casio fx‐50FH II)
 REG mode (MODE 5)
 Select 1 (Lin) for linear relationship
 Input data *: X comes first; separate X and Y with a,;
press M+ after each pair of data input
 Intercept: SHIFT 2 (S‐VAR) 1 (VAR) 1 EXE
 Slope: SHIFT 2 (S‐VAR) 1 (VAR) 2 EXE
 Correlation: SHIFT 2 (S‐VAR) 1 (VAR) 3 EXE

* Before inputting the new data set, always clear previous data: SHIFT 9 1 EXE
67
Applications of Linear Regression

 Hong Kong Population


 Time Series Model
 Centa‐City Index
 Multiple Linear Regression

68
Time Series Model

 Attempt to predict future by using a stream of historical


data
 Assume what happened in the recent past will continue
in the near future
 Time is used as the only independent variable

 Where 𝑌 = Predicted value at time period 𝑡
 For time series data exhibit some trend in a long‐range
time horizon

69
Hong Kong Population

 Census and Statistics Department (C&S) published the


Hong Kong Annual Digest of Statistics so as to provide
detailed annual statistical series on various aspects of
the social and economic developments of Hong Kong
 Yearly data on Hong Kong’s population from 1961 to
2022, totalling 62 observations are downloaded from
C&S’s website
 Let denote the population size (in thousands)
= 1, 2, 3… denote the sequence of time
with 𝑋 = 1 representing the year 1961
𝑋 = 2 representing 1962, etc.
70
Hong Kong Population
Cont’d

 A scatter plot of vs. reveals the following


Hong Kong Population Size (1961 - 2022)
8000

7000

6000
Population Size ('000)

5000

4000

3000

2000

1000

0
1960 1970 1980 1990 2000 2010 2020
Year

 The association between and appears to be


approximately linear
 It therefore makes sense to write 71
Hong Kong Population
Cont’d

 Using the aforementioned least squares method and


Excel, the following regression output has been obtained

 What can you tell from this output? 72


Hong Kong Population
Cont’d

1.

2.
 So, 𝑏 3465.7677 thousands is the predicted Hong Kong
population size for the year 1960 (𝑋 = 0)
 𝑏 72.2625 thousands is the predicted average annual
increment in population size
3.

73
Hong Kong Population
Cont’d

4. has a high significant linearly relationship to , as


and ‐value is close to zero for testing
vs.
5. The predicted Hong Kong population sizes for 2023 – 2026
are
 2023 (𝑋 63): 𝑌 3465.7677 72.2625 63
 According to C&S, the actual population size was 7503.1 thousands
 2024 (𝑋 64): 𝑌
 2025 (𝑋 65): 𝑌
 2026 (𝑋 66): 𝑌
 By the end of 2026, the Hong Kong population size is expected to excess 8.2
millions

74
Hong Kong Population
Cont’d

 Of course, the accuracy of these forecasts depends, among


other things, on the legitimacy to extend the linear relationship
established based on the sample values beyond the estimation
period
 It is a commonly used method for predicting time series data
 These forecasts are called “ex‐ante” forecasts since the actual
values of the variable being predicted are unknown at the time
of prediction

75
Multiple Linear Regression

 In many situations, two or more independent variables


may be included in a regression model to provide an
adequate description of the process under study or to
yield sufficiently precise inferences
 For example a regression model for predicting the
demands for a firm’s product in different countries uses
socioeconomic variables (mean household income,
average years of schooling of head of household),
demographic variables (average family size, percentage
of retired population), and environmental variables
(mean daily temperature, pollution index), etc.
76
Multiple Linear Regression
Cont’d

 Linear regression models containing two or more


independent variables are called multiple linear
regression models
 The simple linear regression model can be extended to
include independent variables

77
Centa‐City Index
Source: [Link]
home‐market‐smashes‐through‐record

78
Centa‐City Index
Cont’d

 Why Property Price Indices?


 Investors and potential homebuyers are in need of indicators to
study the current movement of property prices in Hong Kong
 The creation of the “Centa‐City Index” aims to provide such
information to the public as a source of reference on trends in
Hong Kong’s property market
 How are the Index constructed?
 Regression analysis is used to determine the effect of various
attributes on property price
 Attributes such as floor area, years of occupancy, location,
direction, view, floor level, etc. are considered

79
Centa‐City Index
Cont’d

80
Centa‐City Index
Cont’d

81
Centa‐City Index
Cont’d

• More information
• [Link]
82
• [Link]

You might also like