0% found this document useful (0 votes)
14 views18 pages

Constant Variance in Regression Analysis

[1] The document discusses weighted least squares regression as a method to address violations of the constant variance assumption in ordinary least squares regression. [2] It assumes observations have different variances that are inversely proportional to assigned weights, with higher-variance points getting lower weights. [3] The weighted least squares approach minimizes a weighted residual sum of squares to produce estimates of regression coefficients that account for non-constant variances.

Uploaded by

Zhen Wang
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views18 pages

Constant Variance in Regression Analysis

[1] The document discusses weighted least squares regression as a method to address violations of the constant variance assumption in ordinary least squares regression. [2] It assumes observations have different variances that are inversely proportional to assigned weights, with higher-variance points getting lower weights. [3] The weighted least squares approach minimizes a weighted residual sum of squares to produce estimates of regression coefficients that account for non-constant variances.

Uploaded by

Zhen Wang
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 4

STATS 101A Introduction to Data Analysis and Regression


Maria Cha
Model Assumption: Constant Variance
• In the population linear model,
𝑌! = 𝛽" + 𝛽#𝑥! + 𝑒!
, where 𝑒! ~𝑁(0, 𝜎 $).

• Assumption of ‘constant variance’: 𝑉𝑎𝑟 𝑒! = 𝜎 $

• Implication : Each data point provides equally precise


information about the deterministic part of the total
variation of the model. In other words, the standard
deviation of the error term is constant over all values of
the predictor or explanatory variables.
Model Assumption: Constant Variance
• Recall: If the assumption is violated, all inference
tools are no longer valid.

• To fix the problem,


• 1) Transform the variables (either 𝑋 or 𝑌 or both)
– done in Chapter 3.
• 2) Assign a reasonable ‘weight’ to each variance so
that the non-constant behavior is controlled
– Weighted Least Square method
Weighted Least Square
• Idea : If some observations have larger variance in their
error terms, then we make their variance smaller. If
some have smaller variance, we make it bigger.

• Greater variance implies ‘unreliability’ of the observation.


Thus, we want to put a ‘smaller’ weight to such
observations.

• Thus, the weights (𝑤! ) are inversely proportional to the


corresponding variances; points with low variance will
be given higher weights and points with higher variance
are given lower weights.
Weighted Least Square
• Example: Consider the data with non-constant variance.
𝜎$
𝑉𝑎𝑟 𝑒! |𝑥! = 𝑉𝑎𝑟 𝑌! |𝑥! =
𝑤!

• It is evident that the variance increases proportional to 𝑥! :

%! #
𝑉𝑎𝑟 𝑒! |𝑥! = &"
→ 𝑤! = '"
Weighted Least Square
𝜎$ 1
𝑉𝑎𝑟 𝑒! |𝑥! = → 𝑤! =
𝑤! 𝑥!

• For example,
𝑥# = 1, 𝑤# = 1 → 𝑉𝑎𝑟 𝑒#|𝑥# = 𝜎 $
1
𝑥$ = 2, 𝑤$ = → 𝑉𝑎𝑟 𝑒$|𝑥$ = 2𝜎 $
2
1
. 𝑥( = 3, 𝑤( = → 𝑉𝑎𝑟 𝑒(|𝑥( = 3𝜎 $
3

• Again, the weights (𝑤! ) are inversely proportional to the


corresponding variances. Then, how do we control/reflect
the non-constant variance to our estimation of 𝛽" and 𝛽#?
Weighted Least Square
• Recall: Ordinary least square (OLS) estimates 𝛽" and 𝛽#
are minimizing RSS of the model.
𝑅𝑆𝑆 = <(𝑦! − (𝑏" + 𝑏#𝑥#))$

• Weighted least square estimates for 𝛽" and 𝛽# are


minimizing Weighted RSS of the model, where the
model is

𝑌! = 𝛽" + 𝛽#𝑥! + 𝑒!

with 𝑒! ~𝑁(0, 𝜎 $/𝑤! ). . Thus, the Weighted RSS is


𝑊𝑅𝑆𝑆 = < 𝑤! (𝑦! − (𝑏" + 𝑏#𝑥#))$
Weighted Least Square
• Find 𝑏" and 𝑏" that minimizes WRSS and denote them
as 𝛽B"& and 𝛽B#& .

• Slope:
∑ 𝑤! (𝑥! − 𝑥̅& )(𝑦! − 𝑦E& )
𝛽B#& =
∑ 𝑤! (𝑥! − 𝑥̅& )$
, where 𝑥̅& = ∑ 𝑤! 𝑥! / ∑ 𝑤! and 𝑦E& = ∑ 𝑤! 𝑦! / ∑ 𝑤! .

• Intercept:
𝛽B"& = 𝑦E& − 𝛽B#& 𝑥̅&

• Note: The WLS estimates are also unbiased.


How to find 𝑤!
• Experience or prior information using some
theoretical model.

• Some common cases:


• 1. If we expect an increasing relationship between
𝑉𝑎𝑟(𝑒̂! ) and 𝑥! , i.e. Poisson data, then 𝑉𝑎𝑟 𝑒̂! =
𝜎 $ 𝑥! . Thus, 𝑤! = 1/𝑥! .
• 2. If the responses are the average of 𝑛!
observations at each 𝑥! , then
𝜎$
𝑉𝑎𝑟 𝑒̂! = 𝑉𝑎𝑟 𝑦/! =
𝑛!
Thus, 𝑤! = 𝑛! .
Example
• Consider the data [Link]
• The aim of the study was to develop a regression equation to
model the relationship between the number of rooms cleaned,
𝑌 and the number of crews, 𝑋. Below are the scatter plot with
OLS regression line and the corresponding residual plot.
Example
• So it appears that a linear relationship is reasonable,
but the constant variance assumption isn’t.

• In this case, the 𝑥-variable is discrete with values


2,4,6,8,10, 12, and 16. It appears that the standard
deviation of the residuals might be proportional to
variance of 𝑌 for each 𝑋,
𝜎 $ ∝ 𝑆𝐷(𝑌|𝑋)$

• This suggests an analysis with


1
𝑤! ∝
𝑆𝐷(𝑌|𝑥! )$
Example
• We refit the data with WLS model:

OLS model >>>


Example
• It does not show much difference when we compare the
fitted lines from Ordinary LS vs. Weighted LS.
Example
• To check how the variance of the residuals is
constant, consider the variance of the updated
residuals, i.e.
𝑉𝑎𝑟 𝑒! = 𝜎 $ /𝑤!
Then,
𝑉𝑎𝑟 𝑤! 𝑒! = 𝜎 $

• In the residual summary of the weighted least squares


analysis, this is based on 𝑤! 𝑒! instead of the raw
residuals 𝑒! . In addition, to see if a reasonable
weighting has been done, plot 𝑤! 𝑒! instead of 𝑒! .
Example
• The WLS model now suggests the constant variance
in the residuals.
Advantages
• In the transformed model, the interpretation of the
coefficient estimates can be difficult. In weighted
least squares the interpretation remains the same as
before.

• Weighted least squares gives us an easy way to


remove one observation from a model by setting its
weight equal to 0.

• We can also down-weight outlier or influential points


to reduce their impact on the overall model.
Limitation
• The biggest disadvantage of weighted least squares is
the fact that the theory behind this method is based on
the assumption that the weights are known exactly.

• This is almost never the case in real applications, of


course, so estimated weights must be used instead.

• When the weights are estimated from small numbers of


replicated observations, the results of an analysis can
be very badly and unpredictably affected. It is important
to remain aware of this potential problem, and to only
use weighted least squares when the weights can be
estimated precisely relative to one another.
Limitation
• Weighted least squares regression, like OLS regression,
is also sensitive to the effects of outliers. If potential
outliers are not investigated and dealt with appropriately,
they will likely have a negative impact on the parameter
estimation and other aspects of a weighted least
squares analysis. If a weighted least squares regression
actually increases the influence of an outlier, the results
of the analysis may be far inferior to an un-weighted
least squares analysis.

You might also like