Influential Observations in Regression
Influential Observations in Regression
___ • .4""'•
't-·
INFLUENTIAL OBSERVATIONS
IN LINEAR REGRESSION
by
R. Dennis Cook
February, 1977
'.j· .··
'~:; ;~~-
')_·
Abstract
and the convex hull of the observed values of the independent variables.
presented.
lo INTRODUCTION
where Y is an [Link] vector .of observators, Xis an nxp full rank matrix
constructed by first specifying the design space (eogo some closed convex
subset of RP) and then choosing the design points within the space so
the configuration of the design points in the factor space can have an
Behnken and Draper (1972) have noted that a wide variation in the
points. Box and Draper (1975) suggest that for a design to be insensitive
of Davies and Hutton (1975) also reflect the opinion that the residual
least squares estimate of ~-- The measure combines information from the
studentized residuals and the residual variances, and shows that design
points with relatively small residual variances will tend to be the more
the residual variances are constant). However, the role that such an
given by
where
,. ,..
R =Y-Y=Y-~
determined by the design points (ioe. the rows of X). ,We find, it
Cook (1977) is reviewed, some comments on its use are given, and
2. INFLUENTIAL OBSERVATIONS
Cook (1977) proposed that the importance of the ith data point be
without the point an~, second, measuring the distances, D., between the
1.
2
where s = R'R/(n-p). A large value of D. indicates that the associated
-- 1.
50% probability point, then the removal of the ith data point moves the
least squares estimate to the edge of the usual 50% elliptical confidence
,..
region for ·f centered at ~:
Let I
v •. = x.(X I
X) "1x. where x.I is the ith row of X.. The
1.l. -l. - - -l. -l.
2 ,..
t. V(yi)
l.
D.1. = -
p
(2)
v(ri)
- 5 -
function of the likelihood ratio test statistic for the hypothesis that
. I . 1· -1
X (XX)
- - - - <- max. v ..•
X
].
].].
(3)
This follows because the levels of constant value of the quadratic form
on the left of relationship (3) are ellipsoids and the ellipsoid passing
IVH. Expression (3) shows that the point with the largest
,...
Thus, large values of V (y . ) /V ( r . ) indicate "outlying" design
1. J.
pointso Of course, the point with the largest prediction variance need not be the
one whose Euclidean distance from the center of the design is the greatest
since the values of ~'(X'X)-lx depend on the density of the points in the
. I I -1 1
X • (X X) x. < -k (4)
-J - - -J -
for k=l and assuming that it is true for k=k0 • It can be shown to be
(X'X)- 1x.x~(X'X)-l
= (X'X)-1 - - -J J
1
1 + x~(X'X)- x.
-J - - J
of x .•
-J
Generally we may anticipate that the design point corresponding to
~
max~v .. (max V(y.)/V(r.)) will. lie on the boundary of the IVH in a region
1.1 J. 1.
Since the ith design point lies on an ellipse whose value is v.. any
l.l.
gap in the spacing of the design pointso A large gap may be taken as an
space where a final model may be appropri_ate for the purpose of predictiono
Prediction at any point ~ for which (3) does not hold may be
design points are virtually the only points for which the model is
be an outlier (t~ is large) and at the same time lies in a high density
l.
~
same. The same conclusion may hold if a point lies on the boundary of
indicates that the solution is not stable across the full IVH
This should be cause for concern since the final form of the solution
subject for future investigationo Our main contention is that the use of
Let D.J ( -l.• ) denote the value of the distance measure for the jth point
based on the data set from which the ith point has been removed, i~j.
- 9 -
Let
. -1
V. . = X~ (X 1 X)
X.
l.J -1. -- - -J
and
)-1
wki, - ~ ~{-i)!(-i)
- '( I I
~J,
where
x ' . x . = -x 'x- -
-( -1.)-(-1.)
x. x ~
-1.-1.
2
and x' is the·rth row of X. Also, let 8 (-i) denote the usual
-r
2
estimate of a based the data set with the ith point removedo
v .. =w • ./(1 +w .• ) (5)
l.J l.J l.l.
2
v .. = w .• - w• ./(1 + w •• ) (6)
JJ JJ l.J ].].
,._ ,._ I -1 I"°
A-A( ")=(XX)
~ .i::. -1. - -
x.[y.]. - -1.-
-1.
x.~)/(1 - v ].].
.. ) (7)
2 2 ,,.. 2,
(n-p)s = (n-1-p)s ( -1.") + (y.]. - -1.~
x.A),(1 - v ].].
.. ) o (8)
Expressions (5) and (6) are easily verified using the identity
, -1 , ·, -1
(!'X)-1
' -1 <!c -i)!(-i)) ~i~i <!c -i~(-i) >
= (X(-i)~-i)) 1 + wl.1.
••
Expression (7) was shown by Cook (1977) and expression (8) was shown by
Beckman and Trussell (1974)0 The derivations are straightforward and will
2
DJ"( -J.• ) = t.( .)w . ./p(l
J - i JJ
- w •. )
JJ
(9)
2
where tJ·c .)-i
denotes the studentized residual for the jth point in the
data set with the ith point removed. We now relate equation (9) to the
full data set by dealing with t~( ") and w .. separately. Consider
J -J. JJ
first w .. : Using equations (5) and (6) it is easily verified that
JJ
2
V • • = W• • .,. p •• (1 - V •• )
JJ JJ J.J JJ
where denotes the correlation coefficient between the ith and jth
p ••
J.J
residuals in the full data set. It follows that
2 2
v .. (1 - p •• ) + p ••
w • ./(1- w •• ) =
]] J.J J.] (10)
2 0
JJ JJ (1 - v .. )(l - p •• )
JJ J.J
A large value of v.. would have been, detected in the analysis of the
JJ
full data set. Thus, if the variance of the jth predicted value increases
correlation between the ith and jth residuals in the full data set. We
see also that if the residual correlations are negligible then the variances
of the predicted values will remain essentially unchanged when any point
is deleted.
Next,
Using the results in equations (5) - (7) a little algebra will verify that
2 2 2 2
t.J ( -1.. ) = s [t. p •• t.] /s( .)(1
J.J 1. col.
J
where t. and t. are studentized residuals from the full data seto Using
1. J
equation (8) this reduces to
2
2
(n - p - l)(t. - p •• t.)
1 1.J 1.
(11)
J -1..) =
t.( 2 2
(n - p t. )(1 p •• )
]. 1.J
deletion of any point with a studentized residual larger than one will
2 2 2
(n - p - 1) [ t. ""' p •• t.] V •• ( 1 - p •• ) + p 1.· .
= J 1.1 1. 11 · 1.1 1_ (12)
D.J ( -1.• ) 2
(n - p - ti)(l - p~j) (1 - vjj)p
Section 4o
- 12 -
of ! are zero are commonly used to simplify the original model. When
,..
Further, let ~k(-i), Tk(-i), dk(-i), and Fk_(-i) denote the analogous
quantities based on the data set without the ith observationo
sk(
I"'
") = sk
-]. I"'
- ck.r./(1
1. 1. - v 1.1.
.. )
where
and ..=k is a pxl vector with a 1 in the kth position and zeros elsewhereo
Also,
2
dk(-i) = dk + cki/(l - vii) 0
"2 2
Fk(-i) = ~k(-i)/ 8 (-i)dk(-i)
2
Tk V.. \ 2
(n-p-l)t. l t.i
l.
-v <r-~;.:) 1
(12a)
Fk(-i) - 2 2
(1 +y v . ./(1-v .. ))
(n-p-t.) J.l. l.l.
l.
,... I ,..
where y denotes the correlation between ~k and x. ~.
-1. -
measure, D., and will be relatively large for points on the boundary of
l.
Next, consider the deletion of a point that fits the model quite
2
well (t.
].
<
- 1). In this case,
. \ 2
[Tk - yt. (v . ./1-v .. ) ]
~ ]. l.l. l.l.
2
1 +y V• ./1-v ..
u. 11.
2
•
[VTkv . ./(1-v •. )
l.l. .].].
+ t.]
l.
t~].
1 + V 2v l.l.
. ./1-v ..
l.l.
Thus, we can generally expect all partial F-statistics greater than one
is deletedo
- 15 -
4. RESIDUAL CORRELATIONS
2
The squared correlation coefficient, p •• , between the ith and jth
l.J
2 2
p • • = v. ./ ( 1-v .. )( 1-v. . ) o (13)
l.J l.J l.l. JJ
2
To investigate the causes of a large value for p •• , we shall hold the
1.J
jth design point fixed and find how to choose the ith design point so
over sowe convex subset of the factor space that consists of all
permissible values for the ith design pointo The required calculation
2
is facilitated by writing p •• in terms of explicit quadratic forms in
l.J
2
w. . and w •• are quadratic forms in
-1.
x. while w •• is independent
1.J 1.l. JJ
of x. and may be considered constanto If the model contains a constant
-1.
Let x; = (1, z;) and, assuming that the independent variables in the
1
0
/
(X{-i) X(-i))
-1
= ( n:1 A )·
... 16 -
substituting
w ••
J.J
=--
1
n- 1
+z!Az.
-J -1. -
The largest possible value for p7.J.J will obviously depend on the
subset of the factor space over which the supremum is takeno Lacking
seto From .equation (14) it is not difficult to see that the supremum
2
must be obtained on the boundary of G(min[c, n (w .. - 1/(n-l))])o It
JJ
follows that
c' Q~.
2
sup Pij = (15)
[Link]
-l.
( 1-W .. ){ 1+ _1.._l
JJ n-
+ C
1
) + C
I
Q:JJ.
where
1 1 ~
Q .• =
JJ
+ ( w .. -
JJ
-)2
n-1
<'~-l)Jc I
and
2 1
c' = min(c, n (w .. - )) o
JJ n- 1
- 17 -
points whose residuals are the most highly correlatedo It may be,
2
sup p •• ~ w. . •
[Link] l.J JJ
-1.
other pointso
5. OUTLIERS
.!_=!i+!+.!-
consisting of n-1, zeros and .e, < n-p unknown parameterso Without
loss of generality we may assume that we wish to test the last .e,
! = ! ~ + Jze ~ + .! (17)
-1
where ct = (~'X) !1 (X J?.. + ~)
T = (I
....... - ____ X}-l -X')
X(X 1........
z = (o(n-.e,)x1,)
I.e,x.e,
and !.e, is the i,xl vector consisting of the last .e, components of 80
- 20 -
,..
ft = (Z'T'Tz)-l
--~ --- z'T'Y
,.. -1
!t = (!_t) RJ,
/ -1
R:t T.e, R:t (19)
H:8:t = 0 is
. alR
Rif.J. --'J, -1 n-t-p
FJ, = R;R - l',,r,, RJ, t
---
The dependence on the studentized residuals and residual correlations
(20)
-1
n-p-t!o t
-lll-L -t
where ti and .e_L are the matrices of studentized residuals and residual
correlations for the last L observations. For L ·= 1 we have
studentized residual,
the L most likely outlying points. When this is the case the quadratic form in
(21)
most likely outlying points will not, in general, correspond to the two
valueso
approximation
C3 22 -
-1
.e...e, :!= 21 - .e.1,
t. t .p •.
]. J l.J
question.
disturbing the analysis may not be the ones with the larger
6. EXAMPLE
ammonia to nitric acid. The original data set is from Brownlee (1965) and
(1, 3, 4 and 2l)·were-outliers and that one of the explanatory variables was
not needed. Their final [Link] a linear and quadratic term for one
explanatory variable, a linear term for the other and was based on 17
"valid" observationso
In this example we adopt the final model of Daniel and Wood but
include all data points. For ease of reference the data set has been
subsets of the data from Table lo Tables 2, 3, and 4 give the values
of D.,
].
v ].].
.. , and t.,
].
respectively, for the six data sets. Table 5
square error for each data seto Note that the last data set used in
each table consists of the observations that Daniel and Wood judged
valid.
for this importance can be obtained from Tables 3 and 4. The four
and are replicates. The spacing of these values indicates two gaps in
the coverage of the design space. Observation 21 has the third largest
that when inspecting the studentized residuals we are implicitly assuming the
has at most one nonzero component). The two most likely candidates as
4 and 21. These observations also have the two largest studentized
seems intuitively obvious and is easy to show that the tests of this
hypot~esis is the same as the test of H:e. = 0 based on the data set
1
- 25 -
with observation 21 removed. The second column in each table shows the
results after the removal of observation 21. Note that now observation 4
on the residual correlations P!, 21 (i = "i, 20) in the full data seto
confidence ellipseo The reasons for this importance are that observation 2
may be rejected at the Oo05 level.) From Table 5 we see that now the
2 A
coefficient of x , ~ , is highly significanto The increase in the
1 3
A
partial F-statistic for ~ is due to the large value of t and a
3 4
A A
small value of the correlation between ~
3
and ~~ (y = 0957) in
column in each table shows this change. Notice that now the gaps .. in the
In the latest data set (2, 4, and 21 deleted) design points 1 and
the results based on the subset of the data that Daniel and Wood judged
of this observation would move the least squares estimates beyond the
quite well and is important because it stands alone on the edge of the
removed the first three partial F-statistics are all less than one while
F decreases but remains fairly largeo It appears that for the final
4
data set of Daniel and Wood the quadratic term is needed to model a single
observation.
REFERENCES
[2] Behnken, D.W. and Draper, N.R., (1972). Residuals and their variance
patterns. Technometrics 14, 101-lllo
[3] Box, [Link] and Draper, N.R., (1975)0 Robust designo Biometrika
62, 347-3520
[6] Davies, R.B. and Hutton, B., (1975)0 The effects of errors in the
independent variables in linear regression. Biometrika g,
383-3910
[9] Lund, R.E., (1975)0 Tables for an approximate test for outliers
in linear models. Technometrics 17, 473~476.
TABLE 1
1 80 6400 27 42
2 80 6400 27 37
3 75 5625 25 37
4 62 3844 24 28
5 62 3844 22 18
6 62 3844 23 18
7 62 3844 24 19
8 62 3844 24 20
9 58 3364 23 15
10 58 3364 18 14
11 58 3364 18 14
12 58 3364 17 13
13 58 3364 18 11
14 58 3364 19 12
15 50 2500 18 8
16 50 2500 18 7
17 50 2500 19 8
18 so 2500 19 8
19 50 2500 20 9
20 56 3136 20 15
21 70 4900 20 15
TABLE 2
Observations Deleted
Observation None (21) (4,21) (2 ,4, 21) (1,2,4,21) (1,3,4,21)
1 0.162 0.107 0.210 2.042 * *
2 0.193 0.593 1.331 * * 12.175
3 0.125 0.123 0.333 0.403 21.160 *
4 0.304 0.539 * * * *
5 0.003 0.012 0.003 0.006 0.006 0.001
6 0.021 0.037 0.019 0.031 0.033 0.021
7 0.042 0.050 0.007 0.010 0.008 0.003
8 0.014 0.014 0.010 0.023 0.034 0.056
9 0.043 0.040 0.010 0.008 0.002 0.028
10 0.028 0.008 0.013 0.030 0.049 0.051
11 0.028 0.008 0.013 0.030 0.049 0.051
12 0.062 0.009 0.002 0.009 0.019 0.028
13 0.001 0.044 0.107 0.177 0.190 0.213
14 0.001 0.019 0.034 0.054 0.055 0.068
15 0.002 0.007 0.005 0.005 0.002 0.006
16 0.002 0.001 0.007 0.019 0.035 0.027
17 0.004 0.000 0.000 0.001 0.003 0.003
18 0.004 0.000 0.000 0.001 0.003 0.003
19 0.008 0.001 0.008 0.011 0.007 0.006
20 0.008 0.010 0.037 0.078 0.117 0.088
21 0.699 * * * * *
TABLE 3
Values of v .. Based on Selected Subsets
11
of Observations From Table 1.
Observations Deleted
Observation None (21) (4,21) (2,4,21) (1,2,4,21) (1,3,4,21)
1 0.409 0.421 0.421 o. 727 * *
2 0.409 0.421 0.421 * * 0.993
3 0.176 0.199 0.201 0.308 0.983 *
4 0 ._191 0.192 * * * *
5 0.103 0.108 0.125 0.125 0.125 0.131
6 0.134 o. 134 0.164 0.164 0.165 0.169
7 0.191 0.192 0.238 0.239 0.240 0.242
8 0.191 0.192 0.238 0.239 0.240 0.242
9 0.163 0.170 0.206 0.208 0.218 0.208
10 0.139 0.175 0.175 0.176 0.179 0.179
11 0.139 0.175 0.175 0.176 0.179 0.179
12 0.212 0.272 0.275 0.276 0.279 0.280
13 0.139 0.175 0.175 0.176 0.179 0.179
14 0.092 0.110 0.111 0.112 o. 116 0.113
15 0.188 0.189 0.191 0.191 0.195 0.193
16 0.188 0.189 · 0.191 0.191 0.195 0.193
17 0.187 0.195 0.195 0.195 0.198 0.197
18 0.187 0.195 0.195 0.195 0.198 0.197
19 0.212 0.232 0.234 0.234 0.236 0.237
20 0.064 0.064 0.069 0.070 0.076 0.070
21 ' 0.288 * * * * *
TABLE 4
TABLE 5
A
Observations Deleted
None (21) (4 21) (2 4 21)
A : A A A
Term ak Fk ak Fk ak Fk f3k Fk
1 -14.30 0.19 -25.90 1.00 -3.74 0.04 13.26 0.85
xl · -0.46 0.21 0.07 0.01 -0.51 0.80 -1.10 5.95
Term ak Fk ak Flc'
1 33.)4 3.96 -15.41 1.5
xl -1.82 10.61 - 0.07 0.03
X2 0.023 24.09 0.007 4.60
1
x2 0.44 7.81 0.53 12.27