Example G Cost of construction
of nuclear power plants
Description of data. Table G.l gives data, reproduced by permission of the
Rand Corporation, from a report (Mooz, 1978) on 32 light water reactor
(LWR) power ,plants constructed in the USA. It is required to predict the
capital cost involved in the construction of further LWR power plants. The
notation used in Table G.l is explained in Table G .2. The final six lines of data
in Table G.l relate to power plants for which there were partial turnkey
guarantees and for which it is possible that some manufacturers' subsidies
might be hidden in the quoted capital costs.
General considerations. One of the most common problems in advanced applied
statistics is the study of the relation between a single continuous response
varia tie and a number of explanatory variables. When the expected response
can be represented as a linear combination of unknown parameters, with co-
efficients determined by the explanatory variables, and when the error struc-
ture is suitably simple, the techniques of multiple regression based on the
method of least squares are applicable. The formal theory of multiple re-
gression, and the associated significance tests and confidence regions, have
been extensively developed and are described in numerous textbooks; see, for
example, Draper and Smith (1981) and Seber (1977). Further, computer
programs for implementing the methods are widely available.
Nevertheless, there can be difficulties, partly of technique but more impor-
tantly of interpretation, in applying the methods, especially to observational
data with fairly large numbers of explanatory variables. We now mention
briefly some commonly occurring points. Of course, in any particular appli-
cation many of the potential difficulties may be absent and indeed the present
example seems relatively well behaved.
Some issues that arise fairly commonly are the following:
(i) What is the right general [Link] to fit?
(ii) Are there aspects of error structure that seriously affect the analysis?
(iii) Are there outliers or anomalous observations that need to be isolated?
(iv) What can be done if a subset of observations is isolated, possibly not
following the same model as the main body of data?
81
82 Applied statistics
Table G.l. Data on thirty-two LWR power plants in the USA
c D T, T, s PR NE cr BW N PT
460.05 68.58 14 46 687 0 1 0 0 14 0
452.99 67.33 10 73 1065 0 0 l 0 J 0
443.22 67.33 JO 85 1065 I 0 I ·o I 0
652.32 68.00 11 67 1065 0 I I 0 12 0
642.23 68.00 11 78 1065 J I l 0 12 0
345.39 67.92 13 51 514 0 I I 0 3 0
272.37 68.17 12 50 822 0 0 0 0 5 0
317.21 68.42 14 59 457 0 0 0 0 I 0
457.12 68.42 15 55 822 I 0 0 0 5 0
690.19 68.33 12 71 792 0 1 •) I 2 0
350.63 68.58 12 64 560 0 0 0 0 3 0
402.59 68.75 13 47 790 0 l 0 0 6 0
412.18 68.42 15 62 530 0 0 I 0 2 0
495.58 68.92 17 52 1050 0 0 0 0 7 0
394.36 68.92 13 65 850 0 0 0 I 16 0
423.32 68.42 11 67 778 0 0 0 0 3 0
712.27 69.50 18 60 845 0 I 0 0 17 0
289.66 68.42 15 76 530 1 0 I 0 2 0
881.24 69.17 IS 67 1090 0 0 0 0 ' 1 0
490.88 68.92 16 59 1050 1 0 0 0 8 0
567.79 68.75 11 70 913 0 0 I 1 15 0
665.99 70.92 22 57 828 I I 0 0 20 0
621.45 69.67 16 59 786 0 0 1 0 18 0
608.80 70.08 19 58 821 1 0 0 0 3 0
473.64 70.42 19 44 538 0 0 1 0 19 0
697.14 71.08 20 57 !130 0 0 1 0 21 0
207.51 67.25 13 63 745 0 0 0 0 8 I
288.48 67.17 9 48 821 0 0 I 0 7 I
284.88 67.83 12 63 886 0 0 0 1 11 1
280.36 67.83 12 71 886 I 0 0 1 11 I
217.38 67.25 13 72 745 I 0 0 0 8 1
270.71 67.83 7 80 886 J 0 0 1 11 1
Table G.2. Notation for data of Table G.1
c Cost in dollars x I0- 0, adjusted to 1976 base
D Date construction permit issued
T, Time between application for and issue of permit
T, Time between issue of operating license and construction permit
s. Power plant net capacity (MWe)
PR Prior existence of an LWR on same site(= 1)
NE Plant constructed in north-east' region of USA ( = 1)
cr Use of cooling tower ( = 1)
BW Nuclear steam supply system manufactured by Babcock-Wilcox(= 1)
N Cumulative number of power plants constructed by each architect-engineer
PT Partial turnkey plant ( = 1)
Example G 83
(v) Is it feasible to simplify the model, normally by reducing the number of
explanatory variables?
(vi) What are the limitations on the interpretation and application of the
final relation achieved?
All these points, except the key issue (vi), can to some extent be dealt with
formally, for instance by comparing the fits of numerous competing models.
Often, though, this would be a ponderous way to proceed.
Consideration of point (i), choice of form of relation, involves a possible
transformation of response variable, in the present instance cost and log cost
being two natural variables for analysis, and a choice of the nature and form of
the explanatory variables. For instance, should the explanatory variables,
where quantitative, be transformed? Should derived explanatory variables be
formed to inv~stigate interactions? Frequently in practice, any transformations
are settled on the basis of general experience: the need for interaction terms may
be examined graphically or, especially with large numbers of explanatory
variables, may be checked by looking only for interactions between variables
having large 'main effects'. In the present example, log cost has been taken as
response variable and the explanatory variables S, T1 , T 8 and Nhave also been
taken in log form, partly to lead to unit-free parameters whose values can be
interpreted in terms of power-law relations between the original variables. It is
plausible that random variations in cost should increase with the value of cost
and this is another reason for the log transformation.
Complexities of error structure, point (ii), can arise via systematic changes
in variance, via notable non-normality of distribution and, particularly
importantly, via correlation in errors for different individuals. All these
effects may be of intrinsic interest, but more commonly have to be considered
either because a modification of the method of least squares is called for or
because, while the least-squares estimates and fit may be satisfactory, the
precision of the least-squares estimates may be different from that indicated
under standard assumptions. In particular, substantial presence of positive
correlations can mean that the least-squares estimates are much less precise
than standard formulae suggest. A special form of correlated error structure is
that of clustering of individuals into groups, the regression relations between
and within groups being quite different. There is no sign that any of these
complications are important in the present instance.
Somewhat related is point (iii), occurrence of outliers. Where interest is
focused on particular regression coefficients, the most satisfactory approach is
to examine informally or formally whether there is any single observation or
small set of observations whose omission would greatly change the estimate in
question; see also point (iv).
In the present example there is a group of 6 observations distinct from the
main body of 26 and there is some doubt whether the 6 should be included.
84 Applied statistics
This is quite a common situation; the possibly anomalous group may, for
example, have extre:.ne values of certain explanatory variables. The most
systematic approach is to fit additional linear models to test consistency. Thus
one extra paratneter can be fitted to allow for a constant displacement of the
anomalous group and the significance of the resulting estimate tested. A more
searching analysis is provided by allowing the regression coefficients also to be
different in the anomalous group; in the present instance this has been done one
variable at a time, because with 10 explanatory variables and only 6 obser-
vations in the separate group there are insufficient observations to examine
for anomalous slopes simultaneously.
Point (v), the simplification of the fitted model, is particularly important
when the number of explanatory variables is large, and even more particularly
when there is an a priori suspicion that many of the explanatory variables are
measuring essentially equivalent things. The need for such simplification
arises particularly, although by no means exclusively, in observational
studies. More explicitly, the reasons for seeking simplification are that:
{a) estimation of parameters of clear interest can be degraded by including
unnecessary terms in the model :
(b) prediction of response of new individuals is less precise if unn'ecessary
terms are included in the predictor;
(c) it is often reasonable to expect that when many explanatory variables are
available only a few will have a major effect on response and it may be of
primary interest to isolate these important variables and to interpret their
effects;
(d) it may be desirable to simplify future work by recording only a smaller
numb~r of explaratory variables.
Techniques for the retention of variables are, as explained in Section 3.4 of ,
Part I, forward, backward or some mixture. Where some of the parameters
represent effects of direct interest they should be included regardless of the
operation of a selection procedure. It is entirely possible that forward selection
leads to a different equation from backward selection, although this l1as not
happened in the present example. It is therefore important, especially where
interpretation of the particular form of equation is central to the analysis, that
if there are several simple equations that fit almost equaiiy well, all should be
isolated for consideration and not one chosen somewhat arbitrarily.
Supppse now that a representation, hopefully quite a simple one, has been
obtained for expected response as a function of certain explanatory variables.
What are the principal aspects in using and interpreting such an equation? This
is point (vi) of the list above. There are at least five rather different possibilities.
Firstly, an equation such as that summarized in Table G.4, including the
residual standard deviation, provides a concise description of the data, as
regards the dependence of cost on the other variables. Such a description can
Example G 85
be useful in thinking about the data qualitatively and in comparing different,
somewhat related, sets of data.
A second descriptive use is in the study of the individual cases. The residual
from the fitted model is an index for each power station assessing its cost
relative to what might have been anticipated given the explanatory variables.
Thirdly, the equation can be used for prediction. A new individual has given
(or sometimes predicted) values of the relevant explanatory variables and the
equation, and associated measures of variability, are used to forecast cost,
preferably with confidence limits. In such prediction the main assumption,
in addition to the technical adequacy of the model in the region of explanatory
variables required for prediction, is that any unmeasured variable affecting
response keeps the same statistical relationship with the measured explanatory
variables as obtains in the data. Thus, in particular, if the new individual to be
predicted differs in some way from the reference data, other than is directly or
indirectly accounted for in the explanatory variables, a modification of the
regression predictor is worth consideration. For example, a major technolo-
gical innovation between the data analysed and the individual to be predicted
would call for such modification of the predictor.
Fourthly, the equation may be used to predict for a new individual, or some-
times for one of the original individuals, the consequences of changes in one or
more of the explanatory variables. For example, one might wish to predict
not so much the cost for a new individual as the change in cost for that indivi-
dual as size changes. The relevant regression coefficient predicts that change,
provided that the other explanatory variables are held fixed and that any
important unobserved explanatory variables change appropriately with the
change in size. The prediction of changes in uncontrolled observational
systems, e.g. in the social sciences, needs particularly careful specification of the
changes in explanatory variables envisaged.
Finally, and in some ways most importantly, one may hope to gain insight
into the system under study by careful inspection of which explanatory vari-
ables contribute appreciably to the response and of the signs and magnitudes
of the associated regression coefficients. Thus in the present example, why do
certain variables appear not to contribute appreciably, why is the regression
coefficient on log size appreciably less than one, the value for proportionality,
and so on? As indicated in the previous paragraph, the regression coefficients
estimate changes in response under perturbations of the system whose precise
specification needs care.
The last two applications of regressions need considerable thought, especi-
ally if there is any possibility that an important explanatory variable has been
overlooked.
The analysis. As explained in the preceding section, we take log C, logS, log N,
log T1 and log T 2 ; throughout natural logs are used.
86 Applied statistics
A regression of log C on all ten explanatory variables gives a residual mean
square of0.5680/21 = 0.0270 with 21 degrees of freedom. Elimination of non-
significant variflbles successively one at a time removes BW, log T1 , log T 2 and
PR (Table G.3), leaving six variables and a residual mean square of0.6337/25
= 0.0253 with 25 degrees of freedom; the residual standard deviation is 0.159.
Table G.3. Elimination of variables
No. Variables Residual
variables eliminated
included s.s. d.f.
10 0.5680 21
9 BW 0.5709 22
8 log T1 0.5730 23
7 logT. 0.6165 24
6 PR 0.6337 25
None of the eliminated variables is significant if re-introduced. The estimated
coefficients and, standard errors for the six-variable regression are 'given in
Table G.4. The variable PT, denoting partial turnkey guarantee, has a co-
efficient of -0.2261, with a standard error of 0.1135 (25 d.f.), suggesting that
cost tends to be reduced on average by about 20 per cent for these six plants .
•
Table G.4. Multiple regression: full and reduced models
Variable Regression coefficient
Reduced model Full model
Estimate Standard Estimate Standard
error error
Constant -13.26 3.140 -14.24 4.229
PT -0.2261 0.1135 -0.2243 0.1225
cr 0.1404 0.0604 0.1204 0.0663
logN, -0.0876 0.0415 -0.0802 0.0460
logS 0.7234 0.1188 0.6937 0.1361
D 0.2124 0.0433 0.2092 0,0653
NE 0.2490 0.0741 0.2581 0.0769
logT1 0.0919 0.2440
logT, 0.2855 0.2729
PR -0.0924 0.0773
BW 0.0330 0.1011
Residual st. dev. 0.159 0.164
(25 d.f.) (21 d.f.)
Example G 87
To check whether these six plants and the twenty-six others can be fitted by a
model with common coefficients for each of the variables CT, log N, logS and
log D, we include in turn in the regression the interaction of each variable with
PT. This cannot be done for the variable NE since all six PT plants were con-
structed in the same region. Table G.S summarizes the results. None of the
interaction coefficients is significant. We note that the coefficients of the six
Table G.5. Regressions including interactions with PT
Z = CT Z =log N Z =logS Z=D
Variable Estimate s.e. Estimate s.e. Estimate s.e. Estimate s.e.
Constant -13.23 3.193 -13.26 3.225 -13.09 3.239. -13.22 3.231
PT -0.2429 0.1221 -0.2293 0.8265 -2.188 5.854 -1.529 15.17
CT O.JJll2 0.0652 0.1404 0.0629 0.1400 0.0615 0.1412 0.0624
logN -0.0868 0.0422 -0.0876 0.0423 -0.0868 0.0423 -0.0875 0.0423
logS 0.7229 0.1208 0.7234 0.1219 0.7176 0.1222 0.7222 0.1221
D 0.2121 0.0440 0.2124 0.0444 0.2104 0.0444 0.2120 0.0444
NE 0.2490 0.0754 0.2490 O.o757 0.2484 0.0755 0.2489 0.0757
PTxZ 0.0798 0.1887 0.0014 0.3683 0.2916 0.8700 0.0193 0.2246
common variables in Table G.S remain fairly stable, except for PT which, in
two cases, is estimated very imprecisely. A model with common coefficients as
given in Table G.4 seems reasonable. With this model the predicted cost in-
creases with size, although less rapidly than proportionally to size, is further
increased if a cooling tower is used or if constructed in the NE region, but
decreases with experience of architect-engineer.
Table G.6. Comparison of observed and fitted values based on six-variable
regression of Table G .4 fitted to log C
Observed Fitted Residual Observed Fitted Residual
6.131 6.051 0.080 6.568 6.379 0.189
6.116 6.225 -0.109 5.669 5.891 -0.222
6.094 6.225 -0.131 6.781 6.492 0.289
6.481 6.398 0.083 6.196 6.230 -0.034
6.465 6.398 0.067 6.342 6.178 0.164
5.845 5.916 -0.131 6.501 6.651 -0.150 '
5.607 5.934 -0.327 6.432 6.249 0.183
5.760 5.704 0.056 6.411 6.384 0.027
6.125 5.987 0.138 6.160 6.129 0.031
6.537 6.411 0.126 6.547 6.797 -0.250
5.860 5.789 0.071 5.335 5.401 -0.066
5.998 6.262 -0.264 5.665 5.606 0.059
6.021 5.891 0.130 5.652 5.621 0.031
6.206 6.241 -0.035 5.636 5.621 0.015
5.977 6.016 -0.039 5.382 5.401 -0.019
6.048 5.992 0.056 5.601 5.621 -0.020
0.3 X
0.2 xx
X
0.1
x*
X
X X ~~
X X X
X
0
;;; X X
~ X~
1::!
;;; X
~ -0.1
~
X
X
-0.2.
X
X X
-0.3
X
57 50 59 70 71
D-
Fig. G. I. Residuals of log C from six-variable model versus D, date construction permit
issued.
0.2
X X
X
X X X
0.1
X
X
X
X
X
X
X
*
.."
X
0 X
X X
1::!
·;;; X
X
*
~ -0.1
X
X X
X
X
X
X
6.2 6.6
logs-
Fig. G.2. Residuals oflog C from six-variable model ve~sus logS, power plant net capacity.
0.3 X
0.2-
X X
X
X X X
0.1
X X X X
X
~
0
~ X X
iii X X )!!<
::l X
-o
·;;; X
~ -0.1 .x
X X X
-0.2
X
X X
-0.3
X
5.5 6.0 6.5
Fitted value--
Fig. G.3. Residuals of log C from six-variable model versus .fitted values.
0.3 X
0.2 X X
X
X xX
0.1
~x>«
X~
0
,·;;;,
n;
xxxf
X
m
a: -0.1 X
XX
X
-0.2
X
X X
-0.3
X
-2.0 -1.0 0 1.0 2.0
Normal order statistic--
Fig, G.4. Residuals of log C from six-variable model versus normal order statistics.
90 Applied statistics
Fitted values and residuals are given in Table G.6. The residuals give no
evidence of o,utliers or of any systematic departure from the assumed model;
this can be [Link] by plotting in the usual ways, against the explanatory
variables, for example against D (Fig. G.l) and log S (Fig. G.2), against
fitted values (Fig. G.3) and normal order statistics (Fig. G.4).
The estimated standard error of predicted log C for a new power plant,
provided conditions are fairly close to the average of the observed 32 plants, is
approximately 0.159 (1 + 1/32)112 = 0.161 with 25 degrees of freedom. Thus
there is a 95 per cent chance that the actual cost for the new plant will lie within
about ±39 per cent of the predicted cost.
Related reference. Draper and Smith (1981) give a comprehensive account of
multiple regression and discuss (Chapter 3) the examination of residuals.