DIPLOMA IN MATHEMATICAL STATISTICS
1995{96
Applied Projects
(as summarised by their authors)
1
S.G. Christodes Wolfson Cointegration and analysis of economic time-series
(Department of Land Economy)
A.G. Dales Trinity Hall Changes to the compensation recovery unit
and their implications for insurance settlements
(Price Waterhouse)
G. Harper Churchill Modelling of data from a quality control dataset
(London International Group)
T.P. Harris St Catharine's The eect of climate upon the isotopic
composition of tree rings
(Anglia Polytechnic University)
M. Kanaan St Edmund's Adult mortality in English parishes 1540{1840
(History of Population Unit)
J.H.N. Lees Newnham Clinical concordance in siblings with
multiple sclerosis
(MRC Biostatistics Unit)
J.D. Logan Trinity Hall A structural time series approach to forecasting
the space-time incidence of infectious diseases:
post-war measles elimination programmes in the
United States and Iceland
(Department of Geography)
N.M.B. Lohse Trinity Hall Foreign Exchange Risk: Determining an estimate
of the covariance structure of foreign exchange
rates for trading risk control
(NatWest Markets)
P. Panayiotou St John's Analysis of non-linear growth curves in their
application to plant epidemiology
(Department of Plant Sciences)
K. Poovanendiraraj Churchill Dietary and L-carnitine eects on the growth
of small-for-gestational age babies
(MRC Dunn Nutrition Unit)
Y.-C. Teh Queens' Statistical analysis of multimedia trac in
modern communication networks
(British Telecom)
T.F.J.A. Undreiner Trinity Hall Cluster analysis of rms according to
R & D strategies
(Judge Institute of Management Studies)
2
COINTEGRATION AND ANALYSIS OF ECONOMIC TIME-SERIES
For each of the three countries, France, Italy and West Germany, we were given data on the output
of the country (GVA) and on the employment in that country from the years 1975 to 1993. The
data was yearly and therefore consisted of just 19 points per series. The aim of the project is
to investigate the nature of these series, and to use cointegration to attempt to nd a statistical
model relating the GVA and employment series. The series were assumed to be non-stationary
and were tested for cointegration, since omission to do so would lead to `spurious' results with no
practical implications.
Tests were performed between the GVA and employment series for each individual country (the
`single equation' case), and between the GVA and employment series of several countries simulta-
neously (the `multiple equation' case). The Dickey{Fuller and the Phillips{Perron techniques are
two of the methods used in this project to test for non-stationarity. The tests used to check for
cointegration in single equations included the Dickey{Fuller test, the Dynamic Regression test,
and the Kremers test. The Johansen procedure was used to examine cointegration in the multiple
equation case. All the methods were implemented using the statistical package S-PLUS.
It was of course clear from the outset that the number of available data points (19) for each
of our time series is far too small to expect any reasonable degree of condence in the results
and conclusions of this study. Many of the tests used are t-type tests which are known to be
badly aected by small data sets, since in general the standard errors for estimated coecients
are high and most critical values are from asymptotic distributions and based on Monte Carlo
methods. Thus we found several con
icting results from the various tests. Nevertheless, using the
MLE approach of Johansen (which our limited experience suggests is more robust for small data
sets), we have concluded that there does indeed exist a relationship between a given country's
employment and the rates of growth of its economy (GVA) and the economies of its major trading
partners. This is of course not too surprising, but our derived relationship is more complex than
the hypothesised relationship expressed by Verdoorn's Law [Kaldor 1966].
CHANGES TO THE COMPENSATION RECOVERY UNIT AND THEIR IMPLI-
CATIONS FOR INSURANCE SETTLEMENTS
This report, `Changes to the Compensation Recovery Unit and their Implications for Insurance
Settlements (CCRU) concerns data collected by Price Waterhouse, the large accountancy and
actuarial rm, about insurance claims and the amount of government benets that are paid to
recipients of insurance settlements. The report expands upon `The Compliance Cost Assessment,
Compensation Recovery Scheme' (January 1996), a report written by Price Waterhouse at the
request of the Department of Social Security. This report, which is referred to as PW, is included
in CCRU as Appendix C.
The Compensation Recovery Unit (CRU) is the part of the Department of Social Security re-
sponsible for reclaiming government benets paid to individuals who are awaiting the outcome
of insurance claims. At present, these benets are repaid by any claimant who is awarded a
settlement of greater than $2,500. This sum is known as the Small Payments Limit or SPL. The
government is considering removing the SPL. This would mean the CRU would be eligible to
reclaim all benets paid to the recipient of an insurance settlement.
3
CCRU attempts to nd a model for the distribution function of settlements found in the data
set. It comments on the problems of nding such a model. These problems include: the problem
of clustering, the problems associated with a very heavily-tailed distribution, and the possible
dierences in the distribution functions for settlements according to liability type.
The t of the models are tested using Kolmogorov{Smirnov and chi-squared tests. Further, a
modication of the chi-squared approach is developed, which seeks to nd distribution models for
the settlement data that take into account the possible eect on insurance settlements resulting
from the abolition of the SPL.
Finally, in Chapter 6, CCRU discusses the results presented in PW concerning the amount of
benets attached to cases that settle below the SPL. This is benet, which under the present
system is paid to claimants but not recovered by the CRU. CCRU suggests that insucient data
on such cases have been gathered accurately to predict the additional income that the CRU would
receive if the SPL were to be abolished.
MODELLING OF DATA FROM A QUALITY CONTROL DATASET
Quality Control procedures commonly demand that some percentage of any batch of items be
tested in order to ensure the suitability of the product for consumption. In a standard test
procedure for condoms, items are mechanically in
ated until they burst. The volume and pressure
at bursting point are recorded. If either of these values falls below some threshold level, the item
is deemed defective. If there are too many defective items in a sample, then the whole batch is
rejected.
This sampling procedure eectively recodes the volume and pressure data as a simple pass/fail
binary response. Might it be possible to use the actual values of pressure and volume measured
to make stronger inferences based on a smaller sample size? It was thought that it might be if the
underlying distribution of volume and pressure was better understood. Attempts to t standard
distributions to this kind of data set had failed in the past. The aim of this project was to t a
distribution to a Burst Pressure/Burst Volume dataset.
In the project, I consider the Bivariate Gamma Distribution as a possible model, before moving
on to explore some non-parametric density estimative techniques for tting. In particular, Kernel
and Adaptive Kernel Density Estimates are considered. The ts obtained using the dierent
methods are compared, and issues of potential overtting are discussed.
The data used were collated by the London International Group , who are based at the Science
Park in Cambridge. All analyses were carried out using S-PLUS. The nal write-up was produced
using LaTEX.
THE EFFECT OF CLIMATE UPON THE ISOTOPIC COMPOSITION OF TREE
RINGS
The eect of climate on tree growth leads to the expectation that tree rings provide a record
of past climates. In recent years, much of the search for a link has centred on the chemical
composition of the cellulose in each individual tree ring. This report considers data consisting
of the isotopic ratios of carbon, hydrogen and oxygen from tree rings in a single oak tree grown
4
in the Wennington area near Cambridge. The climate of the surrounding area is also known in
terms of various meteorological variables. Assigning dates to the tree rings leads to a nal data
set comprising several time series known over comparable time periods. The auto-dependence of
the isotopic time series is studied before extending the analysis to consider cross-dependence with
the climatic series. Exploratory data analysis provides a general overview of the inter-relatedness
of the data.
A key issue is how most eectively to summarise the climatic variables to ensure the best com-
parisons. Comparisons are made both by reducing the dimension of the problem via principal
components analysis, and by isolating particular variables for analysis, the latter method allowing
for greater interpretability. A related conclusion that comes from the analysis is the importance
of emphasising the climate prevalent at the July and August stage of the growing season. Wher-
ever possible, particular parametric models are considered and evaluated. Indeed, a feature of
the analysis is the allowance for development of the model framework in particular directions;
the consideration of long range dependence models, for example, as an extension of the simpler
ARIMA analysis, or the analysis of nonlinear behaviour which looks beyond the usual linear
framework. The exploratory nature of the whole investigation leads to conclusions that are more
qualitative than quantitative, demonstrating the need for further studies before there can be
eective forecasting.
S-PLUS was used throughout for all statistical analysis and graphical output. Functions from
various S-PLUS libraries were called upon in conjunction with several functions dened specically
for the analysis in hand. The entire report was created using the typesetting system LaTEX.
ADULT MORTALITY IN ENGLISH PARISHES 1540{1840
The aim of this study initially was to investigate the factors that in
uenced adult mortality
during the period 1520{1840. However, there was a problem in the composition of the data: not
all parishes were observed throughout this period. Hence, we focused on the following two periods
1675{1749 and 1750{1789. In the rst period all the parishes were observed while in the second
some parishes were missing.
Adult mortality was studied to see whether any signicant dierences between these two periods
exist. Factors such as gender, major occupation within a parish, parish environment, marriage
interval and numbers of born and dead children of an individual were studied for their impact on
the survival of an individual during the specied periods.
Generally, no signicant dierences were found between the two periods. However, gender dier-
ences were detected females' advantage in survival was observed after they reach the menopause
period when no more children are conceived. In addition, the ratio of the numbers of dead to
born children aected the survival after thirty but not after fty. Moreover, retail and handicraft
occupations proved to be more dangerous than other major occupations within parishes, whereas
urban environments were more hazardous than other parish environments. Finally, the marriage
interval eect was signicant during the rst period but not during the second period. Hence the
eect of these covariates on survival can be linked to the health and welfare of the couple and the
environment.
5
The model tting was accomplished using Cox Proportional Hazards, a semi-parametric approach
to t the data. All the statistical analysis was done using S-PLUS routines for proportional
hazards. The report was written in LaTEX with the gures imported from S-PLUS and inserted
into the LaTEX document using epsf .
The report is divided as follows. Chapter 1 discusses the background of adult mortality and
dierent demographic approaches to assess the parish registers. Chapter 2 introduces the deni-
tion of the variables used and a summary of the data. Chapter 3 introduces the mathematical
background for Cox Proportional Hazards. Variable selection is carried out in Chapter 4. In
Chapters 5 and 6 we analyse the data and t the models for the two periods under investigation.
Chapter 7 summarises the main results and provides suggestions for further work. It is followed
by appendices which include some of the functions used for the statistical analysis.
CLINICAL CONCORDANCE IN SIBLINGS WITH MULTIPLE SCLEROSIS
Multiple sclerosis is a chronic disease of the nervous system, of unknown cause. It aects dierent
parts of the brain and spinal cord, resulting in a myriad of clinical symptoms and varying degrees
of disability.
Sib-pair studies are useful for investigating such diseases of unknown cause and complex inheri-
tance, as they can be used to estimate the relative contribution of inheritance and the environment,
by comparison of age and year of rst onset in siblings. For instance, familial clustering in age of
onset would suggest that inherited factors in
uence onset, whereas familial clustering in year of
onset suggests that onset of disease follows exposure to some external environmental factor.
This study seeks to determine this relative contribution and investigate concordance of other
clinical features in siblings with multiple sclerosis, using data from 177 families in which there
are two or more aected siblings.
Intra-class correlation coecients (for metric variables) and kappa statistics (for categorical vari-
ables) were used to measure the tendency of aected siblings to be more alike than unrelated
individuals. Because of underlying biases arising as a result of the cross-sectional design of the
study, the signicance of these results was determined using an empirical test procedure in which
control bootstrap samples of age-matched `families' of unrelated individuals were generated.
This study demonstrated that while there was evidence of a small genetic contribution to disease
aetiology, there was a larger contribution from exposure to environmental factors. Also, although
this genetic contribution is small, there was evidence to suggest that clinical features of familial
disease are partly due to inherited factors.
Most of the statistical analysis in this report was carried out using the statistical package S-PLUS.
A function was written in C to generate bootstrap samples and the empirical testing procedure
was then implemented by interfacing between S-PLUS and C. Microsoft Word version 2.0 was
used to produce the report.
6
A STRUCTURAL TIME SERIES APPROACH TO FORECASTING THE SPACE-
TIME INCIDENCE OF INFECTIOUS DISEASES: POST-WAR MEASLES ELIM-
INATION PROGRAMMES IN THE UNITED STATES AND ICELAND
In 1979 the global eradication of the disease smallpox was achieved. This was a major success in
the epidemiological world and has naturally prompted questions as to whether other infectious
diseases can also be eradicated. One of the diseases the World Health Organisation has targeted
for eradication by the next millennium is measles and the attention of this project was restricted
to the analysis of this one disease. The main vehicle for eradication is vaccination. However,
as was clearly demonstrated by the eradication of smallpox and by attempts made to eradicate
measles in the United States, the cost of such vaccination policies is very high. These costs only
disappear if global eradication is achieved and this makes eradication an expensive business for
which success will only be possible if ecient vaccination strategies are developed and put into
practice worldwide. To do this, models need to be devised that can forecast accurately the space-
time incidence of the disease then, using these models, spatially targeted vaccination programmes
must be set up to replace the existing `blanket coverage' policy.
Therefore, in this project, we returned to the age-old problem of forecasting. The data considered
were the quarterly number of measles cases reported in each of the six administrative regions of
Iceland, for the period 1945 to 1985. We rst examined the links between disease endemicity and
population size to decide how the critical community size is aected by vaccination. We then
looked at the geographical corridors of measles spread within Iceland and attempted to develop
a new approach to epidemic prediction using structural time series models (STSM) rather than
the more traditional Susceptible{Infected{Removed (SIR) or ARIMA models. The key to this
method is the straightforward transformation of the STSMs into state-space form, for which
estimation can be routinely carried out in the time domain using the Kalman lter. This analysis
was carried out using the STAMP computing package designed specically for STSMs. An added
complication to the problem was the disruption to the time series caused by the introduction of
mass vaccinations against the disease in the mid 1960s. One approach to such events is to use
intervention analysis which, in the structural time series approach, allows the event to be modelled
directly by the use of `dummy' explanatory variables. In addition to the standard intervention
variables, other variables based on work by Griths in the 1970s were tested, and these proved
to be very successful in accurately modelling the intervention eect.
However, it was observed that the STSMs exhibited similar problems to the other standard ap-
proaches by often failing to predict the start of an epidemic and by either over- or under-estimating
the magnitudes of the epidemics. A feature of some of the tted series was that, although they
provided an accurate match to the observed series, they were oset from the observed series by
a lag of one quarter. Since the measles virus was known to follow set geographical corridors of
which the majority begin in Reykjavik, it was decided to create explanatory variables for the
other regions based on the Reykjavik data. Lagged, clipped and trigger variables were tested and
it was observed that some of these variables did indeed reduce the oset exhibited by the tted
series.
7
FOREIGN EXCHANGE RISK: DETERMINING AN ESTIMATE OF THE CO-
VARIANCE STRUCTURE OF FOREIGN EXCHANGE RATES FOR TRADING
RISK CONTROL
The day-to-day
uctuations in foreign exchange rates are problematic for banks wishing to limit
the foreign exchange traders' losses. Consequently, within a bank, it is necessary to check that
the traders do not take too much risk. In this project, the subject of interest is the control of risk
within each bank.
This study examines several approaches to the calculation of the covariance matrix estimate for
the distribution of daily changes in foreign exchange rates. These changes are assumed to be
Log Normal with mean 0. I use this covariance matrix to estimate the risk of foreign exchange
portfolios.
The results are based on daily foreign exchange data over the period October 1993{September
1995 on both the United States (US) and United Kingdom (UK) markets. Calculations are
computed using the statistical computing package S-PLUS.
I compare three methods to compute the covariance matrix estimate:
the maximum likelihood estimation using either 6 or 18 months of data,
the exponential weighting method,
the percentile method, based on the sample 5 percentile of changes in the value of the
rates.
Using this comparative analysis, I conclude that the exponential weighting method leads to the
most precise estimate of the risk. The expression for the covariance matrix estimate with this
method is:
X
n
^ = (1 )i xi xTi
i=0
with the chosen decay factor = 0:989 and xi the observation on day i. This observation method
is then checked through the comparison of estimated volatilities with `implied volatilities' and is
tested over the period October{December 1995.
The eciency of the risk estimation appeared to be limited by several factors:
The covariance structure of foreign exchange is unstable and so it is dicult to get a precise
prediction even for the immediate future.
The dierence of liquidity between currencies (i.e. currencies traded more or less frequently)
leads to incorrect estimates of the corresponding cross-correlation coecients. The results
were very sensitive to this measurement eect.
The historical estimation method of this study ignores economic and political factors that
can signicantly in
uence the behaviour of the rates.
Each of these limiting factors reduces the accuracy of the risk estimation. However, such analysis
rather than producing the `correct' estimate, may serve to achieve two particularly signicant
functions:
{ to alert management in the case of large risks
{ to encourage the traders to reduce their risk.
8
ANALYSIS OF NON-LINEAR GROWTH CURVES IN THEIR APPLICATION
TO PLANT EPIDEMIOLOGY
The aim of the project is to analyse the data collected by T. O'Neill for three distinct horticultural
experiments. The `individuals' under treatment are Fusarium inoculated plants and the aim is
to nd a good t in order to be able to compare the treatments under, of course, the specied
environmental conditions. For each experiment a number of treatments is applied to identical
samples of inoculated plants under controlled conditions. There have been, historically, two main
approaches in the analysis of this kind of data. The rst one considers the use of polynomials in
tting the data whilst the second, the use of non-linear `growth' curves. In the current analysis
the second is judged superior and therefore is the one pursued. The choice of models for this
approach to be biologically consistent is determined by the following two criteria:
1. The parameters of these curves should have some kind of biological interpretation.
2. They should be derived from dierential equations which describe the rate of the underlying
process through time.
The analysis was carried out using growth curves of logistic type and as a result the statistical
method was non-linear regression (growth curves being non-linear).
The model assumptions are that we are in a situation of i.i.d. normal random variables and the
method used for the estimation of the parameters is that of maximum likelihood estimation;
which is precisely the method of least squares, under our assumption of normality.
The goodness-of-t of the models is carried out via examination of residual and probability plots.
As we will see in the analysis, the independence assumption fails to be satised in most cases.
This should not worry us since our main interests are to get good estimators for the parameters
of the model and to test for common parameters among sets of treatments. Even though the
residuals are correlated, this will not aect the parameter estimates as point-estimators of the
true unknown parameters and neither will it aect the reduction in deviance when testing for
common parameters (Seber and Wild 1989). Hence the analysis of the residuals via the dierent
kinds of residual plots is not central and is given for mathematical completeness; it is further
discussed in Chapter 6, with the appropriate ACF and PACF plots in Appendix B.
Finally, examination of the parameter loadings (being the counterpart of leverages in linear re-
gression models) indicates which parts of the range are more in
uential in determining each
parameter. These can be used for improving future experiments since in all cases, for the models
considered, they reveal the same type of information. Chapter 2 describes the choice and theory
of the models considered. Chapters 3, 4 and 5 describe the analysis of the three dierent data
sets and inferences made. Chapter 6 concludes with an overview of the results and suggestions
for further work. Appendix A shows input and output from the tests carried out. Appendix B,
as mentioned above, shows the ACF and PACF plots for the residual series obtained from the
ts, as a side comment to the correlation in the error structure.
The tting and comparison of the curves was carried out using the Maximum Likelihood Program
(MLP) (Ross 1987), and the graphs using S-PLUS.
9
DIETARY AND L-CARNITINE EFFECTS ON THE GROWTH OF SMALL-FOR-
GESTATIONAL AGE BABIES
The aim of this project was to examine whether there is any signicant dierence in growth pattern
between small-for-gestational age (SGA) babies from the two diet groups, and the two dierent
l-carnitine groups respectively. A phase II, double blind, placebo controlled parallel study was
carried out on SGA babies. The babies were randomised into active and placebo groups stratied
for sex, and breast versus bottle feeding. Measurements of the growth parameters (e.g. weight)
on the 54 infants were taken at birth, 2{3 days, and 2, 5, 12, 26, 39 and 52 weeks respectively.
First, some plotting of this longitudinal data set were done, to understand the structure of the
data, and to help us ask the correct questions related to the growth of SGA babies in the dierent
groups.
Therefore, to answer the above mentioned questions, the dimension of the data matrix had to be
reduced. This was achieved by a method known as the Two Stage Method of Summary Measure.
Essentially the method reduces the data matrix into two summary measures, which picks out
certain characteristics of the growth curves. These characteristics could be regression coecients.
Fractional polynomials of degree two were the class of parametric regression models chosen to
be considered due to the fact that historically these were the model that have been used to t
growth data (Royston and Altman 1994). After obtaining the summary measure(s), univariate
statistical methods were used to check whether there was any signicant dierence in the growth
parameters, for example, among breast-fed and bottle-fed babies.
The statistical analysis and plots were done with S-PLUS 3.3. The project was typed using
LaTEX.
STATISTICAL ANALYSIS OF MULTIMEDIA TRAFFIC IN MODERN COMMU-
NICATION NETWORKS
Modern telecommunication networks are capable of using common resources to carry a wide range
of trac streams from dierent sources. To guarantee quality of service for a given connection,
the network must be able to assign and regulate correctly the bandwidth allocated to the service.
Ecient design, management and control of such communication systems therefore depends on
the statistical characteristics of the trac. In the past it has been dicult to assess the validity
and appropriateness of traditional stochastic models for high-speed network trac due to the lack
of empirical data. However, large datasets of trac measurements from operational networks are
now available. A recent development is the notion of eective bandwidth which provides a measure
of resource usage taking into account the characteristics and requirements of dierent sources.
This can be used to generate a graphical descriptor, the eective bandwidth surface , of the trac
stream, providing a summary of those trac characteristics which are important when considering
the statistical sharing of common resources.
Empirical trac data collected from an Ethernet local area network is considered as a typical
example of modern network trac. The aim is to demonstrate the use of both traditional statis-
tical methods as well as eective bandwidth surfaces to describe statistical characteristics of the
real trac stream.
Examination of the long-term correlations of the trac arrival process provides substantial evi-
dence that Ethernet LAN trac is statistically self-similar . This result questions the validity of
10
predicted network performance based on traditional Poisson process models. Investigation of the
short-term correlations, and in particular the lag 1 scatterplot reveals remarkable time-dependent
and packet-length-dependent structure in the dataset through the presence of striking diagonal
lines. A recent data-plotting technique, generating textured dot strips , is also used to reveal an
interesting on-o periodicity of 20{30 seconds in the trac arrival process.
A direct computational method is developed for calculating the eective bandwidth surface,
(s; t), of real broadband trac traces. Comparison of (s; t) for the actual Ethernet trace
with that for a tted Fractional Brownian motion (FBM) model leads to the conclusion that
FBM models the eective bandwidth well only for certain values of the space scale parameter, s,
and time scale, t. The availability of dierent Ethernet traces provides a comparison of eective
bandwidth surfaces for datasets resulting from the use of dierent applications and varying de-
grees of Ethernet utilisation. As a contrast to the Ethernet trac, an MPEG video dataset of the
`Star Wars' movie is examined and the (s; t) surface for this trace compared to a strict periodic
source model.
Much research in high-speed networks is currently focussed on ATM (Asynchronous Transfer
Mode) networks. The process of transferring Ethernet trac onto an ATM network is emulated
and the required buer size determined for a range of gateway transfer bit rates. The result
of this process on the eective bandwidth surface is also investigated, and for certain values of
the time-scale parameter, and in the presence of a large amount of buering, this process can
dramatically reduce the eective bandwidth of the trac stream.
The Ethernet dataset was obtained from the Bellcore Morristown Research and Engineering
Center. The statistical package S-PLUS was used for statistical analysis and the generation of
graphs. Numerically intensive calculations were performed using the C language. Additional
gures were generated using the xg drawing package and the report was produced using the
mathematical typesetting language, LaTEX.
CLUSTER ANALYSIS OF FIRMS ACCORDING TO R & D STRATEGIES
A number of rms in the US, the UK, Germany and Japan completed a questionnaire concerning
the scope and eectiveness of R & D strategies in biotechnology. This questionnaire was devised
and the study launched by Dr Susan Bartholomew of the Judge Institute of Management Studies
at Cambridge University. Cluster analysis techniques were used to determine and relate groupings
amongst the rms. A further study would investigate the association of the groups with factors
describing rms.
The strategy of a rm is dened as the partition of its biotechnology R & D expenditure into ve
options which refer to dierent types of alliances. Hence rms described by these ve variables
were grouped in such a way that those within a group are `close' to one another. The result was
a partition into eighteen groups of rms with similar R & D alliances strategies. This partition
showed that it is also relevant to consider partial strategies taking into account only four options.
Calculating this explicitly, the result was a partition into fourteen groups of rms.
There are 229 rms for which a R & D strategy is dened. Since the data set is rather large,
only hierarchical algorithms and Hartigan's K-means algorithm were eectively implemented.
Both partitions were obtained by cutting the dendrogram structure resulting from a hierarchical
complete linkage of the rms.
11
When a clustering algorithm is implemented on any set of objects, it is always possible to nd
a partition of the objects into any given number of clusters. Therefore, the project develops
techniques to check the validity of the resulting partitions. Firstly, the distortion of the original
dissimilarities between objects by the dissimilarities yielded by the dendrogram was measured
using the cophenetic correlation and the Stress. Secondly, criteria for choosing the `correct'
number of clusters were continually used. Thirdly, principal components analysis provided useful
visual assessments. Programmes for carrying out this work were written. In particular a set of
programmes for the computation of the monotone regression of a sequence on another sequence,
according to the algorithm proposed by Kruskal (1964), was devised. A selection of programmes
for the computation of the cophenetic coecients read o a dendrogram, as dened by Sokal &
Sneath (1963), were also derived.
The statistical package S-PLUS was used throughout for all statistical analysis, including pro-
grammes and graphical output. The entire report was created using the typesetting system
LaTEX.
12