"Holderbank" - Cement Course 2000
Materials Technology / B02 - MT II / C13 - Statistics
C13 - Statistics
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:13 PM Page 1
Query:
"Holderbank" - Cement Course 2000
Materials Technology / B02 - MT II / C13 - Statistics / Statistics
Statistics
1. INTRODUCTION
2. SIMPLE DATA DESCRIPTION
1. Graphical Representation
2. Statistical measures4
1. The p-quantile
2. The Box Plot
3. Measures of location
4. Measures of Variability
5. Statistical program packages
6. Interpretation of the standard deviation
7. Outliers
3. THE NORMAL DISTRIBUTION (ND)
4. CONFIDENCE LIMITS
1. Confidence limits for the mean
2. Confidence limits for the median
3. Confidence limits for the standard deviation
4. Other methods for the construction of confidence limits
5. STANDARD TESTS
1. General Test Idea
2. Test Procedures
1. z-Test
2. Sign Test
3. Signed-rank test
4. Wilcoxon-Test
5. t-Test
6. Median-Tests
7. X2-Test
8. F-Test
3. Sample Size Determination
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:14 PM Page 2
Query:
"Holderbank" - Cement Course 2000
6. DATA PRESENTATION AND INTERPRETATION
1. Intelligible Presentation
2. Interpretation
3. Data Interpretation related to Problems in Cement Application
4. Control Charts
5. Comparative representation
7. CORRELATION AND REGRESSION
1. Correlation coefficient
2. Linear Regression
1. Regression line with slope 0:
2. Regression line with y-intercept 0
3. Intercept = 0. Slope = 0
4. Comparison of models
5. Standard deviation of the estimates
6. Coefficient of determination
7. Transformations before Regression Analysis
8. Multiple and non-linear regression
8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN
1. The Five Phases of a Statistical Investigation
2. Sample Surveys and Experiments
3. Fundamental Principles in Statistical Investigations
1. Experiments
2. Sample Surveys
9. OUTLOOK
1. Time series and growth curves analysis
2. Categorical and Qualitative Data Analysis
3. Experimental Designs and ANOVA
4. Multivariate Methods
5. Nonparametric Methods
6. Bootstrap and Jack-knife Methods
7. Simulation and Monte Carlo Method
8. General Literature
10. STATISTICAL PROGRAM PACKAGES
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:14 PM Page 3
Query:
"Holderbank" - Cement Course 2000
PREFACE
For appropriate process and quality control in the cement and concrete industry, a large number of
data are derived. Optimum benefit is, however, only achieved if these are adequately processed and
interpreted. Statistics is one of the important means to make best use of the data be it by application of
numerical methods and/or by graphical representation.
The present handbook describes the relevant basic definitions, formulae and applications of statistical
methods which are useful in the cement industry. Emphasis is put on adequate data description and
graphical representation to ensure reasonable processing and interpretation of statistical data. Most of
the described procedures are illustrated with practical examples.
Chapters 1 and 2 are concerned with the basic ideas of statistics, the rules for the representation of a
given set of data (graphical representation, numerical measures) and the treatment of outliers.
Chapters 3 to 8 present some statistical methods, useful for decision making, experimentation and
process control.
Of special practical significance is the Application Section (chapter 5.2), which includes a procedure
manual with a general check list and a collection of important test procedures. Chapter 9 gives an
outwork to more sophisticated statistics.
Appendix I contains a selection of practical examples to illustrate applicability and interpretation of the
demonstrated methods. A useful work sheet to construct frequency tables and to check the data for
normality is given in Appendix II. Further Appendices contain: the required statistical tables for the
determination of confidence limits and the application of test procedures, a list of recommended
literature and an index of examples used in the text.
A subject index (English, German, French) of statistical terms is provided in Appendix VI.
The copyright for this documentation is reserved by "Holderbank" Management and Consulting Ltd.
The right to reproduce it entirely or in part in any form or by any means is subject to the authorisation
of "Holderbank" Management and Consulting Ltd.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 1. INTRODUCTION
1. INTRODUCTION
Statistics is concerned with methods for collecting, organising, summarising, presenting and analysing
data, obtained by measurement, counting or enumeration.
With descriptive statistics a given set of observations is summarised or presented to get a quick survey
of the corresponding phenomena.
In a more sophisticated analysis of representative samples statistical inference (inductive statistics)
allows conclusions to be drawn about the entire population.
A sample is considered to be representative only if it is drawn from the population at random. Such a
group of n observations is called a sample of size n.
Some main topics of interest in statistical analysis are:
a) Location of the data:
Where are the observations located on the numerical scale? This question leads to the use of
central values as the mean or the median.
b) Variability or dispersion:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:14 PM Page 4
Query:
"Holderbank" - Cement Course 2000
Problems concerning the degree to which data tend to spread about an average value.
c) Correlation:
Degree of dependence between paired measurements, e.g.: Is there a real dependence of mortar
strength on the alkali content in clinker?
d) Regression:
Fitting lines and curves to express the relationship between variables in a mathematical form
especially used for prediction and calibration.
e) Splitting variability:
Looking for the importance of several causes to the variability of observations. Variability arises due
to different components of random errors. Special experiments must be designed to split these
components.
The use of statistical methods and the interpretation of statistical results requires a certain
comprehension of variability. If ten pieces of coal from a delivery are analysed as to their water
content, the results are not identical with the water content of the full quantity. Rather we have ten
different results in a certain range. This variability of data arises not only from the fact that the pieces
really have different water contents, but also because the results are influenced by three different types
of errors which may occur in every set of observations:
Random errors cannot be avoided. They are due to imprecise measurement, rounding of data,
environmental effects and not identically repeatable preparation. The amount of random errors may be
expressed by statistical measures of variability.
Systematic errors lead to a bias in location, but not necessarily in dispersion. A bias in dispersion is
obtained if several systematic errors are mixed.
Example: Every laboratory assistant has usually his own systematic error. This error may be relevant or
not and it may change with time. But we always expect a greater variability of the observations if more
than one person have performed the measurements. On the other hand, we recognise in this example
that a small variability does not necessarily lead to a better average value. Environmental effects may
be systematic too (e.g. air humidity).
Gross errors are "wrong" values in the set of observations. Experience shows that 5% to 10% of gross
errors have to be expected in a data set. Reasons may be: wrong reading of scales, errors in copying
data, data not legible, miscalculations, gross error in measurement. Gross errors have a considerable
influence on statistical results.
Careful measurement and data handling is, therefore, important. Data have to be inspected for outliers
before a statistical analysis is performed.
Conclusion: The deviation of the single results from the true mean of the full quantity originate from
real differences between the samples on one hand, and from the occurrence of several types of errors
on the other.
It is often stated that "everything can be proved with statistics", this reasoning clearly is wrong; false or
often only misunderstood statistics originate from insufficient representations, application of wrong
procedures or assumptions, or misleading interpretation of the results. Especially graphical
representations can easily be manipulated. It is therefore essential that the applicant of statistical
techniques knows what he can and what he cannot do! An amusing booklet dealing with such
statistical "lies" is "How to Tell the Liars from the Statisticians".
In our days of growing computer use, a lot of powerful (and sometimes less powerful or even poor)
statistical software packages are available, leading to extensive use of statistics procedures by
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:14 PM Page 5
Query:
"Holderbank" - Cement Course 2000
nonstatisticians or people not having enough statistical background. These packages manage nearly
any instruction without being able to decide whether the statistical procedure is appropriate. This bears
a great danger of use of statistics by statistical amateurs or ignorants. It is therefore absolutely
necessary that the user of statistics has a solid statistical education.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION
2. SIMPLE DATA DESCRIPTION
In this section some descriptive methods are presented to get a quick survey of the data. The methods
are illustrated with an example of concrete strength.
Example 1
The following data represent the compressive strength of 90 cubes (20 x 20 x 20 cm) of concrete. The
data listed in chronological order of measurement is called set of observations.
35.8 39.2 36.8 32.4 30.7 30.8 23.5 22.8 23.7 31.7 34.6 27.6 29.9 28.4 29.3
33.0 37.6 38.1 33.3 38.9 37.1 33.3 33.4 36.4 44.3 48.9 40.1 43.4 35.4 36.6
32.8 34.1 37.4 27.9 30.2 32.0 45.3 45.8 41.0 26.1 27.9 24.4 35.3 34.5 36.1
30.1 40.2 37.9 25.0 23.0 27.8 33.5 34.2 30.0 29.0 35.2 35.8 23.9 34.9 31.5
35.9 39.7 39.4 32.4 33.6 35.2 32.8 30.2 31.6 28.5 28.5 30.3 31.4 31.8 35.5
27.1 24.5 20.9 24.6 27.2 31.7 32.2 38.6 32.8 37.8 36.8 35.3 41.9 34.4 35.5
Generally a set of n observations is denoted by x1, x2, ..., xn where the index j of xj corresponds to the
number of the observation in the set.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.1 Graphical Representation
2.1 Graphical Representation
Although the set of observations gives the complete information about the measurements, it is little
informative for the reader. A better survey is obtained by grouping the data in classes of equal length,
presented in a frequency table with tally.
Rules for classification:
1) The mid points should be impressive values with a few number of digits.
2) Choose 8 to 20 classes of equal length (approx. n; classes)
3) Class boundaries should not coincide with observed values (if possible)
mid-point class tally absolute relative
frequency frequency
20.0 18.75 - 21.25 I 1 0.011
22.5 21.25 - 23.75 llll 4 0.044
25.0 23.75 - 26.25 llll l 6 0.067
27.5 26.25 - 28.75 llll llll 9 0.100
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:14 PM Page 6
Query:
"Holderbank" - Cement Course 2000
30.0 28.75 - 31.25 llll llll 10 0.111
32.5 31.25 - 33.75 llll llll llll llll 19 0.212
35.0 33.75 - 36.25 llll llll llll ll 17 0.189
37.5 36.25 - 38.75 llll llll ll 11 0.122
40.0 38.75 - 41.25 llll ll 7 0.078
42.5 41.25 - 43.75 llll 2 0.022
45.0 43.75 - 46.25 ll 3 0.033
47.5 46.25 - 48.75 0 0.000
50.0 48.75 - 51.25 l 1 0.011
Total 90 1.000
In this representation we recognise directly a minimum of 20.0 N/mm2, a maximum of 50.0 N/mm2, and
an average value between 32.5 and 35.0. The measurements are distributed symmetrically about the
average value.
The graphical representation of the tally with rectangles is called a histogram.
Fig. 1 Histogram.
The area of the rectangles must be proportional to the tally not the height of the bar, e.g. if two classes
contain the same number of observations and the width of class 1 is twice the one of class 2, then the
height in class 1 is half the height in class 2. Same number of observations same area of rectangle!
Two further informative graphs are the frequency curve and the cumulative frequency curve. They are
constructed as follows:
Frequency curve
plot the class frequency (absolute or relative) against the class midpoint (again take into account
the note for the histogram, same number = same area).
Fig. 2 Frequency Curve.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:15 PM Page 7
Query:
"Holderbank" - Cement Course 2000
Cumulative frequency curve
cumulate the absolute frequencies less than the upper class boundary for every class
plot the cumulative frequency in percent against the upper class boundary.
Fig. 3 Cumulative Frequency Curve.
From this graph we can determine the portion of observations smaller than any given strength value x.
Example: Percentage of the observations smaller than 30.0 N/mm2: 26 %. Inverse
problem: Half of the measurements are smaller (resp. greater) than 33.5 N/mm2.
These two frequency curves are especially suited to compare two or more different distributions with
one another.
Fig. 4 Compare Distributions.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:15 PM Page 8
Query:
"Holderbank" - Cement Course 2000
Stem and Leaf Plot
The stem and leaf plot is very similar to the histogram but allows to reconstruct the individual data.
[The stem and leaf plot is shown in the next picture, the comment does not belong to it, but shows how
the data are reconstructed.]
STEM LEAF COMMENT
20 9 20.9
21 no observation
22 8 22.8
23 0579 23.0, 23.5, 23.7, 23.9
24 456 24.4, 24.5, 24.6
25 0
26 1 etc.
27 126899
28 455
29 039
30 0122378
31 456778
32 0244888
33 033456
34 124569
35 2233455889
36 14688
37 14689
38 169
39 247
40 1
41 029
42
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:16 PM Page 9
Query:
"Holderbank" - Cement Course 2000
43 4
44 3
45 38
46
47 9
48
49
50
Small Samples
Individual values are marked on the scale
Fig. 5 Small Sample Table.
A further attractive possibility to describe graphically the distribution of a variable, the Box Plot, will be
given in the next section. Often it is appropriate (especially in quality control) to plot the observations in
chronological order to show a possible change of the level during the experiment.
Fig. 6 Chronological Order.
If only few data are available, the single values may be represented as points on the measurement
scale (cf. example A1, Appendix I).
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:16 PM Page 10
Query:
"Holderbank" - Cement Course 2000
The scatter diagram or scatter plot is used to represent the relationship between two variables (paired
observations concerning the same individual sample). The following diagrams show the dependence of
concrete strength in N/mm2 from cement/water ratio after 2 days and 28 days. Different symbols can
be used to discriminate groups of observations. In the present case two groups are considered:
Portland cement (PC) and blended Portland cement (BPC).
Fig. 7 Scatter Diagram.
If we are dealing with more than 2 variables a suggested graphical representation is the draftmansplot,
which plots all pairwise scattergrams in one picture.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures
2. Statistical measures
As statistics deals with the behaviour of random variables, i.e. variables which take certain values with
certain probabilities, we are interested in describing this behaviour of the random variable by a few
characteristic measures. The first measures we will be considering are the p quantiles (or p fractiles),
which are roughly said those values for which 100.p % ( 0 = p = 1) of the data are smaller or equal
to. The p quantiles allow a full description of the data, similar e.g. to the cumulative frequency curve
from Section 2.1. The use of the p quantiles gives way to construct a simple and attractive graphical
representation of the data, the box plot.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.1 The
p-quantile
1. The p-quantile
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:16 PM Page 581
Query:
"Holderbank" - Cement Course 2000
xp is called the p-quantile (or p-fractile) if
100 • p%(0 = p = 1)
of the values of the random variable are smaller or equal to xp.
Based on a set of observations from a random variable x, a first method to estimate these quantiles is
given by the cumulative frequency curve.
Given a sample x(i), i=1,..,n the empirical cumulative distribution function (cdf) is defined as
F(x): = (number of the x(i)'s smaller or equal to x) / n
and the following relations between the p quantile y(p) and the cdf F(x) hold
a) F(y(p)) = p
b) Y(p) = min {x(i) : F(x(i)) >= p}
The p quantile y(p) is roughly said the value for which 100*p% of the observed data are smaller or
equal to.
Estimation of P quantile v(P)
Denote by z(i), i=l, ..,n the ordered sequence of the x(i)'s, e.g.
z(1) <= z(2) <= z(3) <= .... '= z(n 1) '= z(n)
and let p(i) = (i 0.5)/n.
Then the P quantile can be computed as follows:
1) 1) If P equals a p(i), then z(i) is the P quantile
Y(P) = z(i)
2) Otherwise compute
j(P) = nP + 0.5
and split this value j(P) into a whole number j and the remaining part B (e.g. 1.347 is split into j=1 and
B=0.347). The P quantile is then
y(P)=(1− B)z( j )+Bz( j + 1)
An example
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:17 PM Page 582
Query:
"Holderbank" - Cement Course 2000
Consider the following ordered sample
i Z(I) p(i)
1 156 0.0417
2 158 0.1250
3 159 0.2083
4 160 0.2917
5 161 0.3750
6 161 0.4583
7 163 0.5417
8 166 0.6250
9 166 0.7083
10 168 0.7917
11 172 0.8750
12 174 0.9583
12.5 % quantile (P=0.125)
As P=p(2), our 12.5 %-quantile y(.125) equals the second ordered observation, thus y(.125) = 158
10 % quantile (P=O.1)
Compute j(P)=nP + 0.5 = 12*0.1 + 0.5 = 1.7
As this yields no whole number our 10 % quantile is (j=1 and B=0.7)
y(.1) = (1 B) z(1) + B z(2)
= 0.3*156 + 0.7*158 = 157.4
25 %-quantile (P=0.25)
Compute j(P)=nP + 0.5 = 12*0.25 + 0.5 = 3.5
As this yields no whole number our 10 % quantile is (j=3 and B=0.5)
y(.25) = 0.5*159 + 0.5*160 = 159.5
50 % quantile (P=0.5)
Compute j(P)=nP + 0.5 = 12*0.5 + 0.5 = 6.5
As this yields no whole number our 10 % quantile is (j=6 and B=0.5)
y(.5) = 0.5*161 + 0.5*163 = 162
Some of the quantiles have special names, e.g.
50 % quantile median (M)
25 % quantile lower quantile (LQ)
75 % quantile upper quantile (UP)
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:17 PM Page 13
Query:
"Holderbank" - Cement Course 2000
10 % quantile lower decile (LD)
90 % quantile upper decile (UD)
OQ UQ Inter quantile range (IQ)
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.2 The
Box Plot
2. The Box Plot
Today the box-plot is perhaps the most frequently used graphical representation for univariate data. It
offers a condensed picture of the data's distribution and shows location, variability and extremes of the
data.
In order to construct the box plot we need three quartiles:
the median
the lower quantile
the upper quantile
The construction is very simple:
where L = max (LQ 1.5 x IQ, min xi)
U = min (UQ + 1.5 x IQ, max xi)
IQ = UQ IQ
x denotes extreme values.
This construction is based on the fact that for a normal distribution approximately 1 % of the data lie
outside the interval (L,U).
An alternative method also frequently used, would be not to use L and U but instead use the lower and
upper decile. Especially for non normal distributions this representation seems preferable.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.3
Measures of location
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:17 PM Page 14
Query:
"Holderbank" - Cement Course 2000
2.2.3 Measures of location
The arithmetic mean is the mean of all observations and corresponds to the centre of gravity in
physics.
x1 + x2 + ...xn
x
n
1
x= = i
n n i =1
35.8 + 39.2 + 36.8 + ... + 34.4 + 35.5
x= = 33.199N / mm2
90
The median is the central value. Half of the observations are smaller, respectively greater than the
median.
Place the measured values in an ordered array starting with the smallest value
X (1) X(2) ... X(n)
From this array we obtain the median by
x~ = x (n +1)
2 if n is odd
1
x~ = x n + x n
2 (2) ( +1)
2 if n is even
For a great number of observations the ordering is very laborious. In this case the median may be
determined graphically from the cumulative frequency curve by looking for the point on the x axis
corresponding to a cumulative frequency of 50 %.
Mean or Median?
If the histogram is symmetrical, mean and median are approximately the same. In this case the mean
is a better estimate of central tendency if the distribution is not too long tailed.
If the histogram is skew, mean and median are different.
Use the mean if you are interested in the centre of gravity or the sum of all observations
Use the median if you are interested in the centre, i.e. with equal probability a future observation
will be smaller, resp. greater than the median.
Example 2
The following histogram represents the income of 479 persons
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:17 PM Page 15
Query:
"Holderbank" - Cement Course 2000
In this extremely skew distribution the arithmetic mean is quite different from the median. Whereas the
median represents a typical income, the mean can be used in a projection to estimate the total income
of the population if the sample is drawn at random.
The trimmed mean is used to estimate the mean of a symmetrical distribution if gross errors in the data
are suspected.
To calculate the -trimmed mean the -percent largest and smallest values are deleted. The trimmed
mean is then the arithmetic mean of the remaining observations. Notation for the 5%-trimmed mean:
xtrimmed 0 05
Usual values for : 5 % to 10 %
The weighted mean is used if certain weighting factors wi are associated with the observation xi.
Reasons may be:
- samples of unequal weights
- observation are measured with unequal precision
xw= w x i i
w i
Schematic survey on the use of measures of location:
Type of distribution
Measure to be used: arithmetic mean x
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:18 PM Page 16
Query:
"Holderbank" - Cement Course 2000
Measure to be used: trimmed mean xtrimmed
Measure to be used: median ~
x
Measure to be used: median or arithmetic mean. The choice depends on the interest of the user.
Remark: The median is equal to the 50%-trimmed mean.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.4
Measures of Variability
2.2.4 Measures of Variability
The variance is the mean square deviation of single observations from the mean.
n
s =
2 1
(x i − x )2
n −1 i =1
For practical computations use
1 2
i x 2 − ( x i )
s2 = 1
n − 1 n
The standard deviation is the positive square root of the variance.
s = + s2
In our example:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:18 PM Page 17
Query:
"Holderbank" - Cement Course 2000
n
s=
1
n − 1 i =1
(x i − x )2 = 55.596N / mm2
The coefficient of variation is a relative measure of the variability (relative to the mean).
s
v=
x
v is often used for comparing variabilities. It is especially useful if s increases proportionally with the
mean x , i.e. v = constant.
Example of compressive strength:
5.596
v= = 0.169 = 16.9%
33.199
The range is the difference between the largest and the smallest value.
R = X(n) - x(1) = Xmax - Xmin
Useful with small sample sizes.
Note: The range increases in general with increasing sample size n (greater probability to get extreme
values).
If the variables under consideration are assumed to be normally distributed then
the range can be used to provide an estimate of the standard deviation, according to
s = R k where k ~= n for 3 n 10
For sample size n >10 divide the set of observations in random sub samples of size n', calculate the
arithmetic mean R of the range in these sub-samples, and estimate s according to
s ~R k
using a k-value corresponding to n'. . For a first rough estimation the following k values may be used:
k~4 for n = 20
k ~ 4.5 for n = 50
k ~5 for n = 100
The interquartile range is the difference between the upper and lower quartiles.
Q = x.75 − x.25
4
s Q.
If the distribution is symmetric 3
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.5
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:19 PM Page 18
Query:
"Holderbank" - Cement Course 2000
Statistical program packages
5. Statistical program packages
There are a lot of statistical programs on PCs, microcomputers and mainframes that offer the
possibility to compute these statistical measures including quartiles and even to plot the box plots. The
perhaps most prominent under many others are SAS, SYSTAT/SYGRAPH, SPSS, BMDP,
STATGRAPHICS.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.6
Interpretation of the standard deviation
6. Interpretation of the standard deviation
In the histogram of example 1 (chapter 2.1) frequency of observations decreases on both sides of the
mean. The distribution law seems to be symmetrical. In the present case we consider the 90
measurements to be a random sample of a specific distribution model, the normal distribution (or
Gaussian distribution). This type of distribution is represented by a symmetrical, bell shaped curve, and
often observed in practical applications.
The normal distribution is determined by the mean and the standard deviation.
For this distribution about 68 % of the observation are expected in the interval mean 1 standard
deviation.
If the distribution is not normal, the standard deviation allows no direct interpretation. In this case we
can determine the interval that contains a certain portion of the observations graphically from the
cumulative frequency curve.
Another possibility to get an interpretation of the standard deviation in a skew distribution is
“normalising” of the distribution by transformations. In the case of example 2 (chapter 2.2.3), the
logarithms of incomes show approximately a normal distribution. So called variance-stabilising
transformations are especially used to fulfil normality conditions in higher statistical analysis.
For further information consult e.g. Natrella (1963).
Further properties of the normal distribution are given in section 3.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 2. SIMPLE DATA DESCRIPTION / 2.2 Statistical measures / 2.2.7
Outliers
2.2.7 Outliers
As mentioned in the introduction we have to expect about 5 % to 10 % gross errors in a set of
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:19 PM Page 19
Query:
"Holderbank" - Cement Course 2000
observations. Most of them may not be recognisable in the region of all other values. Some may be
extremely outlying values with an important influence on statistical results.
In modern statistics robust methods are studied, which are not sensitive to a certain portion of gross
errors, as for example the median or the trimmed mean. More sophisticated robust procedures are in
general rather complicated.
If classical measures as the standard deviation and the arithmetic mean are used, we have to check
the data for outliers and to eliminate them from the set of observations (cf. example A2, Appendix I).
Note: If outliers are eliminated, they must be recorded separately in the report.
To detect outliers, check the data plot (tally, histogram or chronological order) for suspicious values. If
outliers are suspected and the data show a normal distribution, use the Dixon criterion for rejecting
observations (n<26).
The Dixon Criterion
Procedure
1) Choose a, the probability or risk we are willing to take of rejecting an observation that really belongs
in the group.
2) If
3n7 Compute r10
8 n 10 Compute r11
11 n 13 Compute r21
14 n 25 Compute r22
1) where rij is computed as follows
rij If X (n ) is suspect If X (1) is suspect
r10 ( X(n ) − X( n−1) ) ( X(n ) − X(1) ) ( X(2) − X(1) ) ( X(n ) − X(1) )
r11 ( X(n ) − X(n−1) ) ( X(n ) − X(2) ) ( X(2) − X(1) ) ( X(n −1) − X(1) )
r21 ( X(n ) − X( n−2) ) ( X(n ) − X(2) ) ( X(3) − X(1) ) ( X(n −1) − X(1) )
r22 ( X(n ) − X( n−2) ) ( X(n ) − X(3) ) ( X(3) − X(1) ) ( X(n−2) − X(1) )
3) Look up r1−α / 2 for the rij from Step (2), in Table A 2
4) If rij ri −α / 2 reject the suspect observation; otherwise, retain it.
In the case of a sample size n>25 use the following procedure:
1) Choose , the probability or risk we are willing to take of rejecting an observation that really
belongs to the group
x( n) − x(1)
ZB =
2) Calculate s
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:19 PM Page 590
Query:
"Holderbank" - Cement Course 2000
3) Look for ZT,1−α in Table A 3, Appendix III
(4) If Zb ZT,1−α reject the suspect observation, otherwise retain it.
Note: The presented outlier rejecting rules are only valid in normal distributions. In a skew distribution,
the elimination of outliers is very dangerous and should be avoided. In this case the reason for the
extreme observation must be known.
A check for outliers is not necessary if the trimmed mean or the median is used and if we are only
interested in a location measure.
For further tests on outliers cf. "Wissenschaftliche Tabellen Geigy, Statistik".
The standard deviation is extremely sensitive to outliers.
A check is therefore important, because it is not allowed to calculate a standard deviation from a
trimmed set of observations.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 3. THE NORMAL DISTRIBUTION (ND)
3. THE NORMAL DISTRIBUTION (ND)
As mentioned in section 2, the normal distribution is a theoretical model of a statistical universe
(population), defined by the mean and the standard deviation . Graphically it is represented by a
smooth, symmetric mean x and the standard deviation s of the samples are estimated for the true, but
unknown values and .
By the standardisation formula
x−μ x−x
Z=
σ (respectively s if n is large)
the ND is transformed in a normalised form with mean 0 and standard deviation 1.
Formulae for the standard normal distribution:
Density function (bell shaped curve):
1 z2
f (z) = exp(− )(− z )
2π 2
Especially of interest is the area under the curve (distribution function), which corresponds for every
given value z to the probability of an observation to be smaller than z.
z
1 z2
ϕ(z) =
2π
exp(− 2
)dz,ϕ(−z) = 1− ϕ(z)
−
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:20 PM Page 591
Query:
"Holderbank" - Cement Course 2000
Numerical values for z and (z) are given in Table A 1, Appendix III. The probability to observe a
measured value between two given limits T1 and T2 can be calculated as follows:
1) transform the limits T1 and T2 in a standardised form
2) calculate the probability with help of Table A 1 by
F(z1,z2 ) = ϕ(z2 ) − ϕ(z1)
Interpretation of the standard deviation in a normal distribution:
We expect about 95 % of the observation between -2 and +2 . More general we have:
The statement that the values (of the population) lie between +z1-/2.. and -z1-/2.. is right with the
probability S = 1 - and wrong with probability . One sided or two sided regions may be considered.
The corresponding z-values are taken from a table of the standard normal distribution.
Usual percentiles for the statistical confidence S:
a) Two sided, S = 1− α;zα 2,z1−(α / 2)
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:20 PM Page 22
Query:
"Holderbank" - Cement Course 2000
S(%) 90 95 99 99.9
-z /2 = z1 -- /2 1.64 1.96 2.58 3.29
b) One sided, S = 1− α;zα ,z1−α )
S(%) 90 95 99 99.9
-z = z1 -- 1.28 1.64 2.33 3.09
How to check normality?
Before using the characteristics of the ND the validity of the model has to be checked.
Method: Draw the cumulative frequency curve in the normal probability paper (Appendix II). Normality
can be assumed if the resulting curve is approximately a straight line between 5% and 95%.
Example: Normal probability plot of compressive strength data
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:20 PM Page 23
Query:
"Holderbank" - Cement Course 2000
Conclusion: The observations of compressive strength follow approximately a normal distribution.
Short cut rule for rejecting normality: If no negative values are allowed in the observations and the
coefficient of variation is greater than 30 %, the distribution is not normal.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 4. CONFIDENCE LIMITS
4. CONFIDENCE LIMITS
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 4. CONFIDENCE LIMITS / 4.1 Confidence limits for the mean
1. Confidence limits for the mean
The arithmetic mean is an estimate for the true, unknown mean of the population of all possible
observations. We may be interested in the precision of this estimation. Therefore we calculate a
confidence interval that contains the true value with high probability (confidence level).
a) If the distribution is normal:
Two sided confidence interval at confidence level 1-
s
x t (1−α / 2;f ) .
n
with
f = n-1 degrees of freedom
t1-/2;f given in Table A 4, Appendix III (Student t-Distribution)
Interpretation:
With probability 1- the true mean lies between the two confidence limits.
For n>50 t1-/2;f may be replaced by z1-/2 given on page 15 (corresponding to normal deviates of
Table A 1, Appendix III). In our example of compressive strength we calculate the 95% confidence
interval by replacing the t- by the z- value:
s ~ s
x t 0.975;89 .
=x z0.975
n n
5.6
→33.2 1.96 = 33.2 1.16
90
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:20 PM Page 24
Query:
"Holderbank" - Cement Course 2000
b) If the distribution is not normal:
Approximate confidence intervals can be obtained by
s
x z1−α / 2 .
n
This approximation is derived from the central limit theorem and can be used if
10s2
n
x2
c) Range method:
In the case of a normal distribution, confidence limits may be calculated by use of the range instead of
the standard deviation (often used in quality control for small sample sizes n).
x λ1−α / 2R
Values for are given in Table A 6, Appendix III.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 4. CONFIDENCE LIMITS / 4.2 Confidence limits for the median
4.2 Confidence limits for the median
Confidence limits for the median can be determined directly from the ordered array of observations.
x(1) x(1) .... x(n−1) x(n)
(1-)-confidence interval:
x(k) < Median < x(n-k+1)
n − 1 z1−α / 2
k= − n −1
with 2 2 rounded to the next lower integer value
z1-/2 is given in Table A 1, Appendix III.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 4. CONFIDENCE LIMITS / 4.3 Confidence limits for the standard
deviation
4.3 Confidence limits for the standard deviation
If the population follows a normal distribution, two sided confidence limits are given by
f f
s 2
σ s 2
x
1−α / 2;f
x α / 2;f
with f = n 1 degrees of freedom
2 2
x1−α / 2;f ,xα / 2;f critical values of the chi-square distribution given in Table A 5, Appendix III
Confidence intervals are reduced with increasing sample size n, i.e. the more observations available,
the better is the estimation.
The degree of improvement for the arithmetic mean can be derived from the fundamental central limit
theorem. If several samples of size n are drawn from the same population, the arithmetic means of
these samples are approximately normal with mean and standard deviation s/n, the so called
standard error of the mean (SEM).
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:22 PM Page 25
Query:
"Holderbank" - Cement Course 2000
For increasing n to infinity, the standard error of the mean tends to zero, i.e. the estimation tends to be
absolutely precise if no systematic errors are present.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 4. CONFIDENCE LIMITS / 4.4 Other methods for the construction of
confidence limits
4.4 Other methods for the construction of confidence limits
Sometimes it happens that we have to deal with very complicated functions of random variables, for
which we can't derive or know the underlying distribution function, but we would like to have confidence
limits for the values of this complicated function. A newer method, the Bootshap method (B. Efron,
1983), allows to obtain such results, but is rather intensive in calculation.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS
5. STANDARD TESTS
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.1 General Test Idea
1. General Test Idea
Tests are used to make decisions in the case of incomplete information. We want to know if some
observed difference is significant or only an effect of random errors.
Significant:
We decide that a difference really exists. The probability of this decision to be false is known and can
be chosen by the decision maker. A usual choice of this error probability is 5%.
Not significant:
The observed difference may be realized by chance alone. The data give no argument to suppose the
existence of a real difference. If such a difference really exists the sample size n is too small to detect
it.
Example 3:
At a cement plant A, 618 titration samples gave the mean value x1 = 82.67 and the calculated
standard deviation s1 = 4.13 .
At a later stage a further series of 525 samples were taken and after processing, yielded the
information of calculated mean value x2 = 85.58 and calculated standard deviation s2 = 3.79
Your decision problem is:
State whether a significant change has occurred under plant conditions!
Example 4:
The production rate of a cement mill in tons/hr was measured as:
28.3 27.2 29.3 26.7 29.9 24.6 25.0 30.0 26.3 27.8
After modification, the production rate of the same mill in tons/hr was measured as:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:22 PM Page 26
Query:
"Holderbank" - Cement Course 2000
28.0 30.0 30.5 26.0 31.0 30.3 24.6 25.4 26.7 29.3
Your decision problem is:
Has the modification made a significant improvement?
There are many problems in which we are interested in whether the mean (or another parameter value)
exceeds a given number, is less than a given number, falls into a certain interval, etc.
Instead of estimating exactly the value of the mean (or another parameter), we thus want to decide
whether a statement concerning the mean (or other parameter value) is true or false, i.e. we want to
test a hypothesis H0 about the mean (or another value).
In example 3 the hypothesis H0 could be
“No significant change has occurred in the plant conditions"
In example 4 the hypothesis H0 could be
“No significant improvement has been made by the modification”
We will solve the two given decision problems later in section 5.2 (test procedures).
Now we will consider another specific example introducing at the same time the important parts of all
similar decision problems.
Example 5:
In the manufacture of safety razor blades the width is obviously important. Some variation in dimension
must be expected due to a large number of s mall causes affecting the production process. But even
so the average width should meet a certain specification. Suppose that the production process for a
particular brand of razor blades has been geared to produce a mean width of 0.700 inches. Production
has been underway for some time since the cutting and honing machines were set for the last time,
and the production manager wishes to know whether the mean width is still 0.700 inches, as intended.
We call set of all blades coming from the production line in a certain time interval (t, t+h) the statistical
population to be studied. For example t = January 5th, 0oo and t+h = January 6th, 000, if the statistical
population we are interested in is the set of all blades produced on January 5th.
If the production process was initially set up on that day to give a mean width of 0.700 inches, we can
say that the hypothesis H0 "the produced mean width of the regarded population is 0.700" should be
tested. In symbols this is: µ0 = 0.700 = hypothesized mean.
Accepting the Hypothesis H0:
Suppose we draw a simple random sample of 100 blades from the production line. We measure each
of these carefully and find the mean width of the sample to be 0.7005 inches. The standard deviation in
the sample turns out to be 0.010 inches. That is,
n = 100
x = 0.7005 inches
s = 0.010 inches
For the hypothesis µ0 = 0.700 to be true, the sample mean x = 0.7005 inches would have to be drawn
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:22 PM Page 27
Query:
"Holderbank" - Cement Course 2000
from the sampling distribution of all possible sample means whose overall mean is 0.700 inches.
Now the important question arises: If the true mean of the population really were 0.700 inches, how
likely is it that we would draw a random sample of 100 blades and find their mean width to be as far
away as 0.7005 inches or farther? In other words, what is the probability that a value could differ by
0.0005 inches or more from the population mean by chance alone?
If this is a high probability, we can accept the hypothesis that true mean is 0.700 inches, because it is
very easy to get it (high probability).
If the probability is low, however, the truth of the hypothesis becomes questionable because the
sample we got is in reality very seldom.
To get at this question, compute the standard error of the mean from the sample:
s 0.010
sx = = = 0.001inches
n 100
Since the difference between the hypothetical mean and the observed sample mean is 0.0005 inches
and the standard error of the mean is 0.001 inches, the difference equals 0.5 standard errors. By
consulting Table A-1 (zp = 0.5)*(3) we find that the area within this interval around the mean of a
normal curve is 38%, so that 100 - 38 = 62 % of the total area falls outside this interval (cf. dashed
lines below).If 0.700 inches were the true mean, therefore, we should nevertheless expect to find that
about 62% of all such possible means would, by chance alone, fall as far away as 0.5 sx or farther.
Therefore, the probability is 62% that our particular sample mean could fall this far away. This is a
substantial reason to accept the hypothesis and attribute to mere chance the appearance of a 0.7005
inches mean in a single random sample of 100 blades.
Naturally we reject in the same time the contrary of H0 namely that:
μ 0 :μ 0.700
In case H0 can not be rejected, avoid saying “it is proved that H0 is correct”. or "there are no
differences in the means”, say "there is no evidence that H0 is not true" or "there is no evidence H0
should be rejected”.
Rejecting the Hypothesis H0:
Later, after production has gone on for some time, the query again arises:
Is it reasonable to believe that the true mean width of blades produced remains 0.700 inches? Since
the process was adjusted to yield that figure the hypothesis still seems reasonable. We could then test
it by taking another random sample of 100 blades.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:23 PM Page 28
Query:
"Holderbank" - Cement Course 2000
This time the standard deviation is still 0.010 inches, so the standard error of the mean is still 0.001
inches, but the mean is now 0.703 inches:
In order to test the hypothesis that the true mean of the population is 0.700 inches, we again go
through the same line of reasoning. If the true population mean really were 0.700 inches, how likely is it
that we should draw a random sample of 100 blades and find their sample mean to be as far away as
0.703 inches?
Since the difference between the hypothetical mean of 0.700 inches and the actual sample mean of
0.703 inches is 0.003 inches, and the standard error of the mean is 0.001 inches, the difference is
equal to three standard errors of the mean (i.e. 0.003/0.001 = 3).
Now if 0.700 inches really were the population mean, we know from Table A-1, Appendix III that 99.7%
of all possible sample means, for random samples of 100, would fall within three standard errors
around 0.700 inches. Hence, the probability is only 0.3% that we would get a sample mean falling as
far away as ours does.
We have two choices:
1) We may continue to accept the hypothesis (i.e. leave the production process alone), and attribute
the deviation of the sample mean to chance.
2) We may reject the hypothesis as being inconsistent with the evidence found in the sample (hence,
correct the production process).
Either of two things is true and we have to make a decision between them:
1) the hypothesis is correct, and an exceedingly unlikely event has occurred by chance alone (one
which would be expected to happen only 3 out of 1000 times); or
2) the hypothesis is wrong
Type I and Type II Errors
Understandably, the question can be raised: What critical value should we select for the probability of
getting the observed difference (x - 0) by chance, above which we should accept the hypothesis H0
and below which we should reject it? This value is called the
critical probability or level of significance
The answer to this question is not simple, but to explore it will throw further light on the nature and logic
of statistical decision making. Let's study the following example:
in H0 is true H0 is false
In reality H0 is true H0 is false
We decide
Accept H0 3 right decision 2 error II
accept a true accept a
hypothesis false
hypothesis
Reject H0 1 error I 4 error II
reject a true accept a
hypothesis false
hypothesis
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:23 PM Page 29
Query:
"Holderbank" - Cement Course 2000
Example:
H0: the parliament building burns
Decision maker is the commander of the fire brigade.
Another expression for error I is: error of first kind
Another expression for error II is: error of second kind
If we ask here, what is worse
error I or error II,
then error I naturally costs a lot of money because the fire will destroy the whole parliament building.
Error II only moves the fire-brigade.
In a long run of cases which the hypothesis is in fact true (although we do not know it is true, for
otherwise there would be no need to test it), we will necessarily either be wrong as in 1 or right as in 3.
That is to say, if we make an error it will be of Type I.
Suppose we should adopt 5% as the critical probability. accepting the hypothesis when the probability
of getting the observed difference by chance exceeds 5% and rejecting the hypothesis when this
probability proves to be less than 5%. This amounts to the decision to accept the hypothesis when the
discrepancy of the sample mean is less than 1.96 standard deviations, and to reject the hypothesis
when the discrepancy is more than 1.96 standard deviations.
If we take 1% instead of 5%, as above, we will get as limit 2.58 standard deviations.
In fact, the percentage of cases in which we would expect to make an error of the first kind is precisely
equal to the critical probability adopted.
(The probability of error I will be abbreviated quite often by ).
Just significant probability level:
In many studies the critical probability is used to describe the statistical significance of a sample result.
For example, an economist collects some data on, say, interest rates and the demand for money. He
hypothesizes some relationship and wishes to see if the data support his thesis. He tests the
hypothesis to rule out the alternative that the observed relationship occurred by pure chance. He
reports his sample results as "significant at the 1 percent level". Such a statement is a report to the
reader that has the following meaning:
1) if we were to set up a statistical hypothesis
2) if we were to test this hypothesis using a critical probability of 1%
3) then we would reject the hypothesis and rule out a chance relationship
Significance levels of 10%, 5%, 1%, 0.1% are often used in reporting sample data. The smallest of
these probability values is chosen at which the hypothesis can be rejected.
So we see now what is basic for every statistical test:
1) we need a clear hypothesis H0
2) we have to know by which statistic we want to test H0 (in our example it was x )
3) we need a criterion C for decision making in the following form:
reject H0 if C applies
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:23 PM Page 30
Query:
"Holderbank" - Cement Course 2000
accept H0 if C does not apply
The criterion C is usually given in form of a critical limit, called significance limit, which should not be
exceeded by the calculated test statistic.
To perform such a statistical test, it is necessary to select the risk to commit a type I error, i.e. the
significance level (see section 5.2, p. 28).
Often a relevant difference has to be detected with a certain probability, i.e. with a predetermined type
II error . This is only possible with a certain sample size n, as is further outlined in section 5.3.
If we reject a hypothesis when it should be accepted, we say that a Type I error has been made. If, on
the other hand, we accept a hypothesis when it should be rejected, we say that a Type II error has
been made. In either case a wrong decision or error in judgement has occurred.
In order for any tests of hypotheses or rules of decision to be good, they must be designed so as to
minimize errors of decision. This is not a simple matter since, for a given sample size, an attempt to
decrease one type of error increases the other type. In practice one type of error may be more serious
than the other, and so a compromise should be reached in favour of a limitation of the more serious
error. The only way to reduce both types of error is to increase the sample size, which may or may not
be possible.
Note that from the philosophy of testing there is always the possibility to reject H0 although H0
effectively is true (the probability for this is , the type I error probability). So if you are testing say 100
“correct” datasets (i.e. for which H0 holds) to a significance level = 5% you will expect about five
results that reject H0, although it holds. This has to be considered if a lot of (statistical) tests are made
on the same data material (see e.g. Multiple Testing ....... .”Simultaneous Statistical Inference”, Miller,
1981).
One-sided and two-sided tests
In example 5 we are interested in values of significantly higher or smaller than 0. Any such test which
takes account of departures from the null hypothesis H0 in both directions is called a two-sided test (or
two-tailed test H0 :μ = μ 0 ,A : μ μ 0 )
However, other situations exist in which departures form H0 in only one direction is of interest. In
example 4 we are interested only in an improvement of production rate due to the modification of the
cement mill and so a one-sided test is appropriate. It is important that the decision maker should
decide if a one-tailed or two-tailed test is required before the observations are taken.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures
5.2 Test Procedures
Performance of a statistical test requires in advance the recognition and formulation of the present
decision problem (test situation). Use the following checklist for applications of statistical tests:
CHECKLIST
1) Is the decision problem concerned with the mean. median. standard deviation. correlation
coefficient or other?
Formulate a clear hypothesis H0.
2) Is it a one-sample or a two-sample problem?
One-sample problem:
Mean, median or standard deviation of a given sample of size n is compared with a hypothetical
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:23 PM Page 31
Query:
"Holderbank" - Cement Course 2000
value (example 5).
Two-sample problem:
The decision problem is concerned with differences between two given samples.
3) For the two-sample problem: Are the observations independent or paired? Observations are paired
if every value in the first sample can be attached definitely to a value in the other sample. Paired
comparisons are in general considerably more efficient than a comparison of independent samples.
4) Is the sample drawn from a normal distribution?
5) Decide between one-sided or two-sided test.
One-sided tests are used if differences only in one direction may occur or are of interest.
Two-sided tests are used if no preliminary information is available, in what direction the sample
may differ.
6) Choose the test procedure in the table “TEST CONCERNED WITH” (sample sizes n<50 are
considered to be small).
7) Choose the significance level a. Usual choices are = 5% or = 1% (error of type I).
8) Calculate the test statistic T of the chosen test (cf. the formulae of the following pages).
9) Look for the significance limit Tp for the test statistic in the corresponding table with p = 1-/2
(two-sided test) or p = 1- (one-sided test).
10) Decision: If the calculated test statistic exceeds the significance limit, reject the hypothesis H0 (the
observed difference is significant at the level ). Otherwise there is no reason to reject the
hypothesis H0 and the difference is considered to be not significant.
TEST CONCERNED WITH
test situation median mean general normal mean
distribution standard
deviation
0ne-sample problem
small n (n<50) sign test signed-rank-test t-test x2-test
large n (n50) sign test z-test z-test x2-test
Two-sample problem
independent
small n median test Wilcoxon-test t-test F-test
large n median test z-test z-test F-test
paired
small n sign test signed-rank-test t-test -
large n sign test z-test z-test -
More than two x2-test Krukskal- Analysis of Bartlett-test
samples Wallis-test variance (ANOVA)
Friedmann-test
(These tests are not treated i n this paper)
Selected test statistics
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.1 z-Test
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:24 PM Page 32
Query:
"Holderbank" - Cement Course 2000
5.2.1 z-Test
Used to test means with large sample size n.
a) One-sample-problem:
Given is a sample of size n with arithmetic mean x and standard deviation s, we test the
hypothesis that the mean (estimated by x ) is equal to a given or target value 0
Hypothesis H0 = 0
(x − μ 0 ) n
z=
Test statistic σ
In general is not known. It can be replaced by s for large sample sizes.
b) Two independent samples:
Given are two samples of size n1 and n2 with arithmetic means x1,x2 and standard deviations s1, s2.
Hypothesis H0 1 = 2
x1 − x2
z=
s12 s22
+
Test statistic n1 n2
c) Paired comparison
Given is a sample of n paired observations xi, yj. Calculate the arithmetic mean d of all the
differences di = yi - xi and the standard deviation sd.
Hypothesis H0 x = y respectively d = 0
d n
z=
Test statistic sd
Significance limits
Often used significance limits for z1- (one-sided test) and z1-/2 (two-sided test) are
z0 95 = 1.645 z0.975 = 1.960
z0 99 = 2.326 z0 995 = 2.576
For other significance levels look for z1- in Table A-1.
Decision
The difference is significant if
z z1−α / 2 two sided test
z z1−α
z −z 1−α one-sided test
Example
a) In the example 1 we test the two-sided hypothesis H0:
µ = 35.0 N/mm2 = µ0 (one-sample problem).
x = 33.2
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:24 PM Page 33
Query:
"Holderbank" - Cement Course 2000
s = 5.6
n = 90
= 0.05
(33.2 − 35.0) 90
z= = −3.05
5.6
Izl = 3.05 ' 1.96 = Z0.975
Decision: The sample mean differs significantly from the standard strength 35.0 N/mm2.
b) in example 3 (chapter 5.1) we test the hypothesis whether a change in plant conditions has
occurred or not (two independent samples):
H0: µ1 = 2
n1 = 618, n2 = 525
x1 = 82.67, x2 = 85.58
s1 = 4.13, s2 = 3.79
= 0.05 (two-sided test)
85.58-82.67
85.58 − 82.67
z=
(4.13)2 (3.79)2
+
618 525
z = 12.41 1.96 = z0.975
Decision: Highly significant difference between the two samples of titration values.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.2 Sign Test
5.2.2 Sign Test
Test for the median in one-sample problems and paired comparisons.
a) One-sample problem: Given is a sample of size n.
We test the hypothesis that the median of the population is equal to a given value m0
H0: median = m0
Test statistic:
Count the number of observations smaller than m0 and those larger than m0. The test statistic k is
then the smaller of the two numbers.
b) Paired comparison: Given are n paired observations xi, yj.
We test the hypothesis that xi and yj have the same median.
Test statistic:
For each pair of observations xi, yj record the sign of the difference yi - xi .
The test statistic k is then the number of occurrences of the less frequent sign.
Significance limits
n −1 z1−α / 2
kα / 2 = − n −1
Calculate 2 2 (two-sided test).
For the one-sided test replace /2 by . Zp is the significance limit of the z-test
z0.95 = 1.645 z0.975 = 1.960
Decision
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:25 PM Page 34
Query:
"Holderbank" - Cement Course 2000
If k is less than k/2 (resp. k) conclude that the medians are different, otherwise, there is no reason to
believe that the medians differ. 44
Example
Test the hypothesis that 50% of the population has an income of more than 3000 (example 2).
H0: m0 = 3000
k = 217
n = 479
= 0.05 two-sided test
478 1.96
k0.025 = − 478 = 217.6
2 2
The test is just significant at the 5% level. Because the sample median is 2700, we conclude that less
than 50% of the population has an income of 3000.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.3 Signed-rank test
5.2.3 Signed-rank test
Test for the mean in symmetrical distributions with small sample sizes n (one-sample or paired
comparison).
a) One-sample problem: Given is a sample of size n from a symmetrical distribution.
Hypothesis H0 : µ = µ0
Disregarding signs, rank the di according to their numerical value, i.e., assign rank 1 to the smallest
observation, assign the rank of 2 to the di which is next smallest, etc. In case of ties, assign the
average of the ranks which would have been assigned if the di's had differed only slightly. (If more than
20% of the observations are involved in ties, this procedure should not be used).
To the assigned ranks 1, 2, 3, etc., prefix a + or a - sign, according to whether the corresponding di is
positive or negative. The test statistic T is then the sum of these signed ranks.
a) Paired comparisons: Given are n paired observations xi, yi.
The hypothesis is tested that both have the same mean.
Test statistic: Compute di = yi - xi for each pair of observation and continue as in a).
Significance limits
Look for T0.95 or T0.975 in Table A-7. If the number m of differences di exceeds 20 perform a z-test with
T
zT =
m(m + 1)(2m + 1)
6
Decision
Conclude that the means differ if
T T1−α / 2 (or zT z1−α / 2 ) two-sided test
T T1−α (orzT z1−α ) one-sided test
T −T1−α (orzT −z1−α ) one-sided test
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:25 PM Page 35
Query:
"Holderbank" - Cement Course 2000
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.4 Wilcoxon-Test
5.2.4 Wilcoxon-Test
Test for a comparison of means in two independent samples:
Given are two independent samples of size n1 and n2.
We test the hypothesis that both samples have the same mean
H0:µ1 = µ2
Test statistic:
Combine the observations from the two samples, and rank them in order of increasing size from
smallest to largest. Assign the rank of 1 to the lowest, a rank of 2 to the next lowest, etc. (Use
algebraic size, i.e., the lowest rank is assigned to the largest negative number, if there are negative
numbers). In case of ties, assign to each the average of the ranks which would have been assigned if
the tied observations had differed only slightly. (If more than 20% of the observations are involved in
ties, this procedure should not be used).
Let n1 = smaller sample
n2 = larger sample
n = n1 + n2
Compute R, the sum of the ranks for the smaller sample. (If the two samples are equal in size, use the
sum of the ranks for either sample).
Compute W = 2R - n1(n + 1)
Significance limits
Look for W0.95 or W 0.975 in Table A-8. For sample sizes not mentioned in the table perform a z-test with
W
ZW =
n1n2 (n + 1)
3
Decision
Conclude that the means differ if
W W1−α / 2 (or zw z1−α / 2 ) two-sided
W W1−α (orzw z1−α ) one-sided
W −W1−α (orzw −z1−α ) one-sided
Example
In example 4 (chapter 5.1) we are interested in an improvement of production rate after a modification
of the cement mill.
= 0.05 (one-sided test)
Combined sample (values in italics correspond to the sample after modification):
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:25 PM Page 36
Query:
"Holderbank" - Cement Course 2000
tons/hr 24.6 24.6 25.0 25.4 26.0 26.3 26.7 26.7 27.2 27.8
rank 1.5 1.5 3 4 5 6 7.5 7.5 9 10
tons/hr 28.0 28.3 29.3 29.3 29.9 30.0 30.0 30.3 30.5 31.0
rank 11 12 13 14 15 16.5 16.5 18 19 20
The sum R of ranks in the second sample is
R = 1.5 + 4 + 5 + 7.5 + 11 + 13 + 16.5 + 18 + 19 + 20 = 115.5
W = 2 . 115.5 - 10 . 21 = 21
W = 21 < 46 = W0.95
Decision
The improvement is not significant. We have no reason to assume a real improvement due to the
modification.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.5 t-Test
5.2.5 t-Test
This test is used for comparisons of means in normal distribution with small sample sizes n. If the
distribution is not known to be normal, prefer rank tests (Wilcoxon, signed ranks).
The test situation is the same as in 5.2.1 (z-test) refer for remarks on hypothesis and decision. Instead
of z1- look for significance limits t1- in Table A-4, with n-1 degrees of freedom (df) in the one-sample
case or for paired comparison, and with n1 + n2 - 2 degrees of freedom (df) for the two-sample case
respectively.
One-sample problem: Given is a sample of size n drawn from a normal distribution.
(x − μ 0 n
t=
s
Two independent samples: Given are two independent samples of size n1 and n2 from two normal
distributions with means µ1, µ2 and equal standard deviation .
x2 − x1 n1n2
t=
(n 1 −1)s21 + (n 2 − 1)s22 n1 + n2
(n1 + n2 − 2)
Paired comparison: Given are n pairs of observations xi, yi, when x and y follow a normal distribution.
d
t= n
sd
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.6 Median-Tests
5.2.6 Median-Tests
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:26 PM Page 37
Query:
"Holderbank" - Cement Course 2000
Median tests for two independent samples are not given here in an explicit form. For small sample
sizes n Fisher's test for 2 x 2 contingency tables and for large n x2-test for contingency tables may be
used (Natrella 1963, Noether 1971).
In a comparison of two independent samples we are often not interested in differences only of means
or medians, but generally in location differences of the two samples.
In this case use the tests for means, which are more efficient than median tests (Wilcoxon-test or
z-test).
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.7 X2-Test
5.2.7 X2-Test
Several X2-tests are available for different test situations. The test we explain here is used to compare
the standard deviation s of a sample with hypothetical value 0.
X2-tests are also used for
Goodness of fit test
Independence test
Loglinear models
(see e.g. Haberman, 1978)
Given is a sample of size n drawn from a normal distribution with mean and standard deviation . We
test the hypothesis that the standard deviation , estimated by s from the sample is equal to a
hypothetical standard deviation 0.
(n − 1)s2
x2 =
H0: = 0 Test statistic: σ 02
Significance limits
Look for significance limit X2 in Table A-6 with m = n-1 degrees of freedom (df).
p,m,
Decision
If X X
2 2
or X2 X2 conclude: = 0 (two-sided)
1-/2;m /2;m
If X2 X2 conclude: > 0 (one-sided)
1-;m
X2 X2 ;m conclude: < 0 (one-sided)
Otherwise we have no reason to believe that differs from 0.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.2 Test Procedures / 5.2.8 F-Test
5.2.8 F-Test
Comparison of the standard deviation in two independent samples.
Given are two independent samples of size n1 and n2 with standard deviations s1 and s2. Both samples
are drawn from a normal distribution. We test the hypothesis that both populations have the same
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:26 PM Page 608
Query:
"Holderbank" - Cement Course 2000
standard deviation.
H0: 1 = 2
Test statistic:
Let be s1 > s2, then compute
s12
F=
s22
Significance limits
Look for F1-;m1,m2 (one-sided) or F1-/2;m1,m2 (two-sided) in Table A-9 with m1 = n1-1 and m2 = n2-1
degrees of freedom.
Decision
Conclude 1 2 if F F1-/2;m1,m2 (two-sided)
or 1 > 2 if F F1-;m1,m2 (one-sided)
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 5. STANDARD TESTS / 5.3 Sample Size Determination
5.3 Sample Size Determination
As stated at the end of chapter 5.1, the probability of making a type II error in a given test with
significance level (type I error) depends upon the sample size n.
How can we determine the sample size that is necessary to hold the type II error within certain
boundaries (probability )? Or in other words: What sample size n is necessary to detect a relevant
difference with great probability (1-)?
Let us first examine a one-sided one-sample-test for testing a mean:
Suppose we are interested in the mean µ of a random variable X. We want to test the hypothesis H0: µ
= µ0 against the alternative hypothesis A: µ > µ0 with a z-test. The error-type-one shall be .
The probability to detect a deviation of from µ0 shall be at least (1 - ), if this deviation exceeds a
preselected relevant difference ( is now an upper bound for the type II error).
The following figure shows the distribution of the sample mean x under H0 and under the (special)
alternative hypothesis A: = 0 + :
Interpretation of the graph:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:27 PM Page 39
Query:
"Holderbank" - Cement Course 2000
CHo is the criteria (significance limit) of the test:
The hypothesis H0 is rejected with probability even if it is true.
On the other hand, the probability to accept H0: = 0, even if the true mean is = 0 + (type II
error), is greater than .
The relevant difference may only be detected with probability (1-), if CHo CA; in this case the
risk to commit a type II error is .
σ2
The variance n of the two distributions becomes smaller with increasing n and hence CA moves to
the right and CHo to the left. Therefore, we can find an n such that CHo CA:
CH CA
0
σ σ 6 (1)
z1−α + μ 0 z β + μ 0 + δ =− z1−β + μ0 + δ
n n n
σ
(z1−α + z1−β ) n
δ
n σ (z + z )2 (2)
2
or δ 2 1−α 1−β
n corresponds with the required sample size to detect a desired difference with probability (1-) by a
z-test with significance level .
For the two-sided one-sample-test we only have to replace by /2, and we get:
n σ (z (3)
2
+ z )2
δ 2 1−α / 2 1−β
Note: is not replaced by /2!
In a similar (but somewhat more complicated) way we find lower bounds for the sample sizes n1 and n2
satisfying our conditions in the one-sided two-sample-test:
σ 2σ 1 + σ 22 (4)
n2 (z1− β + z1−α ) 2
δ 2
σ1
n1 n2
σ2
For equal standard-deviations (1 = 2 = ) we obtain from (4):
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:27 PM Page 40
Query:
"Holderbank" - Cement Course 2000
2σ 2 (5)
n 1 = n2 (z + z1−α ) 2
δ 2 1− β
In the two-sided two-sample-test we simply have to replace by /2 in (4) or (5).
The following table shows short rules for determining sample sizes when = = 5% and 1. = 2 =
in the two-sample-case. The numbers in parentheses refer to the formulae from which the rules are
derived.
one sample two samples
2
one-sided 11σ 11σ 2
n n 1 = n2
δ2 δ2
(2) (5; for 1 2 see (4))
2
two-sided 13σ 26σ 2
n n 1 = n2
δ2 δ2
(3) (5)
with = relevant difference to be detected with probability (1-)
= = 5%
Note:
The variance 2 is usually not known, but often some knowledge about 2 is available from former
experiments (standard deviation). Otherwise 2 may be estimated in a pilot study.
In the case of small sample sizes (t-test, Wilcoxon-test), add 5% to the calculated n.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION
6. DATA PRESENTATION AND INTERPRETATION
(Chambers, Cleveland, Kleiner and Mikey, 1983).
The problem of data presentation and interpretation is common to cement manufacturers and users. In
every stage of cement, aggregate, concrete production and application, data concerning materials,
equipment, energy consumption, market situation, costs, etc. are generated. All this information must
increase the knowledge about what happens in the process and market, thus providing the basis of
decision. Moreover, communication between supplier and consumer has to rely on this information. We
consider it, therefore, vital not only for optimum manufacture of cement and cement based products,
but also with respect to the mutual relation between manufacturer and user that appropriate attention is
given to data handling, i.e. presentation and evaluation.
We may differentiate between three levels of information:
data needed for direct process and quality control (routine decision, off- or on-line)
data required for day-to-day management decisions based on quality reports
additional data allowing long-term improvements and developments.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:28 PM Page 41
Query:
"Holderbank" - Cement Course 2000
With respect to presentation and interpretation of test results, the main problems for plant
management are:
how to organize the required information flow so that decision can be made at all levels of
competence
to reproduce and evaluate the corresponding data in a specific situation.
The relevant data of al1 domains cannot be made available without a functional information system.
But availability alone does not necessarily result in a rational decision based on the data. To achieve
this, two further aspects must be considered: first, the data should be presented in an intelligible form,
i.e. a high transparency of the results enables the management to recognize certain relationships or
critical results in time. Secondly, the decision maker must be aware of accuracy, significance and
relevance of the considered data in order to provide a realistic interpretation.
Availability of data
A quick availability of data does not necessarily require a fully integrated data based system with
electronic data processing, but is rather a matter of organization. Important data (e.g. for process and
quality control) should circulate with little loss of time. Consequently, an appropriate reporting system
(organization) must be established. Additional data should not disappear in some drawer where a later
retrieval is impossible or at least very inconvenient. To avoid such a disorder, all data are recorded in a
similar way, including a note on the circumstances of measurement and provenance of samples and
test results. Remarks about circumstances and provenance are used to judge the comparability of
different sets of observations. A standardised recording procedure makes data surveying easy and
simplifies a later change to electronic data processing.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION / 6.1 Intelligible
Presentation
1. Intelligible Presentation
The main aim of data presentation is transparency rather than secrecy, i.e. graphs should be employed
instead of tables. A graphical presentation gives a quick survey on relevant information such as
changes in time, extreme values, relationship between variables. A short survey on frequently used
graphs can be found in section 6.4. It is recommended to produce the graphs directly at the source of
the data in the form of tallies and/or control charts. A control chart may already be included in the
laboratory journal and reports, directly behind the columns for sample identification and test results.
The possibility of obtaining a quick survey by consulting a well presented graph will not only help the
management in its decision making, but transparency will also improve interest and motivation of the
personnel at any level of competence.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION / 6.2 Interpretation
2. Interpretation
First of all. we must know how much we can rely on the data.
Test results are never "true" values, but are rather subject to errors of three types (see also section 1):
random errors caused by sampling, imprecise measurement, environmental effects, etc.
systematic errors caused by bias in sampling or process measurement, by the use of inadequate
experimental design
gross errors caused by recording wrong or not comparable values.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:28 PM Page 42
Query:
"Holderbank" - Cement Course 2000
Usually, we expect 5 to 10% gross errors in a set of observations. Random errors may be denoted as
reproducibility and expressed as the corresponding standard deviation. In general, these errors are
underestimated. Reproducibility should be known for every important analyzing method. Systematic
and gross errors are difficult to characterize. They should be minimized by careful experimentation and
data handling. A periodic check for systematic errors should be done (calibration, comparison with a
standard, inter-laboratory test).
How to compare several data sets or groups of data?
Before any comparison is made, we have to seriously check the comparability of data. Often, data sets
differ in provenance of samples or circumstance of measurement, so that a comparison ma y be
impossible. Even the fact that sample 1 was measured by laboratory assistant Miller and sample 2 by
Brown will lead to a biased comparison if there is any relevant systematic error between the two
persons. If the data sets are comparable and we observe a certain difference, we have to ask the
following questions before taking any action:
a) Is the difference significant?
b) Is the difference relevant?
The problem of significance is answered by a statistical test. If it is not significant, we have no reason
to take any action because the observed difference may occur by chance alone. If the test indicates a
significant difference, it is not necessarily relevant for the problem we are concerned with. Of course,
the decision whether or not it is relevant is not a statistical problem. Perhaps a decision maker is
alarmed when an observed difference, considered to be relevant in the present problem, does not lead
to a significant test result. In this case, the sample size used may be too small or the testing
procedures may not be sufficiently sensitive to solve the given problem.
Consequently, experience in product manufacture and application is necessary to assess the
relevance of difference between target quality and experimental values. However, to judge whether it is
significant requires a training in statistical technique, particularly test procedures. Decisions based
simply on either practical experience of many years (relevance) or on statistical technique and
procedures (significance) will lead to too frequent and unnecessary actions in both the manufacturing
process and product application.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION / 6.3 Data
Interpretation related to Problems in Cement Application
6.3 Data Interpretation related to Problems in Cement Application
If confronted with the problem of finding the reasons for poor quality of cement related products
(concrete, asbestos cement, etc.), the cement is often suspected to be the cause. This can be
explained by the fact that cement is the binding agent, and in the majority of cases constitutes the most
expensive component, although it is known that the effect of other components, proportioning and
curing conditions etc. are of great importance. A further reason to inspect the cement first is the
difficulty to specify the other effects while cement is well defined by its chemical and mineralogical
composition and its fineness.
To find out the causes for inferiority or changes in quality, a collaboration of consumer, producer and
statistician is indispensable: the statistician may be omitted in uncomplicated problems if delegates of
consumer or producer are well trained in statistics. In a first retrospective analysis of available data,
parallel changes of parameters and quality are studied. This is done by drawing scatter diagrams and
performing a regression analysis. The disadvantage of retrospective studies is the difficulty to find out
causal relationships. The effects of several variables are mixed and cannot be separated due to their
causal origin. On the other hand, an observed correlation indicates a possible causal effect. The
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:28 PM Page 43
Query:
"Holderbank" - Cement Course 2000
decision about what cause is really responsible must be made by an experienced specialist and not by
a statistical test. Often a decision is not possible because dependencies are too complex. In this case,
a special experiment has to be planned and performed with a controlled variation of suspected
variables and careful elimination of interfering effects (prospective analysis). In contrast to the
retrospective analysis, an adequately designed experiment renders it possible to evaluate and judge
causal effects with statistical methods. A short survey on the use and interpretation in
regression/correlation analysis and designed experiments is given in chapter 7.
Survey on Graphical Data Representation
Representation of one sample - documentation
Frequency table / Tally
Given is a sample of large size n. Observations are grouped into classes of equal length and marked in
the tally.
Example: 90 values of concrete strength
Histogram
Graphical representation of the tally. Visualization of minimum, maximum, center and shape of the
distribution.
Area of rectangles corresponds to absolute or relative frequencies in the classes.
Frequency curve
Relative (or absolute) frequencies plotted against class mid point. Every point on the curve
corresponds to the relative (or absolute) frequency of observations falling in this class.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:29 PM Page 44
Query:
"Holderbank" - Cement Course 2000
Cumulative frequency curve
Cumulated relative frequencies plotted against upper class boundaries. Every point on the curve
corresponds to the portion of values which are smaller than any given strength x.
Example: Half of the measurements are smaller (respectively greater) than 33.5 N/mm2
Stem-and-leaf Diagram
Similar to the histogram, but allows to see the individual data. Each observation is a leaf of a stem.
Representation of small samples
For small sample sizes n the individual values are plotted directly on the measurement scale. Suitable
to detect outliers and skewed distribution.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:29 PM Page 45
Query:
"Holderbank" - Cement Course 2000
Example: K2O-content of 10 cements
Time-plot
Often it is appropriate to draw the observations in chronological order to show a possible change of the
level in time.
Useful in annual reports.
Cumulative time-plot
Cumulated values are plotted against time. Used to show deviations from a cumulative target.
Example: Actual clinker production is cumulated every month and compared with a target.
Box-Plot (Box-and Whisker-Plot)
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:29 PM Page 46
Query:
"Holderbank" - Cement Course 2000
Special and frequently used plot to numerize data. Shows median, lower and upper quartile (forming
the box) and two lines (the whiskers) extending to the extremes of the data. If unusual extreme values
occur, the whiskers extend only to those points that are within 1.5 times the inter-quartile range.
Scatter-plot or x-y plot
Plots two variables against each other.
Draftmensplot
Plots two or more variables against each other!
Flury-Riedevyl faces and Chemoff faces
Faces are a method to show multivariate data graphically. Each variable is mapped to one or more of
the face parameters. Other methods to represent multivariate data are the star symbol plot, the sun ray
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:29 PM Page 617
Query:
"Holderbank" - Cement Course 2000
plot, castles and trees.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION / 6.4 Control
Charts
6.4 Control Charts
Control charts are special time-plots to show a possible change of the characteristics in a production
process. The essence of a control chart is its clearness and intelligibility. In a quick survey it is possible
to recognize natural subgroups, within which variation is likely to be random but among which
assignable causes of one sort or another may cause non-random variation.
Summarized statistical errors (such as means or standard deviations) and confidence limits determined
from overall sample data can be very misleading, if the data are not free from the effects of assignable
causes. In cement industry annual means and standard deviations are mostly misleading due to a
nonstationarity of the process, i.e. the occurrence of systematic changes at the production level.
(Example: mean and standard deviation of the lime saturation of raw meal and/or clinker in annual
reports. Changes in the target value required by the process or the product lead to a high
non-interpretable standard deviation.)
If possible, the data should originally be collected with this subgrouping in mind; in such cases, the
analyses can be simplified by arranging for an equal number of observations in each subgroup. If the
data must be analyzed as they come, subgrouping may still be possible with knowledge of the data's
source, i.e. obvious changes in the process conditions must be registered together with the data.
Control limits
To judge a significant change at the production level, the charts are completed with control limits. As a
general rule it can be assumed that such a change is present, if a sample point falls outside of 3
-limits, where is the standard deviation of a homogeneous production phase.
This engineering rule has been found to work well. No exact probability is given for a chance variation
beyond 3 -limits, but in general it is very s mall, in the order of perhaps 0.3%.
It may be of advantage to use two limits, such as a 2 -warning-limit and a 3 -action-limit. In this case,
an action should already take place when two subsequent values lie between the two limits.
The use of 3 -limits bases on pure statistical considerations. It may occur that these limits are in
conflict with the required tolerances of an external standard (e.g. process, standard or market
requirements). If the required tolerances are smaller than the statistical ones, a modification of the
process is necessary to improve its precision. In the opposite case, the statistical limits may be used in
order to detect changes as early as possible.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:30 PM Page 618
Query:
"Holderbank" - Cement Course 2000
Control Charts
Several types of charts may be used to detect different types of changes in the process:
Control chart
Cause of Change Mean Range Standard Cumulative
X R deviations
Gross error (blunder) 1 2 - 3
Shift in average 2- - 1
Shift in variability - 1 2 -
Slow fluctuation (trend) 2 - - 1
Rapid fluctuation (cycle) -1 2 -
- = not appropriate / 3 = least useful
2 = useful / 1 = most useful
x − chart
Samples of fixed size n are taken from time to time. The arithmetic mean of the samples are plotted
versus time.
Control limits (for warning and/or action) can be computed with the help of special procedures.
R-chart of s-chart
Given are small samples of fixed size n. The range or the standard deviation of the sample is plotted.
Values exceeding a certain control limit indicate an increase of variability.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:30 PM Page 49
Query:
"Holderbank" - Cement Course 2000
Chart of individual values
Instead of means, individual values ma y be plotted in the order of measurement.
CUSUM-chart
Cumulative-sum-chart. The deviations of individual values from a target value are cumulated.
This chart is especially sensitive to slow fluctuations in the process.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 6. DATA PRESENTATION AND INTERPRETATION / 6.5 Comparative
representation
6.5 Comparative representation
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:30 PM Page 50
Query:
"Holderbank" - Cement Course 2000
Frequency polygon
Comparison of the distribution of two samples. Relative frequencies corresponding to the histogram
are plotted on a line graph against the mid-points of the classes.
Example: Mortar strength in the first and the second half-year.
Cumulative Frequency plot
Comparison of the distribution of two samples. Relative cumulative frequencies are plotted against
upper class limit.
Grouped Box-Plots
Shows two or more box-plots in the same graph. May be used to show the distribution of different
variables or the distribution of the same variable in different groups (samples). The notched box plot
may be used for the latter comparison.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:31 PM Page 51
Query:
"Holderbank" - Cement Course 2000
Time-plot
Comparison of simultaneous changes of several variables in an observed period.
Example: Silica- and alumina ratio of a cement type in a specific year.
Scatter-plot
Illustration of the relationship between two variables.
Example: Relation between strength and density of concrete. The relationship is not linear.
Don't use the correlation coefficient in this case, because it is only a measure for linear dependence.
Comparison of small samples
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:31 PM Page 52
Query:
"Holderbank" - Cement Course 2000
Individual values are marked on the scale separately for each group of observations.
Example: Inter-laboratory-test concerned with elite-content.
Asymmetrical Flury-Riedwyl Faces
This is a special version of the faces. E.g. the variables before and after treatment, respectively, are
mapped to the same face parameter on the left and right face side, respectively. 0r the left side shows
the specification values of the variables and the right side the actual measurements.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION
7. CORRELATION AND REGRESSION
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.1 Correlation coefficient
1. Correlation coefficient
The degree of linear dependence between two random variables X and Y can be expressed by the
correlation coefficient rxy.
sxy
rxy =
sx sy
n
with
s xy =
1
(x i − x )(y i − y ) the covariance of X and Y
n −1 i =1
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:31 PM Page 53
Query:
"Holderbank" - Cement Course 2000
and sx, sy standard deviations of X and Y.
For practical computations use
1
s xy =
n −1
x i y i ( x i )( y i )
1
n
The following figure shows the scatter diagrams for several degrees of dependence.
Properties of rxy:
a) rxy = ranges from -1 to +1
rxy = +1 : All measured values lie on an increasing line
-1 : All measured values lie on a decreasing line
0 : No linear relationship between the measured values
b) rxy is a measure of linear dependence. If the scatter diagram indicates a non-linear relationship,
then the correlation coefficient will be misleading and should not be calculated.
c) rxy is very sensitive (not robust) against outliers.
The necessity of drawing a scatter diagram is illustrated in the following figures. Completely different
graphs may result with equal correlation coefficients (r = 0.82).
rxy can be close to zero even though the variables are clearly non-linear dependent and rxy is not
defined if sx or sy is zero.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:32 PM Page 54
Query:
"Holderbank" - Cement Course 2000
Interpretation of rxy
A high correlation coefficient between two variables does not necessarily indicate a causal
dependence. There may be a third variable not under control which is causing the simultaneous
change in the first two variables, and which produces a spuriously high correlation coefficient. In order
to establish a causal relationship it is necessary to run a carefully controlled experiment (see chapter
8). Unfortunately it is often impossible to control all the variables which could possibly be relevant to a
particular experiment, so that the experimenter should always be on the lookout for spurious
correlation.
The following is an example for confusion of correlation with causation:
The following figure shows the population of Oldenburg at the end of each of 7 years plotted against
the number of storks observed in the corresponding year. Although in this example few would be led to
hypothesize that the increased number of storks caused the observed increase of population,
investigators are sometimes guilty of this kind of mistake in other contexts. Correlation between two
variables Y and X often occurs because they are both associated with a third factor W. In the stork
example, since the human population Y and the number of storks X both increased with time W over
this 7-year period, it is readily understandable that a correlation appears when they are plotted together
as Y versus X.
Fig.: A plot of the population of Oldenburg at the end of each year against the number of
storks observed in that year.
By using sound principles of experimental design and, in particular, randomization, data can be
generated that provide a more sound basis for deducing causality.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:32 PM Page 55
Query:
"Holderbank" - Cement Course 2000
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression
7.2 Linear Regression
If the scatter diagram indicates that the two variables are linearly related, then we may want to predict
the value of one of the variables from a given value of the other variable. For this purpose a regression
line is fitted to the data (method of least squares).
Model assumptions
y j = α + β j + E j (1)
where the residuals Ej's are independent normal random variables with mean 0 and constant variance
2.
= y -intercept
= slope
x = independent variable - Y dependent variable
This assumption is essential if parameter reductions (see later), i.e. reduced models are tested and the
assumptions have to be checked when a model is fitted.
The estimates of the parameters and in model (1) are obtained by the method of Least Squares
(MLS), i.e. the estimates are obtained by minimizing the sum of squared residuals
n n
(y j − α − βx j ) 2 = Eˆ j
2
j =1 j =1
with respect to and !
This method is very wide spread in statistics and the estimates have good statistical properties under
the above model assumptions! Today there are also procedures which do not minimize the sum of
squared residuals but use other criteria (Robust Methods!).
In the following let
1
s xx = (x j − x )2 = x j − ( x j )2 = (n − 1)sx
2 2
n
1
s yy = (y j − y )2 = y j − ( y j )2 = (n − 1)sy
2 2
n
= (x j − x )(y j − y ) = x j y j − x j y j = (n − 1)sxy
1
s xy
n
The resulting estimated regression line is
s
βˆ = xy
yˆ = αˆ + βˆ , where sxx
αˆ = y + βˆx
2
smin = syy − s xy
with a resulting minimal sum of squares (MSSQ) s xx with n-2 degrees of freedom (2
parameters and are estimated). yˆ is the predicted value for the dependent variable y.
Note:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:32 PM Page 626
Query:
"Holderbank" - Cement Course 2000
1) In the mentioned problem there are two regression lines, one to predict y from x and one to predict
x from y. The two lines are not identical and therefore it is not allowed to invert the regression
equation.
2) The prediction equation is valid only in the observed range of observations. Extrapolations may
give misleading results. We have no information that linearity holds outside of the present
observations.
For each pair (xj, yj) we can therefore compute the predicted value yˆ of y by
means of the regression function yˆ = αˆ + βˆx i.e.
yˆj = αˆ + βˆxj
and the estimated residual
ê j = y − yˆj
αˆ2 = Smin /(n − 2) is an estimate for the variance of the residuals. The analysis of the residuals ê j
gives us the possibility to validate the model assumptions (Independence, Normality and constant
Variability). The first check is done by plotting the residuals against the predicted yj. The residuals ê j
should be randomly scattered around the line e = 0 and show no pattern.
For example the following structures of the residuals would indicate, that the model is not correctly
specified (e.g. that the variables x and y should be transformed before calculation of the regression
line).
The second check would be a normal probability plot of the residuals.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:33 PM Page 57
Query:
"Holderbank" - Cement Course 2000
If the residuals suggest that the assumptions do not hold, then further investigations are necessary
(suitable transformations, other influencing variables, time-dependence => growth curves,
time-series-analysis).
After confirmation of the assumption and calculation of the regression line (sometimes also before
calculation) it is possible that we want to investigate whether perhaps the y-intercept equals zero, or
the slope equals zero or whether there is no relation between y and x ( = = 0) or whether the or
equal some specified values 0 or 0 (e.g. if a value of βˆ = 1.1 may be replaced by β 0 = 1.0 ). This
means that we want to test if a reduced (simplified) model is sufficient to describe the relationship
between y and x. For this it is essential that the model assumptions hold.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.1
Regression line with slope 0:
1. Regression line with slope 0:
Model: yj = + Ej
Estimate αˆ for : αˆ = y (= mean of the yj's)
Minimal Sum of Squares (MSQ): Smin = Syy
Degrees of freedom (df): df = n-1
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.2
Regression line with y-intercept 0
2. Regression line with y-intercept 0
Model: yj = xj + Ej
βˆ = x yj j1
Estimate βˆ for : x
2
j
( x j y j )2
Smin = y j 2 −
x
2
MSQ: j .
df = n-1
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.3
Intercept = 0. Slope = 0
7.2.3 Intercept = 0. Slope = 0
Model: yj = Ej
no parameters to estimate
MSQ = y j 2
df = n
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:33 PM Page 58
Query:
"Holderbank" - Cement Course 2000
Other models:
y = α 0 + βx
y = α + β0 x
y = α 0 + β0 x
The estimates and MSQ can be obtained by differentiation of the corresponding Sum of Squared
Residuals!
These special models are only useful if the data really speaks for them!
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.4
Comparison of models
7.2.4 Comparison of models
This is done by performing an Analysis of Variance (ANOVA). ANOVA is a widely spread statistical
technique. (Comparison of more than two means, Experimental Designs etc.). ANOVA compares the
Minimal Sum of Squares of the models to each other. The basic results for the comparison are filled in
a table, the ANOVA-Table.
ANOVA-TABLE
Model MSQ df
H0 S0 min df0
A Smin df
Reduction S0min - Smin df0 - df
Ho denotes the null hypothesis, A the alternative model. By H0 the test-statistic
(S0min − S ) (df 0 − df )
F= min
Smin df
is distributed according to a F-distribution with m1 = (df0 - df) and m2 = df degrees of freedom.
Significance limits: Look up F1−α;(df 0 −df ),df in Table A-9 with m1 = (df0-df) and m2 = df degrees of
freedom.
Decision: If F > F1−α;(df 0 −df ),df then conclude that the simplification to the null-model H0 is not permitted,
therefore the alternative model has to be used; otherwise (FF1-) there is no evidence that the
null-model should be rejected.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.5
Standard deviation of the estimates
7.2.5 Standard deviation of the estimates
The ANOVA of the null hypotheses H’0 : y = and H'’0 : y = x, resp. against the alternative A : y = +
x allows us to compute the standard deviations (sd) of the estimated parameters αˆ and βˆ in the
alternative model A : y = + x: let F( = 0) be the calculated F-statistic for the test H’0 against A (e.g.
ANOVA for = 0) and correspondingly F( = 0) for H’’0 against A.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:34 PM Page 59
Query:
"Holderbank" - Cement Course 2000
Then
αˆ
sd(αˆ) =
F(α = 0)
βˆ
sd(βˆ) =
F(β = 0)
so that the regression line of the alternative (full) model is
y = αˆ + βˆ x
(sd (αˆ )) (sd ( βˆ ))
In order to obtain the standard deviation in a model A : y = x, this model has to be compared with the
null-model H0 : y = 0.
For A : y = , compare A to H0 : y = 0
For A : y = 0+ x, compare A to H0 : y = 0
For A : y = +0x, compare A to H0 : y = 0x
1- confidence intervals for the estimated parameters can be obtained by calculating
αˆ t1−α / 2,msd(αˆ)
and
βˆ t1−α / 2,msd(βˆ)
where m equals the degrees of freedom in the alternative model and t1_a/2,m can be looked up in Table
A-4.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.6
Coefficient of determination
7.2.6 Coefficient of determination
The prediction of y with a regression line is more or less accurate dependent on the degree of linear
dependence. How well does the regression line fit the data? A measure to express the relative
accuracy of prediction compared with the total variation of y is the coefficient of determination r2. In the
case of a regression line the coefficient of determination is the square of the correlation coefficient
s2xy S2xy
r 2 = rxy2 = =
s2x s2y S xx S yy
The total variation of y can be partitioned into two components. Total variation = explained variation +
unexplained variation
where
total variation = (y i − y )2
unexplained variation = (y i − ŷ )2
(observed - predicted)
explained variation
= (yˆ − y ) 2
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:34 PM Page 60
Query:
"Holderbank" - Cement Course 2000
The coefficient of determination is the ratio
explainedvariation
r2 =
totalvariation
In this form r2 is defined also for non-linear curves and multiple regression.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.7
Transformations before Regression Analysis
7.2.7 Transformations before Regression Analysis
Kind of transformation Type of relation Regression after
transformation
y' = ln y, x(y 0) y = αe βx y' = ln+ βx
(exponential)
y, x' = ln x,(x 0) y = α + β ln x y = α + βx'
(logarithmic)
y' = ln y, x' = ln x,(x, y 0) y = αx β y' = lnα + βx'
(exponentiation)
y' = 1 y , x(y 0) y = (α + βx)−1 y' = α + βx
y, x' = 1 x (x 0) 1 y = α + βx'
y =α + β
x
y' = 1 y , x' = 1 x (x, y 0) x y' = α + βx'
y=
αx + β
(hyperbolic)
y' = ln y, x' = 1 x (y 0,x 0) y = αe β x y' = lnα + βx'
y' = 1 y , x' = e− x (y 0) 1 y' = α + βx'
y=
α + βe−x
Model Overview for Linear Regression
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:35 PM Page 631
Query:
"Holderbank" - Cement Course 2000
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 7. CORRELATION AND REGRESSION / 7.2 Linear Regression / 7.2.8
Multiple and non-linear regression
7.2.8 Multiple and non-linear regression
Often the variable to be predicted is not only dependent on one but on several independent variables.
Model:
Yj = β 0 + β1x1j + β 2 x2 j ++ β p xpj + E j
where Ej are independent normal random variables with mean 0 and constant variance 2. Again these
assumptions have to be checked after a model has been fitted.
In this case prediction can be improved in a multiple (or multivariate) regression expressed as
yˆ = βˆ0 + βˆ1x 1+ βˆ 2x 2 + βˆ3x 3 + + βˆ xp p
Or the scatter diagram shows a non-linear relationship which demands to fit a non-linear curve to the
data. Often a logarithmic or some other transformation of one or both variables may lead to a linear
relationship, and a regression line can be fitted to the transformed data. In more complex situations
further methods are available.
For multiple, non-linear and mixed (multiple/non-linear) procedures consult literature (ref. Chatfield
1975)
Flury + Riedwyl (1988) give a very good praxis oriented introduction to multiple linear regression and
multivariate analysis, including Discriminant Analysis, Principal Components, Identification and
Specification analysis.
Example 6
The following data represent
y: water requirement (%) and
x: grain fraction 10-32 (%)
of n = 34 different cements.
x y x y x y x y
32.9 25.0 38.6 27.2 29.8 23.3 37.2 26.8
35.4 26.3 32.5 25.4 30.3 25.6 42.2 25.6
33.1 24.9 31.3 25.1 36.0 25.2 35.8 24.6
38.7 27.9 35.1 26.9 30.2 24.5 39.1 25.2
33.8 24.8 32.2 24.4 35.6 24.0
32.9 25.0 31.5 24.8 40.6 26.4 40.1 26.6
29.6 24.0 31.1 25.4 40.3 25.2 43.2 27.4
29.7 25.0 38.8 25.6 38.1 25.5 36.7 24.5
36.1 26.2 42.1 27.0 34.7 25.8
The scatter diagram shows a positive linear relationship between water requirement and grain size.
Regression lines to predict x or y are clearly distinct. The lines are identical if rxy is 1 or -1.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:36 PM Page 632
Query:
"Holderbank" - Cement Course 2000
If r = 0 the lines are perpendicular, i.e. the best prediction of y is the arithmetical mean y . In the same
way x is the best prediction of x if no dependence exists.
Computations
In a first step five auxiliary sums are provided:
A = xi = 1'205.3 n = 34
B = x 2
= 43'253.19
i
C = yj = 867.1
D = y2 = 22'150.93
i
E = xjyj = 30'828.44
1 AC
s xy = E− = 2.719545
n −1 n
1 A2
s x2 = B − n = 15.918333
n −1
1 C2
s y2 = − = 1.131203
n
D
n −1
x = A n = 35.45
y = C n = 25.05
Sxy = 89.7550
Sxx = 525.3050
Syy = 37.3297
Smin = 21.9973
df = 32
σˆ 2 = 0.6874
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:37 PM Page 632
Query:
"Holderbank" - Cement Course 2000
σˆ = 0.8291
Sxy
rxy = = 0.64
Correlation coefficient: Sxx Syy
S
βˆ = xy = 0.171
Regression line: Sxx
αˆ = y − bx = 19.45
Test: = 0
Model MSQ df
H0 : y = x 178 l527 33
A : y = + x 21.9973 32
Reduction 156.1553 1
156.1553
F= = 227.16 = F(α = 0)
21.9973 32
F.95;1,32 = 4.17 (Table A-9)
→ Simplification not permitted
Test = 0
Model MSQ df
Ho : y = 37.3297 33
A : y = + x 21.9973 32
Reduction 15.3324 1
15.3324
F= = 22.304 = F(β = 0)
21.9973 32
F95,1,32 = 4.17 (Table A-9)
→ Simplified model not allowed !
We obtain the following standard deviations for the parameter estimates αˆ and βˆ
aˆ 19.45
sd(αˆ) = = = 1.2905
F(α = 0) 227.16
βˆ 0.171
sd(βˆ ) = = = 0.0362
F(β = 0) 22.30
and the regression line is:
y = 19.45+ 0.171x
(1.29 ) (0.036 ) σˆ = 0.8291
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:37 PM Page 64
Query:
"Holderbank" - Cement Course 2000
95% confidence intervals for αˆ and βˆ
From table A-4: t.975,32 = 2.042
αˆ :19.45 2.042 * 1.29 = (16.82, 22.08)
βˆ : 0.171 2.042 * 0.036 = (0.097, 0.245)
Coefficient of determination: r2 = 0.41
Comment:
Only 41% of the total variation of Y can be explained by the dependence of the grain fraction. The
prediction can be improved when further variables are considered in a multiple regression. The
coefficient of determination increases to 69% if C3A- and alkali-content is included in the equation
y = 17.44 + 0.15x1 + 0.26 x2 + 0.52x3
x1 = grain fraction 10-32%
x2 = C3A- content %
x3 = K20 + Na2O%
The improvement is visible in a scatter diagram of observed against predicted values.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN
8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN
In practice numerous problems and questions cannot be answered unless special quantitative
investigations are performed. This applies to every practical field, starting with the raw material
exploitation and ending with the sale of the finished product. In many cases, it is not possible to attain
complete information as an effective basis for decision making. For this reason, it is becoming more
and more common to resort to statistical investigations in form of sample surveys and experiments.
The proper performance and interpretation of statistical investigations is essential. The so-called “lies
of statistics” generally have to be attributed to the wrong selection of data, a misinterpretation, or
procedures which do not relate to the objective. Experience has shown that very rarely data are
manipulated intentionally. However, people who employ “wrong statistics”, are convinced, due to a lack
of in-depth knowledge, that their argumentation is objective and correct.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN / 8.1 The Five
Phases of a Statistical Investigation
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:37 PM Page 65
Query:
"Holderbank" - Cement Course 2000
8.1 The Five Phases of a Statistical Investigation
Despite the variety of potential applications, a statistical investigation can roughly be divided into five
phases in the following order:
1) Formulation of problem:
exact definition of purpose for which information is acquired
2) Planning:
purpose-oriented planning of investigation according to statistical principles with the objective to
acquire optimum information at given expenses. In case of doubt and problems contact an
experienced statistician.
3) Performance:
procurement of data strictly in accordance with the established planning
4) Evaluation:
summary and presentation of results; inference from sample to population
5) Conclusions:
realistic interpretation of results
Essential basic principle:
Statistical investigations will only provide truly efficient results if knowledge in the investigated field is
optimally combined with thorough statistical knowledge. Usually this will lead to teamwork, since the
investigator who lacks extensive knowledge in statistics is usually just as unsuccessful in his attempt to
carry out the investigation on his own as the statistician who is entrusted with the entire problem
complex.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN / 8.2 Sample
Surveys and Experiments
8.2 Sample Surveys and Experiments
Statistical investigations are subdivided into two groups sample surveys and experiments - which differ
significantly with respect to their objective and interpretation.
Experiment
In an experiment response variables are investigated, which depend on various factors. The results are
produced by a controlled variation of the factors of primary interest, whereas all other effects on the
response have to be eliminated.
Factors that are being studied may be quantitative (e.g. temperature, concentration) or qualitative (e.g.
type of ample preparation, origin of samples).
In an experiment, causal effects on the response variable can be determined quantitatively if the
necessary provisions have been made in the planning and performance.
Sample survey
Samples are taken from one or several populations (the set of sample units is called sample of size n).
Certain interesting characteristics are measured in order to evaluate (estimate)unknown characteristics
of the populations or to compare populations. Contrary to the experiment, the results of a sample
survey do not permit any conclusions regarding cause and effect. This is often ignored when
correlations between two characteristics are assessed.
Reliable results of both experiments and sample surveys can only be obtained if the respective
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:37 PM Page 66
Query:
"Holderbank" - Cement Course 2000
statistical principles have been adhered to in the planning and performance of the investigation.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN / 8.3
Fundamental Principles in Statistical Investigations
3. Fundamental Principles in Statistical Investigations
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN / 8.3
Fundamental Principles in Statistical Investigations / 8.3.1 Experiments
1. Experiments
The five criteria of a good experiment are:
1) The experiment should serve a well defined purpose
In a first step the problem has to be clearly identified for each experiment, and the hypotheses to
be investigated have to be determined.
2) Factors. which are not of primary interest. should not influence the results
Influencing factors, which are not included in the investigation, must be under control at fixed
levels. Otherwise several effects become mixed up and cannot be separated by statistical methods.
Systematic effects of factors, which cannot be controlled, are eliminated by random allocation of
samples to treatments or factor levels and by random measuring sequence (randomization).
Practical performance of randomization: In example A 1 we assign a number to each sample, No. 1
- 8 to tablets and No. 9 - 16 to beads. Each number is written on a leaflet and mixed in a box. The
order of measurement is given by blind drawing of the leaflets from the box.
3) The experiment should be free of systematic errors
This requirement is partially connected with 2). Systematic errors, due to any change of effects in
time, are eliminated by randomization. An often underestimated source of systematic errors is the
prejudice of the experimenter. In order to avoid this, mainly “blind" experiments should be
performed. The samples are coded with random numbers, the decoding key only known to a
confidential person, who herself is not in a position to perform analyses.
4) The experiment should provide a measure of its precision
An estimate of the precision is obtained by a replication of the experiment. For each step of the
experiment several measurements should be taken. In example A 1 it would be impossible to state
whether a systematic difference exists, if only one tablet and one bead are measured. (Appendix I)
5) Precision and efficiency of an experiment should be high enough to reach the set goals
There are various measures to improve the precision and efficiency of an experiment:
• Reduction of variability by using homogeneous materials and by carefully controlling all the
factors, as well as by strictly observing the analytical regulations.
• Increasing the number of replications; this will lead to the problem of determining the sample
volume required to achieve the desired precision.
• Blocking: Measurements are performed in homogeneous groups (blocks). Such blocks may be
'measurements made by the same operator' or 'measurements made on the same day'.
A special blocking procedure is the paired comparison. In example A 3 both laboratories
measure the compressive strength on samples drawn from the same cement bag.
Block experiments should be balanced and randomized, i.e. the number of measurements and
treatments is equal in each block and assignment of samples or treatments to blocks is
random.
Special cases of blocking: paired comparison.
• Analysis of covariance: factors, which are not subject to the experimenter's control, are
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:37 PM Page 67
Query:
"Holderbank" - Cement Course 2000
recorded in order to eliminate their effects in a later analysis.
The principles above are valid for any experimentation. Special designs for various problems are given
in the literature. Problems may be: Effects of one or several factors on a response, additivity of factor
effects, splitting of components of variability in a measurement procedure in order to achieve a
prescribed precision at minimum cost, inter-laboratory tests, calibration, evaluation of systematic errors
etc.
Note:
In contrast to a regression analysis with a given set of observations, a carefully designed experiment
allows the evaluation of causal relationships. In regression the effects are usually mixed and can only
be separated due to some mathematical model and not due to their real origin (ref. Interpretation of rxy
section 7.1).
STRUCTURE OF AN EXPERIMENT
1) Recognize problem
2) Describe problem in detail and determine hypotheses
3) Define experimental area; contact statistician; determine factors, levels and number of replications;
elaborate exact analytical regulations; determine a reference basis if required. It is often necessary
to analyse what happens if no treatment is applied (e.g. Placebo in clinical experiments).
4) Establish experimental design, taking into account methods to improve precision (blocking).
Randomization:
- random selection of sample units
- random allocation of treatments
- coding of samples
5) Specify variables, which cannot be maintained constant (covariables)
6) Determine number of replications
7) Determine procedural organization
8) Determine procedures of statistical data analysis; exact regulation concerning the manner in which
data are to be supplied (form)
9) Perform experiment
10) Evaluate results (statistical data analysis)
11) Draw conclusions
12) Take measures
Step 2) to 7) are designated as experimental planning. The results of this planning should absolutely
be recorded in an experimental plan. Since experiments are usually quite expensive, careful planning
will definitely be worthwhile.
The request to prepare a written experimental plan entails significant advantages:
the necessity to formulate statements in precise terms
clear conditions and thus less uncertainty in the performance of the experiment
the circumstances of the experiment can be reconstructed, if later on the results are used for
comparison with new results.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 8. STATISTICAL INVESTIGATIONS AND THEIR DESIGN / 8.3
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:39 PM Page 638
Query:
"Holderbank" - Cement Course 2000
Fundamental Principles in Statistical Investigations / 8.3.2 Sample Surveys
8.3.2 Sample Surveys
Sample surveys and experiments are performed in an analogous way. Differences are due to the
different situation. While the purpose of an experiment is to directly produce results, sampling is carried
out to analyse existing characteristics of a population.
Special attention should be paid to the following points:
1) Target population and sampled population
The target population is the aggregate about which the investigator is trying to make inferences
from his sample. It is usually helpful to focus the attention on differences between the population
actually sampled and the population that is attempted to be studied.
2) Sample unit
The sample units have to be described in clear terms. They should not overlap and the sum of all
the sample units must be equal to the investigated population.
In the cement production for instance, if an individual sample is analysed, the results and
interpretations are different from those of an analysis performed on a daily composite sample.
3) Representativity and sampling method
Samples only supply unbiased results if they represent the entire population. It is often mistakenly
believed that the large sample size n ensures its representativity, or it is simply maintained that the
sample is representative in order to prevent rejection of the results.
Representativity can only be ensured by an appropriate, random sampling method, giving each
sample unit an equal chance to be included in the sample (exceptions in special cases: sampling
with unequal probabilities, systematic sampling).
In literature a number of selection procedures are described. Depending upon the problem situation,
various methods ma y be considered. They differ in their practicability, financial consequences
(expenses) and efficiency.
In most cases an optimum selection procedure for a specific problem situation can be found.
STRUCTURE OF A SAMPLE SURVEY
1) Recognize problem. What information is required about what populations?
2) Detailed definition of problem
- definition of target populations
- what information with what accuracy?
3) Determine the sample population. Where, when and how are the sample units extracted from the
population? Determine the variables to be recorded
4) Determine sampling method (equal probability sampling, stratified sampling, cluster sampling,
multi-phase sampling, multi-stage sampling, systematic sampling, unequal probability sampling)\
5) Determine sample scheme (in accordance with the required accuracy)
6) Outline procedural organization
7) Determine procedures of statistical data analysis; exact regulation concerning the manner in which
data are to be supplied
8) Perform sampling
9) If possible perform measurements in random sequence
10) Evaluation and presentation
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:39 PM Page 69
Query:
"Holderbank" - Cement Course 2000
11) Conclusions
12) Take measures
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK
9. OUTLOOK
This section will give a short outlook to further and perhaps more sophisticated statistical applications.
It also contains references for further readings. All these applications require more than elementary
knowledge of statistical methods.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.1 Time series and growth curves analysis
1. Time series and growth curves analysis
This is the analysis of chronological observations, (e.g. daily temperatures at a certain location or
observations over a certain time period of a patient obtaining a certain medicament). The characteristic
of these data are that the observations are not independent from each other and the method of linear
regression cannot be applied.
Literature:
Box, G.E.P. and Jenkins, G.M. (1976). Time Series Analysis, Forecasting and Control, second edition.
San Francisco: Holden-Day, Inc.
Cox, D.R. and Lewis, P.A.W. (1966). The Statistical Analysis of Series of Events. London: Methuen.
Nelson, C.R. (1973). Applied Time Series Analysis for Managerial Forecasting. San Francisco:
Holden-Day, Inc.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.2 Categorical and Qualitative Data Analysis
2. Categorical and Qualitative Data Analysis
This is the analysis of counts. Many experiments contain qualitative and nonmetric variables, e.g. when
analyzing the number of smokers and non-smokers beyond male and female persons. Smoking
behaviour and sex have no natural numeric values associated with them. Used statistical methods for
such data are
Crosstabulation
Contingency tables
Goodness of fit tests
Log linear models
Logistic regression
Correspondence analysis
Literature:
Agresti, A. (1984). Analysis of ordinal categorical data. New York: Wiley-Interscience.
Bishop, Y.M.M., Fienberg, S.E. and Holland, P.W. (1975). Discrete Multivariate Analysis: Theory and
Practice. Cambridge, MA: MIT Press.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:39 PM Page 70
Query:
"Holderbank" - Cement Course 2000
Haberman, S.J. (1978). Analysis of Qualitative Data, Vol. 1: Introductory Topics. New York: Academic
Press.
Haberman, S.J. (1979). Analysis of Qualitative Data, Vol. 2: New Developments. New York: Academic
Press.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.3 Experimental Designs and ANOVA
3. Experimental Designs and ANOVA
This is a wide field containing among others:
One-Way ANOVA
Multifactor ANOVA
Analysis of Nested Designs
Full and Fractional Designs
Response Surface Analysis
Covariance Analysis
and a lot more.
Literature:
Box, G.E.P., Hunter, W.G. and Hunter, J.S. (1978). Statistics for Experimenters. New York: Wiley.
Neter, J. and Wassermann, W. (1974). Applied Linear Statistical Models. Homewood, Illinois: Richard
E. Irvin, Inc.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.4 Multivariate Methods
9.4 Multivariate Methods
This title involves the analysis of multivariate data. It is not appropriate to analyze multivariate data
univariate, because correlations among the different variables are not taken into account.
Statistical Methods:
Correlation Analysis
Covariance Analysis
Multiple Regression
Principal Components
Factor Analysis
Discriminant Analysis
Cluster Analysis
Canonical Correlations
Multidimensional Scaling
and a lot more.
Literature:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:40 PM Page 71
Query:
"Holderbank" - Cement Course 2000
Sieber, G.A.F. (1984). Multivariate Observations. New York: Wiley.
Flury, B. and Riedwyl, H. (1988). Multivariate Statistics. A practical approach. London, New York:
Chapman and Hall.
Johnson, R.A. and Wichern, D.W. (1982). Applied Multivariate Statistical Analysis. London:
Prentice-Hall.
Morrison, D.F. (1976, 2nd edition). Multivariate Statistical Methods. New York: McGraw-Hill.
Everitt, B.S. (1980). Cluster Analysis, 2nd edition, London: Heinemann Education Books, Ltd.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.5 Nonparametric Methods
9.5 Nonparametric Methods
Parametric methods base on certain assumptions on the data (e.g. normality of residuals in linear
regression, normality of the observations in testing etc.). If these assumptions do not hold, it is often
more efficient to use methods that do not use a particular underlying distribution function, so called
nonparametric methods (e.g. the Wilcoxon Test known for Section 5.2.4 is a nonparametric test).
Some nonparametric procedures exist for:
Tests of binary sequences
Tests for randomness
Tests for location
Comparison of 2 samples
Rank correlation analysis
Goodness of fit tests
Literature:
Conover, W.J. (1980). Practical nonparametric statistics, 2nd edition. New York: John Wiley and Sons,
Inc.
Lehmann, E.L. (1975). Nonparametrics. San Francisco: Holden-Day, Inc.
Gibbons, J.D. (1976). Nonparametric Methods for Quantitative Analysis. New York: Holt, Rinehart and
Winston.
Hollander, M. and Wolfe, D.A. (1973). Nonparametric Statistical Methods. New York: Wiley.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.6 Bootstrap and Jack-knife Methods
9.6 Bootstrap and Jack-knife Methods
In dealing with complicated functions of data, it is mostly not possible to derive the underlying
distribution function. Bootstrap and Jackknife sometimes provide the possibility to derive the
distribution functions, statistical measures and a lot more at such complicated data.
Literature:
Efron, B. (1982). The Jackknife, the Bootstrap and Other Resampling Plans. Philadelphia: Soc. for
Industrial and Applied Math.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.7 Simulation and Monte Carlo Method
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:40 PM Page 72
Query:
"Holderbank" - Cement Course 2000
7. Simulation and Monte Carlo Method
By use of simulation and Monte Carlo methods it is often possible to investigate and analyse complex
functions of random variables or systems of events. Simulation incorporates also the generation of
random variables.
Literature:
Fishman, G.S. (1978). Principles of Discrete Event Simulation. New York: John Wiley & Sons, Inc.
Rubinstein, R.Y. (1981). Simulation and the Monte Carlo Method. New York: John Wiley & Sons, Inc.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 9. OUTLOOK / 9.8 General Literature
8. General Literature
Seber, G.A.F. (1984). Multivariate Observations. New York: Wiley.
Belsley, D.A. Kuh, E. and Welsh, R.E. (1980). Regression Diagnostic: Identifying Influential Data and
Sources of Collinearity. New York: Wiley.
Chambers, J.M., Cleveland, W.S., Kleiner, B. and Tukey, P.A. (1983). Graphical Methods for Data
Analysis. Boston: Duxbury Press.
Flury, B. (1980). Construction of an asymmetrical face to represent multivariate data graphically. Tech.
Rep. No. 3, University of Berne, Dept of Statistics.
Schupbach, M. (1984). ASYMFACE - Asymmetrical Faces on IBM-PC. [Link]. No.16. University of
Berne, Dept of Statistics.
Draper, N.R. and Smith, H. (1981), 2nd ed.). Applied Regression Analysis. New York: Wiley.
Grant, E.L. and Leavenworth, R.S. (1980). Statistical Quality Control, fifth ed. New York: McGraw Hill.
Montgomery, D.C. (1985). Introduction to Statistical Quality Control. New York: Wiley.
Recommended literature pp. 90 ff
Huff, D. (1974). How to lie with statistics.
Miller, R.G. (1981). Simultaneous Statistical Inference, 2nd ed. New York: Wiley.
Morrison, D.F. (1983). Applied Linear Models. Englewood Cliffs, NJ: PrenticeHall Inc.
Materials Technology / B02 - MT II / C13 - Statistics / Statistics / 10. STATISTICAL PROGRAM PACKAGES
10. STATISTICAL PROGRAM PACKAGES
Very complete and sophisticated packages are (all programs are available for PC's under DOS):
BMDP Statistical Software Ltd., Cork Technology Park. Cork,
Ireland (phone: 021-542722)
SAS: SAS Inst. GmbH, Cary NC, USA
(phone: (919)467-8000)
SPSS SPSS Inc., Chicago Il, USA (phone: (312)329-3300)
SYSTAT SYSTAT Inc., Evanston Il, USA (phone: (312)864-5670)
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:40 PM Page 73
Query:
"Holderbank" - Cement Course 2000
STATGRAPHICS STSC Inc., Rockville MD, USA (phone: (301)984-5000)
A review of 49 statistical packages is given in the PC magazine, Vol. 8, No. 5, March 1989.
Appendix I
Example A1:
In the application of XRF-analysis of raw meal it was required to know to what extent the results are
dependent on sample preparation, by pressed powder tablets and by fusion with Litetraborate,
respectively.
With 8 preparations of tablets and beads of the same raw meal the following results for CaO and Si02
were obtained.
CaO SiO2
Tablets Beads tablets beads
42.63 42.42 13.87 14.15
42.59 42.44 14.01 13.86
42.63 42.45 14.04 13.92
42.80 42.42 14.09 13.83
43.02 42.59 14.26 14.01
42.61 42.49 13.70 13.78
42.40 42.62 13.89 14.24
42.79 42.26 13.77 13.88
We are interested in differences between the preparation methods. To get a quick survey on the data
we mark every observation on the measurement scale for tablets and beads respectively.
Looking at the graph we suppose a significant difference of CaO-results between
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:40 PM Page 74
Query:
"Holderbank" - Cement Course 2000
beads and tablets. For Si 02 results no difference is visible.
The final decision whether the differences are significant is made with an appropriate test.
Test situation:
We test the hypothesis that the mean results of both preparation methods are the same
H0: tablets = beads
Two independent samples with small sample sizes N = 8
two-sided test: we have no idea in what direction a possible difference may occur before the
observations are taken.
Test procedure: Wilcoxon-test at = 5% level
As an alternative representation we can mark the data on both side of one line. In this case ranks of
the observations can be read directly from the graph.
The ranksum of beads is R = 45.5
W = 2 . 45.5. - 8 . 17 = -45
|W| = 45 => 38 = W0.975 (from Table A-8, Appendix III).
Decision:
The difference of CaO results between tablets and beads is significant.
|W| = 1 < 38 = W0.975
There is no significant difference for SiO2 results.
Example A2:
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:41 PM Page 75
Query:
"Holderbank" - Cement Course 2000
Control of homogenization efficiency.
In intervals of 30 minutes, 20 samples are taken from the raw ma serial stream both before and after
homogenization.
Measured values: CaCO3 content
Total capacity of homogenization silo: 1000 tons
Sample No. CaCO3 before CaCO3 after
homogenization homogenization
1 76.4 76.9
2 76.7 77.0
3 77.5 77.0
4 77.5 76.8
5 76.5 76.9
6 76.0 76.8
7 76.8 76.9
8 77.0 77.0
9 77.1 77.0
10 77.2 76.8
11 77.2 76.9
12 77.1 76.9
13 76.9 76.9
14 76.7 77.0
15 76.8 77.0
16 77.0 76.0
17 77.1 76.8
18 77.0 76.9
19 76.7 77.0
20 76.8 76.9
The graph is chronological order (time-plot) shows a systematical variation of CaCO3 content before
homogenization. Except for one outlying value the variation is small after homogenization.
Before Homogenization
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:41 PM Page 76
Query:
"Holderbank" - Cement Course 2000
After Homogenization
Using the Dixon criterion to check for outliers we compute
x(3) − x(1) 76.8 − 76.0
r22 = = = 0.8
x(n−2) − x (1) 77.0 − 76.0
with x(1) the smallest, X(3) the third smallest and X(n-2) the largest value.
Since r22 = 0.8 > r0.995 = 0.562 (Table A-2, Appendix III) the extreme value is considered to be a real
outliner (with significance level = 1%). It is deleted for further analysis.
Statistical data description
CaCO3 before CaCO3 after homogenization
homogenization
n 20 (20) 19
x 76.9 (76.87) 76.92
x 76.95 (76.9) 76.9
s 0.36 ( 0.22) 0.076 one outlier deleted: 76.0
v 0.0046 ( 0.0028) 0.0010
xmax 77.5 (77.0) 77.0
xmin 76.0 (76.0) 76.8
R 1.5 ( 1.0) 0.2
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:42 PM Page 77
Query:
"Holderbank" - Cement Course 2000
Comment:
Mean and median are almost equal. The distribution of observations seems to be therefore
symmetrical about the mean either before and after homogenization. Due to homogenization the
standard deviation is reduced from 0.36 to 0.08 corresponding to a factor of four to five.
Note:
Elimination of the outliner reduces the standard deviation from 0.22. to 0.08 !
Example A3:
In order to control mortar strength a national control laboratory takes every month a sample of cement
in a plant. It was supposed that these control results tend to be lower than results of internal quality
control.
To find out a suspected syste ma tic error between results of control office and plant laboratory during
a year the samples were measured in both laboratories:
Plant x Control y Difference
d=x-y
January 548 526 22
February 540 524 16
March 574 591 -17
Apri1 469 477 -8
May 540 531 9
June 515 431 84
July 520 485 35
August 531 476 55
September 530 490 42
October 464 452 12
November 524 498 26
December 519 516 3
The strength values are paired. Each pair of observation is concerned with the same sample. To test
for a significant difference we use the singed-rank test at a significance level of = 5% (two-sided)
H0: = 0
Plot of differences di = xi - yi
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:42 PM Page 78
Query:
"Holderbank" - Cement Course 2000
Sum of signed ranks:
T = (-6) + (-2) + 1+ 3+ 4+ 5+ 7+ 8+ 9+ 10+ 11+ 12 = 62
Since T = 62 > 52 = T0 975 we decide that the difference between results is systematic.
Example A4:
The following data represent 115 measurements of Schmidt-haamer-strength (abutment of a bridge).
A detailed interpretation of these data requires their representation in a frequency table and histogram.
With the probability paper we check the data for normality.
The cumulative frequency curve plotted on the probability paper is not linear and the distribution is
therefore not normal.
Histogram and tally show that even values of Schmidt-hammer strength are more frequent than odd
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:43 PM Page 79
Query:
"Holderbank" - Cement Course 2000
values. This seems to be caused by a reading error of scale. Probably only even values are marked on
the measurement scale.
EXTENDED PROBABILITY PAPER
Table A-1 Standard Normal Distribution - Values of P
Values of P corresponding to zp for the normal curve.
z is the standard normal variable. The value of P -zp equals one minus the value of P for +zp, e.g. the P
for -1.62 equals 1-0.9474 = .0526.
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:43 PM Page 80
Query:
"Holderbank" - Cement Course 2000
Table A-2 Critical values for the Dixon Criterion
Table A-3 Critical values for Outlier-tests, (large sample size n)
Table A-4 Critical values of the t-distribution
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:44 PM Page 81
Query:
"Holderbank" - Cement Course 2000
Table A-5 Critical values of the x2-distribution
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:44 PM Page 82
Query:
"Holderbank" - Cement Course 2000
Table A-6 Values of λ1-α for confidence limits (Range method)
Table A-7 Critical values of the signed-rank test
Table A-8 Critical values of the Wilcoxon-test
Table A-9 (continued) Critical values of the F-distribution : F0 975
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:45 PM Page 83
Query:
"Holderbank" - Cement Course 2000
Recommended Literature
The following list gives a selection of books on applied statistics for further reading. The list is restricted
to books which are easily comprehensible for users without profound knowledge in mathematics.
C. Chatfield "Statistics for Technology" Chapman and Hall, London
(1975), 350 pages
General survey on statistical methods with good
comments on the interpretation of statistical results.
Special chapters on regression, design of experiments and
quality control
Noether "Introduction to Statistics: A Fresh Approach" Houghton
Mifflin Company, Boston (1971) 230 pages
Modern statistical methods, especially for the analysis of
experiments. Good explanation of basic statistical ideas
and problems based on intuition.
M.G. Natrella "Experimental Statistics" US Department of Commerce,
NBS Handbook 91 (1963)
"Cookbook" with many test procedures and designs of
experiments. Often it is somewhat difficult to know what
procedure has to be chosen.
M.R. Spiegel "Theory and Problems of Statistics" Schaum's Outline
Series,(1961) New York 350 pages
Definitions and procedures of classical statistics
accompanied by many examples (875 solved problems)
© Holderbank Management & Consulting, 2000 6/23/2001 - 4:07:45 PM Page 84
Query: