Win Stat
Win Stat
WinSTAT for Excel
R. Fitch Software
Copyright © 2009 by R. Fitch Software.
All rights reserved worldwide.
WinSTAT is a registered trademark of Robert K. Fitch. All other trademarks are acknowledged.
Contents
Installation 1
Installing WinSTAT .................................................................................................................. 1
Basic Statistics 21
Overview ................................................................................................................................. 21
Basics/Descriptive ................................................................................................................... 21
Mean.......................................................................................................................... 22
Standard Error ........................................................................................................... 22
Standard Deviation .................................................................................................... 23
Variance..................................................................................................................... 23
Variation Coefficient ................................................................................................. 23
Relative Variation Coeffizient................................................................................... 23
Skewness ................................................................................................................... 23
Kurtosis ..................................................................................................................... 24
Minimum, Maximum, Range, Sum ........................................................................... 24
Median....................................................................................................................... 24
Compare 2 Groups 41
Overview................................................................................................................................. 41
Compare 2 Groups/Indepenent t-test....................................................................................... 41
Compare 2 Groups/Dependent t-test ....................................................................................... 42
Compare 2 Groups/U-test (Mann-Whitney)............................................................................ 44
Compare 2 Groups/Wilcoxon test ........................................................................................... 45
Compare 2 Groups/McNemar test........................................................................................... 47
Compare N Groups 49
Overview................................................................................................................................. 49
Compare N Groups/Analysis of Variance............................................................................... 50
One-factor ANOVA.................................................................................................. 51
Multiple comparisons................................................................................................ 51
Two-factor ANOVA ................................................................................................. 53
Compare N Groups/Repeated measures.................................................................................. 56
Split Plot Repeated Measures ................................................................................... 56
Compare N Groups/H-test (Kruskal-Wallis)........................................................................... 56
Compare N Groups/Friedman ................................................................................................. 58
Compare N Groups/Cochran Q-Test ....................................................................................... 59
Correlation 61
Overview................................................................................................................................. 61
The Correlation Dialog Box .................................................................................................... 61
Correlation/Pearson................................................................................................................. 62
Correlation/Spearman rank ..................................................................................................... 63
Correlation/Kendall's tau......................................................................................................... 64
Correlation/Partial ................................................................................................................... 65
Correlation/Cross-correlation.................................................................................................. 66
Discriminant Analysis 79
Overview ................................................................................................................................. 79
Discriminant Analysis.............................................................................................................. 79
Cluster Analysis 85
Overview ................................................................................................................................. 85
Agglomeration ......................................................................................................................... 86
Single linkage ............................................................................................................ 86
Complete linkage....................................................................................................... 86
Average linkage (UPGMA)....................................................................................... 86
Centroid method ........................................................................................................ 86
Ward's method........................................................................................................... 86
Agglomeration results ............................................................................................... 87
Cluster separation .................................................................................................................... 89
Factor Analysis 91
Overview ................................................................................................................................. 91
Factor Analysis ........................................................................................................................ 91
Factor Scores ............................................................................................................. 95
Survival Analysis 97
Overview ................................................................................................................................. 97
Kaplan-Meier........................................................................................................................... 97
Cox Regression...................................................................................................................... 103
Direct....................................................................................................................... 104
Stepwise .................................................................................................................. 105
Forward ................................................................................................................... 105
Backward................................................................................................................. 105
Output...................................................................................................................... 106
Index 129
Installing WinSTAT
In order to install WinSTAT, you will need to have “Administrator Rights” on the computer. If this is
not the case, then have WinSTAT installed by another user who can log in with “Administrator
Rights”. After a successful installation, it is not longer necessary to have “Administrator Rights” to
work with WinSTAT.
Place the distribution CD-ROM into the drive. Normally, the setup process will then begin
automatically. If not, start the Windows Explorer and double-click on [Link] in the CD-ROM
folder. Follow the instructions as they appear. When asked to enter the WinSTAT serial number, enter
the number printed on the sticker on the outside of your CD-ROM envelope. Once the installation is
complete you will notice that, whenever Excel is started, the WinSTAT toolbar has been added to the
other menu and toolbars.
WinSTAT Toolbar
In the latest version of Excel (2007), you will have to click on the “Add-Ins” tab to see the WinSTAT
toolbar.
A First Example
Using a simple example, we will demonstrate the major features of WinSTAT and show you how to
navigate among these features. To this end, we will use the example file [Link], which is
included under \samples in your WinSTAT installation folder. After opening the file, your screen
will look something like this:
An Excel workbook may contain many tables (worksheets) and each worksheet may have data entered
in any number of positions. Thus, WinSTAT needs some help in knowing exactly which data you are
interested in analyzing. When you see this picture, click on "Yes" to open the appropriate dialog box.
Using this dialog, you can tell WinSTAT exactly which data you are interested in analyzing. The most
common option is All used cells of the active worksheet. However, you may be dealing with a
worksheet containing several tables, or one in which there is some introductory text before the actual
data table begins. In such cases, do the following: First, select the cells in the worksheet that contain
the data you wish to analyze, including the variable names. Then select the WinSTAT command
DATA/SET DATABASE RANGE. This will open the above dialog box, and using the second option
Selected cells of the active worksheet, you will achieve the desired results.
As in the illustration, select "1 variable plus grouping variable", then select Height as the variable and
Age as the grouping variable. Note also the Help button at the lower right, which is available in all
WinSTAT dialog boxes and leads to information about the function selected.
Results as table
Results as diagram
There are several things to note right away. First of all, the table of results is in fact a normal Excel
worksheet, and as such, you may change colors, fonts, make the cell borders visible, and so on as you
would in any other worksheet. You may select any part of the results and copy/paste to Word, Power
Point, etc. The diagram, too, is a standard Excel diagram with its data taken directly out of the table.
Displayed as columns
Class Definition
Now have a closer look at the X-axis, containing age information. WinSTAT has automatically
grouped the data in ranges of 10 inches each. How did it know to do this? First of all, it knew that
some sort of grouping had to be performed, since we chose the MEANS command and not, for
example, the SCATTERPLOT command. Secondly, it followed some basic rules about grouping data
according to how many points are available and how they are distributed. But where has this grouping
information been stored and what if you want different group ranges?
To find the answer, look at the Excel screen again and note that yet another worksheet has been added
to the left of the others, named "WinSTAT Classes". Let's see what's in there:
WinSTAT Classes
The first entry, obviously, is the variable name. Below that you see the cells "90", "90 to 100", and in
the following line "100". Read this to mean: take all values from "90" up to, but not including "100",
and name this class "90 to 100". Similarly, the last two lines mean: take all values from "180" up to,
Then switch back to the "Means" worksheet. Table and diagram have automatically been updated to
reflect the change:
This experiment was designed to give you a feeling for the meaning of the "WinSTAT Classes"
worksheet. Usually, you will choose an easier method to redefine groups.
Clicking on "Suggestion" fills in the fields as shown. This is the grouping that WinSTAT chooses on
its own. Suppose we want a class width of 20? Just type in "20" in the appropriate field, then close the
box. Then go back to the "Means" worksheet to see how the results have been automatically updated:
Go to database range
Under Boy or Girl, select "girl". This causes Excel to hide all of the rows with data for boys. Now go
back to the "Means" worksheet. The results are recalculated with the chosen subset:
Automatic recalculation
Locking Results
Of course, dynamic updates may be too much of a good thing under certain circumstances. Suppose
we want to "freeze" these results while we perform other functions on the entire data or on other
subsets? This is accomplished with a single click on the "Lock results sheet" symbol:
With this button, you can determine on a sheet-by-sheet basis which results are locked and which are
dynamically linked to the database. To try it out, go ahead and lock the present "Means" worksheet.
Then return to the database and display all rows again using the autofilter arrow in Boy or Girl. A
look at the "Means" worksheet will confirm that it still displays the results for the 24 girls. Now unlock
the sheet with another click on the symbol, and it will recalculate to show all 48 children.
Here, we have chosen a second grouping variable, Boy or Girl, and most importantly, we have
selected the worksheet "Columns, red title" as template. In the template combo box appear all open
worksheets which were created using the same command. As expected, we obtain a new results
worksheet. The title cell is red instead of blue and the diagram uses columns instead of points.
Normal
18
16
14
12
Frequency
10
8
6
4
2
0
00
00
00
00
00
00
00
00
00
00
0
00
50
00 200
50
00
00 350
00
50
00 500
50
00
65
15 to 1
25 to 2
30 to 3
40 to 4
45 to 4
55 to 5
60 to 6
o
to
t
t
0
0
00
00
00
00
00
00
00
00
10
20
35
50
Incom e
The frequency data seems to fit fairly well within the bell-curve, but we have no information about
how significant the differences might be. Also, the steps in the histogram can make a comparison
difficult.
120.00
100.00
Percent (cumulative)
80.00
60.00
40.00
20.00
0.00
0 10000 20000 30000 40000 50000 60000 70000
Income Normal
On a cumulative frequency graph, a normal distribution will plot as a sigmoidal curve, as seen above.
In WinSTAT, you can also transform the Y-axis to a “probability scale”, in which case the normal
distribution will plot as a straight line.
Again, the fit in this example looks pretty good, and now we can try to back up this fact with a
statistical calculation.
120.00
100.00
Percent (cumulative)
80.00
60.00
40.00
20.00
0.00
0 1 2 3 4 5 6 7
Children Normal
Because of the obvious steps in the distribution, the distance to the normal curve is forced to be large.
In such cases, it is best to use a Chi-square test, which uses the absolute frequencies of the individual
variable values. WinSTAT automatically calculates the Chi-square test if the number of different
values (or classes) of the variable is not greater than 100.
120.00
100.00
Percent (cumulative)
80.00
60.00
40.00
20.00
0.00
0 50 100 150 200 250 300
Horsepow er Normal
It is obvious that the distance between the curves is much greater than in the previous example, and in
fact the Kolmogorov-Smirnov test yields a p-value of 0.00256.
Skewness
One way to influence a non-normal distribution is to try to make it symmetric about the median. A
measurement of symmetry is the Index of Skewness, which WinSTAT includes in its descriptive
statistics about each variable. A variable with a symmetric distribution about its mean has a skewness
of 0. A variable with a compact “lower tail” and an extended “upper tail” has a positive skewness.
Negative skewness indicates an extended “lower tail” and a compact “upper tail”.
Many real-life distributions exhibit positive skewness, as does the variable Horsepower, with a
skewness of 1.798. The reason is clear: There is a definite limit to the lowest practical horsepower, but
no limit to the highest. Also, as we go up the horsepower scale, we will find many cars bunched
together at the low end, and getting rarer as horsepower increases, probably because powerful cars are
also more expensive.
120.00
100.00
Percent (cumulative)
80.00
60.00
40.00
20.00
0.00
3 3.5 4 4.5 5 5.5 6
HP1 Normal
The curve is now quite symmetric in its appearance, and the Kolmogorov-Smirnov p-value has been
raised to 0.278. Thus, the new variable is not significantly different from the normal distribution and
could be used for any parametric tests requiring a normal distribution.
Overview
Many of the simpler statistics will be found in the BASICS submenu. They can be used to get
descriptive information about the variables, check distributions, and look for some simple kinds of
dependencies among variables.
Basics/Descriptive
Selecting this command opens the following dialog box:
Mean
1 n
x= ∑ xi
n i =1
Standard Error
This refers to the standard error of the mean, and measures the uncertainty involved in x . Obviously,
the larger the sample size, the more certain one can be that x is a good estimator of the true mean of
the population. The standard error is defined as
s
sx =
n
where s is the standard deviation as defined below.
s=
1
(n − 1)
(∑ x 2
i − n(∑ xi )
2
)
Variance
The variance is the square of the standard deviation and is denoted
s2
Variation Coefficient
The variation coefficient is the standard deviation in units of the arithmetic mean:
s
V =
x
Since an increase in mean often goes hand in hand with an increase in standard deviation, the variation
coefficient allows a more direct means of comparison.
Skewness
First, we define the following moments of the distribution:
1
2nd moment m 2 = (n − 1)s 2
n
1 3
3rd moment m3 =
n
∑ xi3 − x ∑ xi2 + 2 x 2
n
Skewness is then defined as
3
m3 (m2 )
−
2
For a discussion of the meaning of skewness, see the section on The Normal Distribution.
m4 (m2 ) − 3
−2
Median
The median is the value such that 50% of the sample cases lie below the value and 50% of the cases
above.
Percentiles
The percentiles are values such that a given percent of the sample cases lie below the value, and the
rest of the cases above.
1 to N Variables
You can create a frequencies table of one or more entire variables. The following example shows the
table and histogram of just one entire variable:
30
25
20
Frequency
15
10
0
0 1 2 3 4 5 6
Children
If a single variable is specified, the program automatically superimposes the best-fit normal curve for
the data.
Here is an example with two entire variables:
25
Frequency
20
15
10
0
camping pickup sedan sports car station
van w agon
If you specify more than one variable, as in the above example, all must have the exact same class
definition (as definined in the WinSTAT Classes table). This assures that a direct comparison is
meaningful.
New or Used
25
new
used
20
Frequency
15
10
0
camping pickup sedan sports car station
van w agon
Type
Entire Variables
If you want means information about entire variables, use the right-hand side of the dialog box,
choosing whichever variables are of interest. For example, choosing Price and Price(2):
± Standard deviation
14000
12000
10000
8000
Mean
6000
4000
2000
0
Price Price(2)
Means diagram
50000
45000
40000
35000
30000
Income
25000
20000
15000
10000
5000
0
camping pickup sedan sports car station
van w agon
Type
In this example, there seems to be a connection between income and type of car (people who drive
sports cars have a higher average income than people who drive sedans). The actual significance of
such a dependency can be tested with more advanced methods such as Analysis of Variance.
25 new
used
20
15
Mileage
10
0
camping pickup sedan sports car station
van w agon
Type
The test based on Turning Points is calculated as follows: Observe the time series, and count the
number of times there is a turning point (i.e. local minimum or maximum) in the data. If the data are
randomly organized, one would expect a turning point at 2/3 of all points. The difference between the
actual number and the expected number can be transformed to a Z-value (the abscissa on the normal
curve) to calculate the significance of the difference.
The turning points test is particularly suited to finding periodicity, since a time series exhibiting
periodicity will not have as many turning points as would be expected of a random series.
The Rank Test is performed as follows: Compare each value (call it X1) in the time series with all
values which occur later in time (call them X2). Assuming randomness, the probability of X2 > X1 is
1/2, and the expected number of comparisons in which X2 > X1 is
1
N ( N − 1)
4
Basics/Outliers
Outliers are values which lie so far away from the mean that one may suspect that the case in question
is not representative of the population measured. In the following example, we are looking for outliers
in the variables Income, Price, and Horsepower.
For each variable chosen, all cases which are suspected to be outliers are listed. The distance from the
mean (in multiples of sigma, the standard deviation of the variable) is printed. Then, a P-value is
printed, indicating the probability of finding at least one value at this distance from the mean in a
normally-distributed sample. For a given n-sigma, the probability increases with sample size.
All cases are listed for which n-sigma is greater than 4 or for which the P-value is less than 0.05. By
clicking within these fields with the mouse, you can change the parameters and redo the calculation.
Now note the pushbuttons below the results. These are used to manipulate the outliers within the
database. The first button will hide all of the outliers listed (which you can confirm by returning to the
database worksheet). Any calculations will then ignore these cases. The second button will hide only
those outliers which you specifically select. For example, click on the row containing "Income, 29",
then click on the second pushbutton to have only this case hidden in the database. To reinstate any
outliers which you have previously hidden, use the third pushbutton.
Kolmogorov-Smirnov test
This test is valid for continuous distributions. The cumulative frequency distribution should display no
obvious steps.
D is the maximum distance measured between the curve of the actual distribution and the best-fit
normal distribution. P is the probability that the given D-value could arise by random fluctuation in a
sample taken from a normally distributed population. Thus, a non-significant (high) p-value allows us
to assume that the variable is distributed normally. In the example, we may assume that Income is
distributed normally and that Horspower is not.
Note that the distribution curve of a variable may be examined graphically using the
GRAPHICS/CUMULATIVE command.
Chi-square test
This test is valid for noncontinuous (discrete) distributions. Because such a distribution displays
visible steps in the cumulative frequency diagramm, the Kolmogorov-Smirnov test measures an
exaggerated distance to the normal curve, so the Chi-square test yields better results. Again, P is the
probability that the given value could arise by random fluctuation in a sample taken from a normally
distributed population. Thus, a non-significant (high) p-value allows us to assume that the variable is
distributed normally.
A chi-square Worksheet
Look at the entries carefully. The first column, Outcome, could be ommitted, since its data is
irrelevant to the chi-square test, but it helps to keep track while entering the other values. The die was
thrown 65 times, and the observed frequencies are recorded in column Observed. Note that the
Expected values need not add up to 65. It is only necessary that their proportions to one another be as
expected. WinSTAT automatically calculates the true expected frequencies based on the total number
of Observed frequencies. In this case, all outcomes 1 to 6 are equally expected. Fill in the dialog box
as follows:
Basics/Crosstabs
This command produces a crosstabulation (bivariate frequency distribution) table for a pair of
variables. It is used to look for dependencies between nominal-level variables.
The output is a matrix of cells, each containing a set of numbers. The contents of each cell are
determined by the headers to the left of the output, which in turn depend on the statistics chosen in the
dialog box. Outside the matrix are row and column totals for the frequencies.
The cell chi-square is a measure of the significance of the difference between actual cell frequency and
expected frequency (that frequency which one would expect to observe if the two variables had no
effect on each other). The sum of all cell chi-squares yields the total chi-square, in this case 16.0417.
D.F. is the number of degrees of freedom. P is the probability that the observed chi-square could be a
result of random fluctuations in unrelated variables rather than of a true dependency. The number of
cells with an expected frequency less than 5 is also displayed, since the chi-square evaluation is
considered invalid if this number is greater than 20% (as in the example).
If the table is 2 by 2 (i.e. both variables are dichotomous), Fisher's Exact Test is computed and the p-
value displayed. This statistic is preferred over the chi-square statistic, especially for small sample
sizes.
The Contingency Coefficient and Cramer's V are measures of the dependency between the two
variables, and can range between 0 (no dependency) and 1 (total dependency).
Crosstabs Direct
Usually, as in the example above, the cell frequencies are calculated by the program using the case
data from the file. Sometimes, however, you may already have the frequencies in tabular form, rather
than the individual raw data. Consider the following worksheet:
The worksheet already contains the crosstabulation frequencies. For instance, teacher B was rated
good by 11 students and bad by 5 students. Now fill in the dialog box as follows:
Overview
This next category of statistics functions includes commands which allow you to look for significant
differences between two sets of data. The two data sets can either be the values of two different vari-
ables, or they can be the values of a single variable divided into two groups by the values of a second
variable. Two of the tests (the t-tests) assume that the data are distributed normally, and two of the
tests (U-test, Wilcoxon test) assume only an ordinal (non-parametric) distribution of the data.
For each group, the mean and standard deviation of the dependent variable is calculated. In this
example, we see that the mean price for new cars is higher than for used cars. F is a measure of the
difference in variance of the two groups. The first P indicates the two-tailed significance of this
difference in variance. If the variances are significantly different, the homogeneous (pooled) t-test is
considered invalid. In this case, one relies on the heterogeneous (separate) t-test, which takes into
account the different group variances.
P indicates the two-tailed significance of the t-value. It is the probability that the observed difference
in means could be the result of random fluctuations in the dependent variable rather than of a true
dependency. In the example, the difference in price is significant.
As an example, we have selected the price of a family's first car compared to the price of the family's
second car. It is important to note that we are making comparisons within one family, thus satisfying
the “same object” requirement mentioned above.
The results indicate that the mean difference in price between the first and second car is significant
(low p-value).
Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Horsepower as the dependent variable and New or Used as the independent variable.
Sometimes, however, you may have already sorted the dependent data itself into two columns. To go
along with the example, you would have the horsepowers of new cars in one variable and the
horsepowers of used cars in the second variable. In this case, use the right half of the dialog box to
select the variables. Here are the results:
The number of negative and positive differences is displayed, as well as the corresponding sum of
ranks. This value is then translated to a Z-value, which is the abscissa of the equivalent point on the
Gaussian normal curve. The Z-value is used to compute the significances. P indicates the significance
of the difference in mean ranks, being the probability that the observed difference could be the result
of random fluctuations in the variables rather than of a true dependency.
In the example, we see that a family's second car has, on average, less horsepower than its first car,
and that this result is highly significant.
McNemar worksheet
The McNemar test calculates a chi-square statistic for the differences in response. The P-value reflects
the significance of the difference.
The Kappa index is useful if comparing the judgement of two people pertaining to the same objects.
Suppose, for example, two doctors examine the same set of patients. Each doctor assigns a "yes" or
"no" to each patient, depending on whether he diagnoses a certain symptom or not. The Kappa value
can range from 0 to 1, with high values indicating good agreement between the two variables.
Overview
In the previous section, you were introduced to functions which compare exactly two data sets for
significant differences. This next category includes commands to compare more than two data sets
among each other. The general name for such comparisons is Analysis of Variance (ANOVA), and in-
cludes the pure analysis of variance for variables with a normal distribution as well as non-parametric
methods.
Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Price as the dependent variable and Type as the independent variable. Sometimes,
however, you may have already sorted the dependent data itself into several columns. To go along
with the example, you would have the prices of sports cars in one variable, the prices of station
wagons in a second variable, and so on. In this case, use the right half of the dialog box to select the
variables.
D.F. is the number of degrees of freedom. The Mean Sum of Squares is the quotient of the sum of
squares and D.F. The Between groups F-value is the quotient of the mean S.S. (between) and the mean
S.S. (within) and is a measure of the differences in means among the various groups. P indicates the
significance of this difference, it being the probability that the observed difference could be a result of
random fluctuations in the dependent variable rather than of a true dependency. Under the Bartlett test,
the chi-square value is a measure of the differences in variances among the various groups. P again
indicates the significance. In the example, there is a significant indication that the price of a car
depends on its type.
It should be noted, however, that the significant results indicated in the above example are invalidated
by the Bartlett test, since it indicates a significant difference in the variances. In this case it would be
advisable instead to use the Kruskal-Wallis analysis, described later, since it is not dependent on
homogeneous variances.
Multiple comparisons
If the one-factor ANOVA determines that differences in the means exist, it is often of interest to ask
which of the group means differ significantly from which other group means. Or put another way,
which subsets of groups can be built whose members have means which do not differ significantly
from each other. The answer to this question is provided by a set of methods called “multiple compari-
sons,” also known as “a posteriori” or “post hoc” tests.
All of the methods listed follow the same general pattern: A range statistic is calculated for each
possible subset size, and if the greatest difference in means within a subset is less than the calculated
range, then the subset is considered to contain members which do not differ significantly from each
other. Various theories regarding the calculation of the ranges are reflected in the different methods.
p
pa =
1
g ( g − 1)
2
where g is the total number of groups. The critical range for each pair of groups is then calculated
using the adjusted level in a t-distribution.
Scheffé: The range for any pair of groups is calculated using an F-distribution. This method is the
most conservative in that it produces the largest ranges of all methods, thus requiring a greater
separation of means before the difference is considered significant.
Tukey: The critical range for each pair of groups is calculated using the studentized range distribution,
which takes into account the total number of groups present.
B-Tukey: Also known as modified Tukey, the critical range for each comparison is simply the average
between the standard Tukey value and the value calculated by S-N-K.
S-N-K: Named after the authors Student, Newman, and Keuls, this statistic also calculates the
studentized range, but uses the number of groups in each subset being investigated rather than the total
number of groups.
Duncan: As in S-N-K, this method calculates the studentized range using the number of groups in the
subset. In addition, it adjusts the significance level to account for the fact that many comparisons are
being made. The formula is:
p n = 1 − (1 − p ) n
where n is the number of groups in the subset being investigated.
REGW: Named after the authors Ryan, Einot, Gabriel, and Welsch, this method is similar to Duncan
in that the significance level is adjusted before the studentized range is calculated. The formula is:
n
p n = 1 − (1 − p ) g
where n is the number of groups in the subset being investigated and g is the total number of groups.
Note that you can click within the worksheet itself to change the method and p-value.
The question as to which types of cars are responsible for the significant difference in price is
answered by the S-N-K multiple comparison test. Note that the table is sorted according to price, that
is, from station wagon (least expensive) to sports car (most expensive). To explain the table, let's
look at one example. In the first row (station wagon), we see the number 2781.80 under sedan. Read
this to mean: within the subset containing all car types from station wagon to sedan (and including
pickup, since its price lies between the two), we may call the differences significant only if the means
vary by more than 2781.80. Since the mean of station wagon is 4562.56 and the mean of sedan is
6591.88, and since the difference between the two is less than 2781.80, there are no significant
differences among these three types. Accordingly, there is a no in the corresponding lower-left
position (in the sedan row under station wagon).
By analyzing the table, WinSTAT can calculate subsets of the grouping variable which are
significantly different from each other. In our example, one subset consists of the sports cars and the
other subset consists of all other car types.
Two-factor ANOVA
The goal of a two-factor (or two-way) ANOVA is to look at the effect of one variable after controlling
for the effect of another variable. For instance, we may want to compare the average mileage of the
cars from different manufacturers after controlling for the number of cylinders.
Two-factor ANOVA
D.F. is the number of degrees of freedom. The Mean S.S. is the quotient of the sum of squares and
D.F. F is the quotient of Mean S.S. and the Error Mean S.S. The Between Groups F-value is a
measure of the differences in means among the various variables and is the statistic of interest. P
indicates the significance of this difference, it being the probability that the observed difference could
be a result of random fluctuations in the variables rather than of a true dependency.
The demonstration file has at most two variables that measure “the same thing,” in this example Price
and Price(2). With only two variables, the comparison can be carried out more efficiently with the
T-test (dependent) command. If you compare the results of the two methods, you will see that the p-
values are identical.
Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Horsepower as the dependent variable and Type as the independent variable.
Sometimes, however, you may have already sorted the dependent data itself into several columns. To
go along with the example, you would have the horsepowers of sports cars in one variable, the
horsepowers of station wagons in a second variable, and so on. In this case, use the right half of the
dialog box to select the variables.
H is a measure of the difference in mean rank among the groups. D.F. is the number of degrees of
freedom, and P indicates the significance, it being the probability that the observed H-value could be
the result of random fluctuations in the dependent variable rather than of a true dependency.
Compare N Groups/Friedman
This test is similar to a repeated measures analysis of variance, in that several variables are examined,
each of which must belong to “the same object” and measure “the same thing” under different
circumstances. The variables need only be measured at the ordinal level, and this is the main
difference between repeated measures and the Friedman test. An example of such variables could be
the ratings of three similar products, each case representing the ratings of one expert.
For each case, the values of the given variables are inspected. The variable with the lowest value has
its value replaced by 1, the second-lowest variable becomes 2, etc. After all cases have been handled,
the mean of each variable is computed, this being its mean rank. The Friedman test determines to what
extent the differences in mean rank are significant.
Use the dialog box to select the desired variables. Choosing Horsepower and Horsepower(2) the
results are:
Overview
This submenu gives you access to all correlation functions. Correlation is a measure of the degree of
dependency between two variables. It says nothing about the cause of the dependency, which could be
some third factor entirely. The following sections describe each command in detail.
The results of a correlation are displayed as a table of rows and columns. Each cell contains
information about the correlation between the row variable and column variable. Use the dialog box to
select which variables should appear as rows and which as columns. If you wish, use the Copy
pushbutton to make the column variables the same as the row variables already selected.
The first entry in each cell is the correlation coefficient itself. The second entry is the number of cases
that entered into the calculation (those cases which have non-missing values for both variables in
question). The last entry is the significance of the correlation. The p-value decreases with increasing
correlation and with increasing case count, indicating increased significance.
Note that if the exact same variables for rows and columns are chosen, the two “consistency” measures
Cronbach’s Alpha and Scott’s Homogeneity Quotient are included in the analysis.
Refer to the PEARSON command for a description of the table contents. You may wish to compare this
example with the Spearman rank correlation and note the similarity of the results.
Correlation/Cross-correlation
This command is best described by an example. Suppose we take temperature readings every x
minutes in two neighboring rooms. One of the rooms is heated and cooled directly, the adjacent room
indirectly through its proximity to the first room. Obviously, the two temperature variables will be
highly correlated, but we would expect the temperature in the second room to lag behind the
temperature in the primary room by some number of time units. That is, the correlation would be even
higher if we could shift the variables in time with respect to one another.
This is exactly what a cross-correlation accomplishes. It shifts the variables one unit (case) at a time
forward and backward and calculates the resulting Pearson correlation coefficient, up to a prede-
termined number of lags.
Look at the sample file [Link]. The source of the data is
G.E.P. Box and G.M. Jenkins, Time Series Analysis: Forecasting and Control, Holden Day, 1976.
The variable Indicator represents a financial leading indicator which is supposed to predict the
volume of future sales. The variables Sales is the actual monthly sales figure. We expect a correlation
between the two variables, and would hope to find a maximum correlation at some time lag,
confirming that the indicator can be used as a predictor.
This results in the following graphical display (note that the CROSS-CORRELATION command always
produces a diagram as well as a table):
0.8
Correlation
0.6
0.4
0.2
0
-30 -20 -10 0 10 20 30
Lag
Rule: a high correlation at a negative lag means that the 1st variable can be used to predict the 2nd
variable. A high correlation at a positive lag means that the 2nd variable can be used to predict the 1st
variable.
It is clear that the indicator correlates with actual sales many months in advance. However, an exact
interpretation is impossible, because we may assume that both the indicator and the sales figures
Overview
The various commands on this submenu try to describe one variable (the dependent variable) as a
function of one or more other variables. The following sections describe each command in detail.
Regression/Simple
Use this command to look for simple equations relating a given dependent variable to a single
independent variable. Various classes of equations may be taken into consideration. Fill in the dialog
box appropriately:
In this example, we will model Mileage as a function of Horsepower, looking for the best fit among
all of the equation classes listed.
R is the correlation coefficient between the actual Y values and the values calculated by the given
equation. High values of R (maximum = 1.0) indicate a good fit. The square of R appears in the next
column and, in statistical terms, is the percentage of variance in the dependent variable which can be
explained by the given equation.
35
30
25
20
Mileage
15
10
0
0 50 100 150 200 250 300
Horsepow er
Residual Analysis
Having created the new variable Regression residuals, we can perform some tests which may lead to
more information about the quality of our best-fit function.
Normality: the residuals should be normally distributed. This can be tested with the
GRAPHICS/CUMULATIVE FREQUENCY command or with the BASICS/TEST OF RANDOMNESS command.
In our example, the Kolmogorov-Smirnov p-value of 0.97925 indicates normality.
Randomness: the residuals should exhibit randomness. Even if the input variables show trend or
periodicity, this fact should not influence the residuals. If it does, then we have missed something
important in modelling the data.
12
10
8
Regression Residuals
6
4
2
0
-2 0 50 100 150 200 250 300
-4
-6
-8
-10
Horsepow er
In the above plot, it appears that the variance of the residuals decreases as X (Horsepower) increases.
However, since there are many fewer data points for the higher X-values, this is probably not
significant, especially since the normality of the residuals has already been demonstrated.
See REGRESSION/SIMPLE for a discussion of the Write residues in and Recalculate cursive rows and
overwrite Y-values checkboxes. By checking Constant = 0, the regression equation is forced to pass
through the origin.
R and R square have the same meaning as in simple regression. The corrected values (in parentheses)
include a “punishment factor” for each increased number of terms. Obviously, a better fit (higher R)
can always be found by increasing the order of the polynomial, but at the expense of simplicity and
ability to interpret the curve. Maximizing the corrected value of R is a method of deciding when to
stop.
Std. error is the standard error of the regression curve (compared to the observed values), measured in
units of the dependent variable.
For each term, the calculated coefficient is displayed as well as a confidence interval for the
coefficient. One may be confident that the true value lies within the indicated bounds. Note that the
percent value (here 95%) may be changed by clicking within the field. T is the quotient of the
coefficient and its standard error. P indicates the significance of the given term in the equation.
35
30
25
20
Mileage
15
10
0
0 50 100 150 200 250 300
Horsepow er
See REGRESSION/POLYNOMIAL for a discussion of the Write residues in, Recalculate cursive rows and
overwrite Y-values, Y-weightings and Constant = 0 checkboxes.
The methods available are:
Direct
All of the independent variables chosen are included in the regression equation. This is standard
multiple regression.
Stepwise
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. After each addition, all
variables now in the equation are reinvestigated. The variable with the least significance (highest p-
value) is then removed from the equation, assuming its p-value is higher than p-out from the dialog
Forward
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. The process continues until
no more variables can be added or until the maximum number of variables (as indicated by Maximum
number of variables in the dialog box) has been reached.
Backward
Starting with all independent variables in the regression equation, the program investigates each one.
The variable with the least significance (highest p-value) is then removed from the equation, assuming
its p-value is higher than p-out from the dialog box. The process continues until no more variables can
be removed.
Maximum R-Square
This method demands significantly more calculation time than the others, but it is guaranteed to find
the “best” (in the R-Square sense) one-variable solution, then the best two-variable solution, etc. At
each step it may remove and add variables as it finds better solutions.
Output
As an example, we choose Horsepower as the dependent variable, and Income and Price as the
independent variables. The Stepwise method is selected, with a maximum of two variables. The
results are:
Overview
In Discriminant Analysis, it is assumed that the data at hand are grouped according to a given nominal-
level variable. Then, any number of further variables are examined to see to what extent they can be
used to “predict” the value of the grouping variable. The similarity to multiple regression is apparent,
the difference being that the variable to be “predicted” in multiple regression is interval-level.
Once the analysis has been performed, the derived equations can be used to classify cases which have
not yet been assigned to groups. A classic example is the bond market. First, discriminant analysis is
performed using bonds of known quality (AAA, AA, ..., C), relating these ratings to various financial
indicators of the respective companies. Following this, the financial indicators of previously unrated
companies can be fed into the equations to get bond ratings for those companies.
Discriminant Analysis
Fill in the dialog box:
The Prior Probabilities are only of meaning when using the derived functions for classification
purposes. If all equal, the group closest to the value calculated by the equations will be chosen for the
classification. If proportional to present frequencies, a bias is installed which makes classification
Overview
Cluster analysis refers to methods which attempt to group cases in such a manner that the members of
each group are, in some sense, “close” to one another. Several variables may be chosen for the analy-
sis, and the differences in these variables between two cases determines the “distance” between the
two cases.
The most common method of calculating the distance is, for the two cases in question, to sum the
squared differences of each of the variables in question. This is known as the squared Euclidian
distance. Before the distances are calculated, the values of all variables in question are standardized
(that is, transformed to variables with mean = 0 and variance = 1), so that the variables with greater
values and variances do not dominate the equation. Note that WinSTAT automatically standardizes the
variables before calculating the Euclidian distances.
Once the distances have been obtained, there are several methods for collecting the cases into groups.
Because the groups grow, as cases (and groups) are combined, and because, once in a group, a case is
never removed, these methods are known as agglomeration methods. The various agglomeration
methods are described in detail below.
Select the variables which are to determine the distances, as well as one of the agglomeration methods.
Also, if one of the file's variables contains labels which identify each case, you may indicate this using
the checkbox. This results in the case labels being used for all output instead of case numbers.
Single linkage
This method (also known as nearest neighbor) says that the distance between two groups is simply the
distance between the two closest members of the groups.
Complete linkage
This method (also known as farthest neighbor) says that the distance between two groups is equal to
the distance between the two farthest members of the groups.
Centroid method
This method says that the distance between two groups is equal to the distance between the two group
centroids. A group centroid is its center of gravity measured over all variables.
Ward's method
This method (also known as incremental sums of squares) says that the distance between two groups is
proportional to the change in the within group sum of squares (see Analysis of Variance) which results
when the two groups are combined.
At each step, a combination of two clusters occurs, and so the total number of clusters is reduced by
one. The two columns joining Cluster 1 and with Cluster 2 indicate the two clusters which are being
combined at this step. The number displayed is the lowest case number of all cases in the cluster. If
case names were indicated in the dialog box, then the labels are displayed instead of numbers. The
number in parentheses indicates the number of cases presently in each of the clusters. Finally, the last
column displays the distance between the two clusters as calculated by the method chosen.
0
10
20
30
40
Distance
50
60
70
80
90
100
Dendogram
A main purpose of the dendogram is as a tool in deciding how many clusters one wants to retain.
Looking at the complete dendogram, we see that the reduction of three groups to two groups (at
distance 38) was only possible with a considerable increase in distance compared to previous
combinations. That is, two clusters were combined which were much farther apart than the clusters
which had been combined up to that point. It would appear logical, therefore, to retain three clusters.
Note that cluster separation is not available for the neighbor joining method
Once the number of clusters has been decided upon, click on the Cluster separation button to assign
cases to respective groups.
You must indicate the number of clusters and the name of a variable which will contain the cluster
number of each case. This variable name must previously have been inserted into the worksheet
database. In this example, a new variable Ward's clusters has been added. Each case in cluster 1 is
given the value 1, each case in cluster 2 the value 2, etc. This variable may be used to group cases in
further analyses, such as discriminant analysis.
Use the GRAPHICS/SCATTERPLOT command, using the new variable as Grouping variable, to get a
plot of the data as grouped.
Ward's Clusters
70000
1
60000 2
3
50000
40000
Income
30000
20000
10000
0
0 50 100 150 200 250 300
Horsepow er
Overview
Factor Analysis is used to look for basic, independent dimensions underlying the existing variables.
One tries to find the smallest set of such dimensions which nevertheless explain the variables to a
sufficient degree. There are many options to consider when preparing a factor analysis, and often a
satisfactory solution will only be found by repeating the analysis a number of times with different
parameters. You should have a general idea ahead of time which variables are to be explained and the
approximate number of factors which may lie behind these variables.
Factor Analysis
Fill in the dialog box:
In all, 76 percent of the variance of all selected variables can be explained by these two factors.
WinSTAT also produces a graphical display of the factor loadings. Here is the plot of the factor
loadings from the example:
0.5
Size of Household
Children
Factor 2
0 Income
Horsepow er
Price
-0.5 Mileage
-1
-1 -0.5 0 0.5 1
Factor 1
By viewing the plot, it is easy to ascertain which variables “belong” to only one factor (those that lie
close to an axis), and which variables are determined by information from both factors.
Overview
WinSTAT offers two survival analysis features: Kaplan-Meier and Cox Regression.
Kaplan-Meier calculates and displays survival tables according to the data at hand and can also check
if a single grouping variable has a significant influence on the survival probability curves.
Cox Regression allows the user to supply several independent variables which may or may not have a
significant influence on the survival probability curves. A regression equation for the influence of
these variables in then calculated. As with multiple linear regression, there are several options which
cause WinSTAT to search for those variables which are significant. Another feature allows the user to
plot a custom survival curve for a single row of data representing one patient or other item.
Kaplan-Meier
The sample file [Link] contains typical data for a survival analysis:
The entries are automatically sorted according to survival time. The column At Risk indicates how
many cases were still in the survey at the given time, before the event or censoring occurred. The next
two columns indicate how many events occurred and how many cases were censored at the given time.
The column Survival probability indicates, for the data at hand, the probability of surviving up to the
given time. The final column gives the standard error for the calculated probabilities.
1.2
0.8
Probability
Censored
0.6
Probability
0.4
0.2
0
0 100 200 300 400
Survival Tim e
Survival probabilities
For each group, the expected number of events is calculated under the assumption that there is no
difference between the two groups. Comparing these values with the actual number of events observed
yields a chi-square statistic. P is the significance of this statistic.
1.2
0.8
Probability
Censored
0.6 Drug A
Placebo
0.4
0.2
0
0 100 200 300 400
Survival Tim e
Here we see the same data as in the Kaplan-Meier survival analysis, but we also see two new columns
to the right, Treatment numeric and Age. The first of these, Treatment numeric, is nothing other
than a copy of Treatment, but using the numeric coding 0 and 1 instead of the text values Drug A and
Placebo. This is necessary because the regression calculation depends on the mathematical use of
numeric variables. All independent variables to be used in a Cox regression must be numeric.
The following variables are necessary for any Cox regression analysis:
1. The first variable contains the survival time data. This is usually an integer value representing
a number of days, months, or years. The variable is often defined as the difference between
two dates. The first date is the time when the case (e.g. patient) entered the survey, and the
second date is when the case left the survey.
2. The second necessary variable is called the event variable. This variable indicates whether the
survey ended for the given case because the pre-defined event (e.g. death) occurred, or
whether the case dropped out of the survey for another reason. If the event variable is not
missing and non-zero, it indicates that the event occurred. If missing or zero, it indicates that
the case was censored, that is, dropped out of the survey.
3. Finally, one or more independent variables which may influence the survival time must be
selected.
As in REGRESSION/MULTIPLE there are several options which determine how WinSTAT decides which
of the independent variables are significant enough to be used in the regression equation:
Direct
All of the independent variables chosen are included in the regression equation. This is standard Cox
regression.
Forward
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. The process continues until
no more variables can be added or until the maximum number of variables (as indicated by Maximum
number of variables in the dialog box) has been reached.
Backward
Starting with all independent variables in the regression equation, the program investigates each one.
The variable with the least significance (highest p-value) is then removed from the equation, assuming
its p-value is higher than p-out from the dialog box. The process continues until no more variables can
be removed.
For each variable in the equation, the calculated coefficient is displayed as well as a confidence
interval for the coefficient. One may be confident that the true value lies within the indicated bounds.
Note that the percent value (here 95%) may be changed by clicking within the field. P indicates the
significance of the given term in the equation. Hazard is e to the power of the coefficient and
indicates to what extent the variable increases the risk of the object under examination.
In the example, one can say that receiving a placebo instead of the drug increases the hazard, as does
an increased age. The effect of age is less relevant. It should be noted, however, that neither of these
variables is particularly significant.
A new individual
Thus, we are answering the question: What survival curve can be expected for a 50 year old patient
treated with the placebo?
The column Survival Probability, Cursive case is now filled with the specific probabilities for the
new individual. The probability curve is also displayed, along with the original curve for comparison
purposes:
Process Capability
The process capability ratio compares the allowed variation in a product's specification with the
variation actually found in a series of measurements on the actual product. It is used to estimate the
fraction of defective products which the process will produce.
The command opens the following dialog box (using the file [Link]):
Overview
This menu gives you access to WinSTAT's graphics capabilities. Some graphics, which relate directly
to a given statistical method, are available directly through that statistic (for example, the plot of a
regression curve). Those plots are described along with the corresponding statistics function and not
repeated here.
Graphics/Histogram
The GRAPHICS/HISTOGRAM command is exactly the same as the BASICS/FREQUENCIES command, so
we simply refer you to that section here.
Graphics/Means plot
The GRAPHICS/MEANS command is exactly the same as the BASICS/MEANS command, so we simply
refer you to that section here.
70000
60000
50000
40000
30000
20000
10000
0
Income Price
The short line within the box represents the median of the given variable. The bottom and top edges
represent the 25th and 75th percentiles. In other words, 50% of the data fall within the box, and 25%
each above and below. The “whiskers” extend to the 5th and 95th percentiles. Finally, the minimum
and maximum values in the sample are indicated with a '+' sign.
Here is a sample scatterplot, plotting Income and Price against each other and grouping according to
New or Used:
New or Used
30000
new
25000 used
20000
Price
15000
10000
5000
0
0 20000 40000 60000 80000
Incom e
120.00
100.00
Percent (cumulative)
80.00
60.00
40.00
20.00
0.00
0 50 100 150 200 250 300
Horsepow er Normal
10.00
9.00
8.00
7.00
6.00
Probit
5.00
4.00
3.00
2.00
1.00
0.00
0 50 100 150 200 250 300
Horsepow er Normal
Probability scale
On a probability scale, the Y axis is transformed in such a manner that a normal distribution plots as a
straight line.
Variable
Select the variable of interest. Note that the examples below refer to several different sample files and
to specific variables within them.
Sample Size
The idea of a sample size is common to all chart types. It is assumed that, at regular intervals, a
constant number of items (the sample) is taken from the production process and measured or tested.
For instance (see X-bar, below), we could measure the inside diameter of 5 piston rings taken from the
production line every 10 minutes. If this is repeated 20 times, the variable will contain 100
measurements (cases) in all, and the sample size is 5. The average measurement of the 5 items (thus
X-bar) will be plotted at each interval, so there will be 20 points on the chart.
Warning limits
This optional value determines the limits beyond which we consider the process to be in need of
examination. A value which is widely used is 2-sigma.
Specification
In addition to control and warning limits, the actual specification limits for the article being
manufactured may be included in the diagram. This is only applicable when an X-bar chart (see
below) is being drawn.
74.02
74.015
74.01
Mean
74
Control limits
73.995
73.99
73.985
0 10 20
Sam ple, N = 5
X-bar chart
There are 25 data points, each one showing the average diameter of the 5 rings in the corresponding
sample.
0.06
0.05
0.04
Inside diameter
Range
0.03 Mean
Control limits
0.02
0.01
0
0 10 20
Sam ple, N = 5
R chart
0.025
0.02
Standard deviation
0.015
Inside diameter
Mean
0.01 Control limits
0.005
0
0 10 20
Sam ple, N = 5
S chart
0.6
0.5
Fraction nonconforming
0.4
Bad cans
0.3 Mean
Control limits
0.2
0.1
0
0 10 20 30
Sam ple, N = 50
p chart
In the above example, it is apparent that something has caused the process to be out of control at
samples 15 and 23.
30
25
Number nonconforming
20
Bad cans
15 Mean
Control limits
10
0
0 10 20 30
Sam ple, N = 50
np chart
45
40
35
30
Count of defects
25 Errors
Mean
20
Control limits
15
10
0
0 10 20
Sam ple, N = 100
c chart
0.45
0.4
0.35
0.3
Defects per unit
0.25 Errors
Mean
0.2
Control limits
0.15
0.1
0.05
0
0 10 20
Sam ple, N = 100
u chart
In the dialog box, specify all of the variables to be included in the chart:
70
80
60
50
60
40
40
30
20
20
10
0 0
Size outside Bad parts Insufficient Misaligned Paint out of
specs glue w eld limits
Pareto Chart
Index
Cox Regression 97, 103
Cramer's V 38
Cronbach’s alpha 62
Cross-correlation 66
Crosstabs 37
cumulative frequency plot 17
Cumulative frequency plot 116
D
A Database 4
a posteriori 51 dendogram 88
Agglomeration 86 Descriptive statistics 21
Analysis of Variance 50 Discriminant Analysis 79
arithmetic mean 22 dummy variables 54
average linkage 86 Duncan 52
Dynamic Link between Data and
Results 11
B
balanced experiment 54 E
Bartlett test 50, 51
Basic statistics 21 eigenvalue 92
Bonferroni 52 eigenvalues 82
Box & Whisker 114 Euclidian distance 85
Box-Cox transformation 20 event 98, 104
breakdown of means 30 expected frequencies 36
expected frequency 38
extraction of factors 92
C
c chart 125 F
canonical correlation 82
canonical discriminant functions 82 Factor Analysis 91
censor 98, 104 factor loadings, graph of 94
centroid method 86 Factor Scores 95
chi-square 38, 51 Fisher's exact test 38
Chi-square 36 Frequencies 25
chi-square (Friedman test) 59 frequency 38
Chi-square test 18, 35 Friedman 58
Class Definition 8 F-value 51
Cluster Analysis 85
Cluster separation 89 G
Cochran Q-Test 59
communality 92 Graphics 113
Compare 2 Groups 41
Compare N Groups 49 H
complete linkage 86
confidence interval 30, 74, 78, 106 histogram 16
confusion table 83 Histogram 25, 113
consistency measures 62 H-test 56
L R
Locking Results 12 R chart 121
LSD 52 randomness 33
range 24
M rank sum 46
Regression 69
Mahalanobis distance 83 REGW 52
maximum 24 Repeated measures 56
maximum R-Square regression 77 residual analysis 71
McNemar test 47 residuals 70, 73, 76
mean 22 rotation of factors 92
mean rank 44, 45, 57, 58
Means 29
median 24 S
minimum 24 S chart 122
multiple comparisons 51 sample size 118
Multiple regression 76 Scatterplot 115
Scheffé 52
N Scott’s homogeneity quotient 62
sigma 119
normal curve 26, 116 Simple regression 69
normal distribution 15 single linkage 86
np chart 124 skewness 19, 23
S-N-K 52
O Spearman rank correlation 63
Split Plot Repeated Measures 56
observed frequencies 36 standard deviation 23
one-factor ANOVA 51 standard error of mean 22
Outliers 34 stepwise regression 76
sum 24
Survival Analysis 97
P
p chart 123 T
Pareto Chart 127
Partial correlation 65 Templates 13
Pearson correlation 62 Test of normal distribution 35
percentiles 24 Test of randomness 33
Polynomial regression 73 Tests of Normal Distribution 18
post hoc 51 transforming non-normal data 19
U
u chart 126
U-test (Mann-Whitney) 44
V
variance 23
Variation coefficient 23
varimax 92
W
Ward's method 86
Warning limits 119
Weighted Regression 74
Wilcoxon test 45
Wilks' lambda 82
X
X-bar chart 120