0% found this document useful (0 votes)
4 views137 pages

Win Stat

Uploaded by

workchatgem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views137 pages

Win Stat

Uploaded by

workchatgem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

User's Manual


WinSTAT for Excel

R. Fitch Software
Copyright © 2009 by R. Fitch Software.
All rights reserved worldwide.
WinSTAT is a registered trademark of Robert K. Fitch. All other trademarks are acknowledged.
Contents

Installation 1
Installing WinSTAT .................................................................................................................. 1

Guided Tour of WinSTAT 3


A First Example......................................................................................................................... 3
Selecting a WinSTAT Database ................................................................................................ 4
The Database Range dialog ....................................................................................................... 4
Class Definition ......................................................................................................................... 8
The Class Definition dialog ..................................................................................................... 10
Dynamic Link between Data and Results ................................................................................ 11
Locking Results ....................................................................................................................... 12
Using Templates ...................................................................................................................... 13

The Normal Distribution 15


Importance of the Normal Distribution.................................................................................... 15
Histogram ................................................................................................................................ 16
Cumulative Frequency Plot ..................................................................................................... 17
Tests of Normal Distribution ...................................................................................................18
Transforming non-normal Data ............................................................................................... 19
Skewness ................................................................................................................... 19
The Box-Cox Transformation ................................................................................... 20

Basic Statistics 21
Overview ................................................................................................................................. 21
Basics/Descriptive ................................................................................................................... 21
Mean.......................................................................................................................... 22
Standard Error ........................................................................................................... 22
Standard Deviation .................................................................................................... 23
Variance..................................................................................................................... 23
Variation Coefficient ................................................................................................. 23
Relative Variation Coeffizient................................................................................... 23
Skewness ................................................................................................................... 23
Kurtosis ..................................................................................................................... 24
Minimum, Maximum, Range, Sum ........................................................................... 24
Median....................................................................................................................... 24

WinSTAT User's Manual Installation • i


Percentiles................................................................................................................. 24
Basics/Frequencies (also Graphics/Histogram)....................................................................... 25
1 to N Variables ........................................................................................................ 25
1 Variable plus Grouping Variable ........................................................................... 28
Basics/Means (also Graphics/Means) ..................................................................................... 29
Entire Variables ........................................................................................................ 29
One Variable Grouped .............................................................................................. 30
Basics/Test of randomness ...................................................................................................... 33
Basics/Outliers ........................................................................................................................ 34
Basics/Test of normal distribution .......................................................................................... 35
Kolmogorov-Smirnov test ........................................................................................ 35
Chi-square test .......................................................................................................... 35
Basics/Chi-square.................................................................................................................... 36
Basics/Crosstabs...................................................................................................................... 37
Crosstabs Normal...................................................................................................... 38
Crosstabs Direct ........................................................................................................ 39

Compare 2 Groups 41
Overview................................................................................................................................. 41
Compare 2 Groups/Indepenent t-test....................................................................................... 41
Compare 2 Groups/Dependent t-test ....................................................................................... 42
Compare 2 Groups/U-test (Mann-Whitney)............................................................................ 44
Compare 2 Groups/Wilcoxon test ........................................................................................... 45
Compare 2 Groups/McNemar test........................................................................................... 47

Compare N Groups 49
Overview................................................................................................................................. 49
Compare N Groups/Analysis of Variance............................................................................... 50
One-factor ANOVA.................................................................................................. 51
Multiple comparisons................................................................................................ 51
Two-factor ANOVA ................................................................................................. 53
Compare N Groups/Repeated measures.................................................................................. 56
Split Plot Repeated Measures ................................................................................... 56
Compare N Groups/H-test (Kruskal-Wallis)........................................................................... 56
Compare N Groups/Friedman ................................................................................................. 58
Compare N Groups/Cochran Q-Test ....................................................................................... 59

Correlation 61
Overview................................................................................................................................. 61
The Correlation Dialog Box .................................................................................................... 61
Correlation/Pearson................................................................................................................. 62
Correlation/Spearman rank ..................................................................................................... 63
Correlation/Kendall's tau......................................................................................................... 64
Correlation/Partial ................................................................................................................... 65
Correlation/Cross-correlation.................................................................................................. 66

ii • Installation WinSTAT User's Manual


Regression 69
Overview ................................................................................................................................. 69
Regression/Simple ................................................................................................................... 69
Residual Analysis ...................................................................................................... 71
Prediction of unknown Y-values ............................................................................... 72
Regression/Polynomial ............................................................................................................ 73
Weighted Regression................................................................................................. 74
Regression/Multiple................................................................................................................. 76
Direct......................................................................................................................... 76
Stepwise .................................................................................................................... 76
Forward ..................................................................................................................... 77
Backward................................................................................................................... 77
Maximum R-Square .................................................................................................. 77
Output........................................................................................................................ 77

Discriminant Analysis 79
Overview ................................................................................................................................. 79
Discriminant Analysis.............................................................................................................. 79

Cluster Analysis 85
Overview ................................................................................................................................. 85
Agglomeration ......................................................................................................................... 86
Single linkage ............................................................................................................ 86
Complete linkage....................................................................................................... 86
Average linkage (UPGMA)....................................................................................... 86
Centroid method ........................................................................................................ 86
Ward's method........................................................................................................... 86
Agglomeration results ............................................................................................... 87
Cluster separation .................................................................................................................... 89

Factor Analysis 91
Overview ................................................................................................................................. 91
Factor Analysis ........................................................................................................................ 91
Factor Scores ............................................................................................................. 95

Survival Analysis 97
Overview ................................................................................................................................. 97
Kaplan-Meier........................................................................................................................... 97
Cox Regression...................................................................................................................... 103
Direct....................................................................................................................... 104
Stepwise .................................................................................................................. 105
Forward ................................................................................................................... 105
Backward................................................................................................................. 105
Output...................................................................................................................... 106

WinSTAT User's Manual Installation • iii


Examining a specific new case ............................................................................... 107

Process Capability 111


Process Capability ................................................................................................................. 111

The GRAPHICS Menu 113


Overview............................................................................................................................... 113
Graphics/Histogram .............................................................................................................. 113
Graphics/Means plot ............................................................................................................. 113
Graphics/Box & Whisker ...................................................................................................... 114
Graphics/Scatterplot .............................................................................................................. 115
Graphics/Cumulative frequency............................................................................................ 116
Graphics/Quality control ....................................................................................................... 118
Variable................................................................................................................... 118
Sample Size............................................................................................................. 118
Control limits .......................................................................................................... 119
Warning limits ........................................................................................................ 119
Specification ........................................................................................................... 119
Using a subset of measurements to determine the control limits ............................ 119
X-bar Chart ............................................................................................................. 120
R Chart.................................................................................................................... 121
S Chart .................................................................................................................... 122
p Chart .................................................................................................................... 123
np Chart .................................................................................................................. 124
c Chart..................................................................................................................... 125
u Chart .................................................................................................................... 126
Graphics/Pareto Chart ........................................................................................................... 127

Index 129

iv • Installation WinSTAT User's Manual


Installation

Installing WinSTAT
In order to install WinSTAT, you will need to have “Administrator Rights” on the computer. If this is
not the case, then have WinSTAT installed by another user who can log in with “Administrator
Rights”. After a successful installation, it is not longer necessary to have “Administrator Rights” to
work with WinSTAT.
Place the distribution CD-ROM into the drive. Normally, the setup process will then begin
automatically. If not, start the Windows Explorer and double-click on [Link] in the CD-ROM
folder. Follow the instructions as they appear. When asked to enter the WinSTAT serial number, enter
the number printed on the sticker on the outside of your CD-ROM envelope. Once the installation is
complete you will notice that, whenever Excel is started, the WinSTAT toolbar has been added to the
other menu and toolbars.

WinSTAT Toolbar

In the latest version of Excel (2007), you will have to click on the “Add-Ins” tab to see the WinSTAT
toolbar.

WinSTAT User's Manual Installation • 1


Guided Tour of WinSTAT

A First Example
Using a simple example, we will demonstrate the major features of WinSTAT and show you how to
navigate among these features. To this end, we will use the example file [Link], which is
included under \samples in your WinSTAT installation folder. After opening the file, your screen
will look something like this:

Sample file opened

WinSTAT User's Manual Guided Tour of WinSTAT • 3


Let's see how children's heights change with age. Using the WinSTAT menu bar, select the function
STATISTICS/BASICS/MEANS. Excel may need some time to load the appropriate WinSTAT function,
but this delay occurs only once in any one Excel session.
Now we are informed by the Excel assistant that WinSTAT is not yet ready to perform, since
information about the desired database is missing.

Selecting a WinSTAT Database

WinSTAT needs a database first

An Excel workbook may contain many tables (worksheets) and each worksheet may have data entered
in any number of positions. Thus, WinSTAT needs some help in knowing exactly which data you are
interested in analyzing. When you see this picture, click on "Yes" to open the appropriate dialog box.

The Database Range dialog

Selecting the database

Using this dialog, you can tell WinSTAT exactly which data you are interested in analyzing. The most
common option is All used cells of the active worksheet. However, you may be dealing with a
worksheet containing several tables, or one in which there is some introductory text before the actual
data table begins. In such cases, do the following: First, select the cells in the worksheet that contain
the data you wish to analyze, including the variable names. Then select the WinSTAT command
DATA/SET DATABASE RANGE. This will open the above dialog box, and using the second option
Selected cells of the active worksheet, you will achieve the desired results.

4 • Guided Tour of WinSTAT WinSTAT User's Manual


Important: Please keep in mind that a workbook can only have one WinSTAT database, so be sure to
select all columns that you will be analyzing. A common mistake is to select only those columns to be
used for the first analysis, then to change the database definition for the next analysis, and so on.
Because of WinSTAT’s dynamic link between data and results, changing the database definition
would break the link. The proper way is to include all columns in the single database, then select the
desired columns for each analysis within the appropriate dialog boxes.
Back to our example: Since the present worksheet does not contain any data outside of the table we are
interested in, we can simply accept the default suggestion All used cells of the active worksheet. The
database selection will remain in effect until we change it, so that normally this will be a one-time
operation for any one workbook. However, if after setting the database, you add data to the end of any
column, remember to reselect the database with the command DATA/SET DATABASE RANGE before
continuing. Otherwise WinSTAT will ignore the new data.
At this point, we can try STATISTICS/BASICS/MEANS once again, and get the following dialog box:

Means dialog box

As in the illustration, select "1 variable plus grouping variable", then select Height as the variable and
Age as the grouping variable. Note also the Help button at the lower right, which is available in all
WinSTAT dialog boxes and leads to information about the function selected.

WinSTAT User's Manual Guided Tour of WinSTAT • 5


The results appear in a new Excel worksheet, which is automatically inserted into the workbook.
These results are displayed in tabular form and as a diagramm.

Results as table

Results as diagram

There are several things to note right away. First of all, the table of results is in fact a normal Excel
worksheet, and as such, you may change colors, fonts, make the cell borders visible, and so on as you
would in any other worksheet. You may select any part of the results and copy/paste to Word, Power
Point, etc. The diagram, too, is a standard Excel diagram with its data taken directly out of the table.

6 • Guided Tour of WinSTAT WinSTAT User's Manual


WinSTAT has added some active controls to the results. For instance, you can use the mouse to click
inside the "95%" confidence header and change the value. Change the header e.g. to 99% and the table
values are automatically updated to reflect the change. To the right of the diagram are some option
buttons which allow you to select the meaning of the error bars.
Let's change the diagram to display columns instead of points: Right-click within the diagram area and
choose "Chart type" from the context menu to display the chart type options:

Chart type options

WinSTAT User's Manual Guided Tour of WinSTAT • 7


Select "Column" the get these results:

Displayed as columns

Class Definition
Now have a closer look at the X-axis, containing age information. WinSTAT has automatically
grouped the data in ranges of 10 inches each. How did it know to do this? First of all, it knew that
some sort of grouping had to be performed, since we chose the MEANS command and not, for
example, the SCATTERPLOT command. Secondly, it followed some basic rules about grouping data
according to how many points are available and how they are distributed. But where has this grouping
information been stored and what if you want different group ranges?
To find the answer, look at the Excel screen again and note that yet another worksheet has been added
to the left of the others, named "WinSTAT Classes". Let's see what's in there:

WinSTAT Classes

The first entry, obviously, is the variable name. Below that you see the cells "90", "90 to 100", and in
the following line "100". Read this to mean: take all values from "90" up to, but not including "100",
and name this class "90 to 100". Similarly, the last two lines mean: take all values from "180" up to,

8 • Guided Tour of WinSTAT WinSTAT User's Manual


but not including "190", and name this class "180 to 190". Try the following experiment: overwrite
"90 to 100" with "youngest kids" and overwrite "180 to 190" with "oldest kids":

A custom class definition

Then switch back to the "Means" worksheet. Table and diagram have automatically been updated to
reflect the change:

The custom classes in the diagram

This experiment was designed to give you a feeling for the meaning of the "WinSTAT Classes"
worksheet. Usually, you will choose an easier method to redefine groups.

WinSTAT User's Manual Guided Tour of WinSTAT • 9


The Class Definition dialog
Select STATISTICS/BASICS/CLASS DEFINITION from the WinSTAT menu bar to call up the class
definition dialog box for the variable Age:

Class definition dialog box

Clicking on "Suggestion" fills in the fields as shown. This is the grouping that WinSTAT chooses on
its own. Suppose we want a class width of 20? Just type in "20" in the appropriate field, then close the
box. Then go back to the "Means" worksheet to see how the results have been automatically updated:

Another class definition change

10 • Guided Tour of WinSTAT WinSTAT User's Manual


We have seen how WinSTAT dynamically responds to changes in the class definition. How about if
we change the data itself? Use the "Go to database range" symbol:

Go to database range

in the WinSTAT menu bar to return to the database worksheet.

Dynamic Link between Data and Results


Using the Excel menu (not the WinSTAT menu!), select DATA/FILTER/AUTOFILTER, causing Excel to
insert its autofilter arrows into the header row:

Excel's autofilter arrows

Under Boy or Girl, select "girl". This causes Excel to hide all of the rows with data for boys. Now go
back to the "Means" worksheet. The results are recalculated with the chosen subset:

Automatic recalculation

WinSTAT User's Manual Guided Tour of WinSTAT • 11


Note that the sample now contains only the 24 girls and not the original 48 children. The diagram has
also been updated accordingly. The dynamic update illustrated here will also take place if you simply
change values within the database.

Locking Results
Of course, dynamic updates may be too much of a good thing under certain circumstances. Suppose
we want to "freeze" these results while we perform other functions on the entire data or on other
subsets? This is accomplished with a single click on the "Lock results sheet" symbol:

Lock results sheet

With this button, you can determine on a sheet-by-sheet basis which results are locked and which are
dynamically linked to the database. To try it out, go ahead and lock the present "Means" worksheet.
Then return to the database and display all rows again using the autofilter arrow in Boy or Girl. A
look at the "Means" worksheet will confirm that it still displays the results for the 24 girls. Now unlock
the sheet with another click on the symbol, and it will recalculate to show all 48 children.

12 • Guided Tour of WinSTAT WinSTAT User's Manual


Using Templates
WinSTAT has a standard display form for each of its functions. We saw in the above example that the
MEANS command creates a diagram with each mean value displayed as a point. We wanted to see the
data as columns and edited the diagram accordingly. Suppose we always prefer the column option? Or
suppose we always want the table in a different color or font? Try the following experiment: go to the
"Means" results worksheet and change the color of the first cell from blue to red. Then change the
name of the worksheet from "Means" to "Columns, red title":

Results sheet renamed

WinSTAT User's Manual Guided Tour of WinSTAT • 13


Let's see if we can use this custom worksheet as a template for a new calculation. Select
STATISTICS/BASICS/MEANS once again and fill in the dialog box as follows:

Specifying the template

Here, we have chosen a second grouping variable, Boy or Girl, and most importantly, we have
selected the worksheet "Columns, red title" as template. In the template combo box appear all open
worksheets which were created using the same command. As expected, we obtain a new results
worksheet. The title cell is red instead of blue and the diagram uses columns instead of points.

14 • Guided Tour of WinSTAT WinSTAT User's Manual


The Normal Distribution

Importance of the Normal Distribution


We include this section here, because understanding the nature of the normal distribution is essential to
the correct use of many of the statistical methods that will follow. The discussion itself uses some of
the statistics and graphics which are described later in the manual. Please refer ahead to those sections
if necessary.
Many statistical methods, the so-called parametric methods, assume that the distribution of data within
a given variable conforms to certain prerequisites. In particular, it is usually necessary to show that the
distribution does not differ significantly from a normal distribution.
Every data set has a mean. If the data are normally distributed, then the data points are clustered
around the mean according to the famous bell-shaped (Gaussian) probability curve. The curve can be
wider or narrower, depending on the standard deviation of the data, but other than that its shaped is
fixed. Because of the fixed nature of the curve, the probability of a data point falling beyond a given
distance from the mean can be calculated exactly, and it is upon this calculation that many statistical
methods are based.

WinSTAT User's Manual The Normal Distribution • 15


Histogram
How can we tell if the data at hand can be considered to follow the normal distribution? One
possibility is to break continuous data into classes, then create a histogram (frequency plot) and
compare the histogram by eye with a best-fit normal curve. Look at the following histogram of the
variable Income in the file [Link]:

Normal

18
16
14
12
Frequency

10
8
6
4
2
0
00

00

00

00

00

00

00

00

00

00

0
00
50

00 200

50

00

00 350

00

50

00 500

50

00

65
15 to 1

25 to 2

30 to 3

40 to 4

45 to 4

55 to 5

60 to 6
o

to
t

t
0

0
00

00

00

00

00

00

00

00
10

20

35

50

Incom e

Histogram with best-fit normal curve

The frequency data seems to fit fairly well within the bell-curve, but we have no information about
how significant the differences might be. Also, the steps in the histogram can make a comparison
difficult.

16 • The Normal Distribution WinSTAT User's Manual


Cumulative Frequency Plot
With continuous data (as in this example), we can get a more exact picture of the distribution by
plotting the cumulative frequency. This method avoids the breakdown into classes which were
necessary for a histogram. The X-axis of the graph runs from the minimum to maximum value of the
variable being analyzed. Each Y-value indicates the percent of cases with values less than or equal to
the corresponding X-value. Thus, the Y-values will always run from 0 to 100 over the width of the
graph.

120.00

100.00
Percent (cumulative)

80.00

60.00

40.00

20.00

0.00
0 10000 20000 30000 40000 50000 60000 70000

Income Normal

Cumulative frequency with best-fit normal curve

On a cumulative frequency graph, a normal distribution will plot as a sigmoidal curve, as seen above.
In WinSTAT, you can also transform the Y-axis to a “probability scale”, in which case the normal
distribution will plot as a straight line.
Again, the fit in this example looks pretty good, and now we can try to back up this fact with a
statistical calculation.

WinSTAT User's Manual The Normal Distribution • 17


Tests of Normal Distribution
Two methods are offered which can be used to support the assumption that given data are distributed
normally. The Kolmogorov-Smirnov test calculates the maximum distance between the cumulative
frequency curve of data and the best-fit normal curve, and then determines the significance of this
distance.
In the example (Income), the Kolmogorov-Smirnov test yields a p-value of 0.692 (high). A low p-
value would indicate significant results, so we conclude here that the difference between the two
curves is not significant. Thus, we may claim that the data are distributed normally.
The Kolmogorov-Smirnov test is best used on continuously distributed variables. The cumulative
frequency distribution of a discrete variable could appear as in this example:

120.00

100.00
Percent (cumulative)

80.00

60.00

40.00

20.00

0.00
0 1 2 3 4 5 6 7

Children Normal

Discrete distribution with steps

Because of the obvious steps in the distribution, the distance to the normal curve is forced to be large.
In such cases, it is best to use a Chi-square test, which uses the absolute frequencies of the individual
variable values. WinSTAT automatically calculates the Chi-square test if the number of different
values (or classes) of the variable is not greater than 100.

18 • The Normal Distribution WinSTAT User's Manual


Transforming non-normal Data
Suppose the data in question do not pass the Kolmogorov-Smirnov test. That is, the test yields a
significantly small p-value. Is there any way to transform the data, such that the transformed data are
distributed normally? It turns out that this is often possible.
Consider the variable Horsepower in the file [Link], whose cumulative frequency plot looks
like this:

120.00

100.00
Percent (cumulative)

80.00

60.00

40.00

20.00

0.00
0 50 100 150 200 250 300

Horsepow er Normal

Cumulative frequency of Horsepower

It is obvious that the distance between the curves is much greater than in the previous example, and in
fact the Kolmogorov-Smirnov test yields a p-value of 0.00256.

Skewness
One way to influence a non-normal distribution is to try to make it symmetric about the median. A
measurement of symmetry is the Index of Skewness, which WinSTAT includes in its descriptive
statistics about each variable. A variable with a symmetric distribution about its mean has a skewness
of 0. A variable with a compact “lower tail” and an extended “upper tail” has a positive skewness.
Negative skewness indicates an extended “lower tail” and a compact “upper tail”.
Many real-life distributions exhibit positive skewness, as does the variable Horsepower, with a
skewness of 1.798. The reason is clear: There is a definite limit to the lowest practical horsepower, but
no limit to the highest. Also, as we go up the horsepower scale, we will find many cars bunched
together at the low end, and getting rarer as horsepower increases, probably because powerful cars are
also more expensive.

WinSTAT User's Manual The Normal Distribution • 19


The Box-Cox Transformation
The Box-Cox transformation may be used to reduce skewness. It is defined by
λ
−1
Y= X for λ ≠ 0;
λ
Y = ln X for λ = 0
Values of λ greater than 1 can be used to eliminate negative skewness, while values of λ in the range 0
< λ < 1 can be used to eliminate positive skewness. In practice, values of 0.5 (square root) or 0 (log)
are often sufficient.
WinSTAT supports the Box-Cox transformation as a built-in function. To continue with our example,
simply insert a new column to the right of Horsepower and name it HP1 (in cell G1). Move the cursor
to cell G2 and type the function definition as follows:
=WS_BOXCOX(F2,0)
The definition can be propogated downward by dragging on the lower right corner of the cell. HP1 is
now the Box-Cox transformation of variable Horsepower, with a λ of 0. The variable HP1 has a
skewness of 0.0864. This is quite good (i.e. close to 0), and trying other values of λ does not yield an
improvement. Now look at the cumulative frequency plot:

120.00

100.00
Percent (cumulative)

80.00

60.00

40.00

20.00

0.00
3 3.5 4 4.5 5 5.5 6

HP1 Normal

Cumulative frequency of HP1

The curve is now quite symmetric in its appearance, and the Kolmogorov-Smirnov p-value has been
raised to 0.278. Thus, the new variable is not significantly different from the normal distribution and
could be used for any parametric tests requiring a normal distribution.

20 • The Normal Distribution WinSTAT User's Manual


Basic Statistics

Overview
Many of the simpler statistics will be found in the BASICS submenu. They can be used to get
descriptive information about the variables, check distributions, and look for some simple kinds of
dependencies among variables.

Basics/Descriptive
Selecting this command opens the following dialog box:

Descriptive statistics dialog box

WinSTAT User's Manual Basic Statistics • 21


Select any number of variables from the list box. The results appear as in the following example:

Mean

The arithmetic mean is defined as

1 n
x= ∑ xi
n i =1

Standard Error
This refers to the standard error of the mean, and measures the uncertainty involved in x . Obviously,
the larger the sample size, the more certain one can be that x is a good estimator of the true mean of
the population. The standard error is defined as
s
sx =
n
where s is the standard deviation as defined below.

22 • Basic Statistics WinSTAT User's Manual


Standard Deviation
The standard deviation of the sample is defined as

s=
1
(n − 1)
(∑ x 2
i − n(∑ xi )
2
)
Variance
The variance is the square of the standard deviation and is denoted

s2

Variation Coefficient
The variation coefficient is the standard deviation in units of the arithmetic mean:
s
V =
x
Since an increase in mean often goes hand in hand with an increase in standard deviation, the variation
coefficient allows a more direct means of comparison.

Relative Variation Coeffizient


The relative variation coefficient is given in percent of the maximum possible value:
s
Vr [%] = x 100
N

Skewness
First, we define the following moments of the distribution:
1
2nd moment m 2 = (n − 1)s 2
n
1 3
3rd moment m3 =
n
∑ xi3 − x ∑ xi2 + 2 x 2
n
Skewness is then defined as
3
m3 (m2 )

2

For a discussion of the meaning of skewness, see the section on The Normal Distribution.

WinSTAT User's Manual Basic Statistics • 23


Kurtosis
In addition to the above moments, we define the 4th moment
1 4 6
m4 =
n
∑ xi4 − x ∑ xi3 + x 2 ∑ xi2 − 3x 4
n n
Kurtosis is defined as

m4 (m2 ) − 3
−2

Minimum, Maximum, Range, Sum


These items are self-explanatory.

Median
The median is the value such that 50% of the sample cases lie below the value and 50% of the cases
above.

Percentiles
The percentiles are values such that a given percent of the sample cases lie below the value, and the
rest of the cases above.

24 • Basic Statistics WinSTAT User's Manual


Basics/Frequencies (also Graphics/Histogram)
The term "Frequencies" refers to the number of cases per class of a given variable.
Fill in the dialog box:

Frequencies/histogram dialog box

1 to N Variables
You can create a frequencies table of one or more entire variables. The following example shows the
table and histogram of just one entire variable:

WinSTAT User's Manual Basic Statistics • 25


Normal

30

25

20
Frequency

15

10

0
0 1 2 3 4 5 6
Children

Histogram with normal curve

If a single variable is specified, the program automatically superimposes the best-fit normal curve for
the data.
Here is an example with two entire variables:

26 • Basic Statistics WinSTAT User's Manual


40
Type
35 Type(2)
30

25
Frequency

20

15

10

0
camping pickup sedan sports car station
van w agon

Histogram with two variables

If you specify more than one variable, as in the above example, all must have the exact same class
definition (as definined in the WinSTAT Classes table). This assures that a direct comparison is
meaningful.

WinSTAT User's Manual Basic Statistics • 27


1 Variable plus Grouping Variable
Choose this option if you want to create a frequencies table for a single variable, and subdivide each
class according to the values of a second, grouping variable. Here is an example, with Type as the
main variable and New or Used as the grouping variable:

New or Used

25
new
used
20
Frequency

15

10

0
camping pickup sedan sports car station
van w agon
Type

Histogram with subgroups

28 • Basic Statistics WinSTAT User's Manual


Basics/Means (also Graphics/Means)
This command yields detailed information about the means either for several separate variables or for
one variable “broken down” (grouped) by the values or classes of one or two other variables.

Means dialog box

Entire Variables
If you want means information about entire variables, use the right-hand side of the dialog box,
choosing whichever variables are of interest. For example, choosing Price and Price(2):

WinSTAT User's Manual Basic Statistics • 29


The items N, Mean, [Link], and [Link]. are the same as those delivered by the
BASICS/DESCRIPTIVES command. The Confidence interval is the range within which one may be
confident that the true population mean lies. Note that the percent value (95%) can be set to any other
value by clicking within the edit field. The Means command also creates a diagram of the calculated
means with error bars:

± Standard deviation

14000

12000

10000

8000
Mean

6000

4000

2000

0
Price Price(2)

Means diagram

One Variable Grouped


If you want a “breakdown of means”, use the left-hand side of the dialog box. Enter the variable to be
grouped, plus one or two grouping variables. As a simple example, here is a breakdown of means for
Income (the dependent variable) grouped by the Type of automobile (the independent variable):

30 • Basic Statistics WinSTAT User's Manual


± Standard deviation

50000
45000
40000
35000
30000
Income

25000
20000
15000
10000
5000
0
camping pickup sedan sports car station
van w agon
Type

Means with grouping variable

In this example, there seems to be a connection between income and type of car (people who drive
sports cars have a higher average income than people who drive sedans). The actual significance of
such a dependency can be tested with more advanced methods such as Analysis of Variance.

WinSTAT User's Manual Basic Statistics • 31


Here is an example with two grouping variables:

± Standard deviation New or Used

25 new
used
20

15
Mileage

10

0
camping pickup sedan sports car station
van w agon
Type

Means with two grouping varibles

32 • Basic Statistics WinSTAT User's Manual


Basics/Test of randomness
It is sometimes of interest to investigate whether the values of a variable, in the order in which they
are recorded, display some sort of significant behavior. For example, do the values rise or fall with
increasing case number (trend), or do they exhibit periodicity?
In particular, this question will be posed during a so-called Residual Analysis. In many statistical
methods (e.g. Regression), we try to model the behavior of one variable based on the values of others.
If the model is good, then the difference between a value predicted by the model and the true value
should not depend on the position of the value in the overall series. Thus, we are interested in
demonstrating the randomness of the residuals.
A dialog box appears, allowing you to specify which variables should be tested (note that in
Regression, for example, one can write the residuals to a new variable). Two tests are performed on
each variable chosen.

The test based on Turning Points is calculated as follows: Observe the time series, and count the
number of times there is a turning point (i.e. local minimum or maximum) in the data. If the data are
randomly organized, one would expect a turning point at 2/3 of all points. The difference between the
actual number and the expected number can be transformed to a Z-value (the abscissa on the normal
curve) to calculate the significance of the difference.
The turning points test is particularly suited to finding periodicity, since a time series exhibiting
periodicity will not have as many turning points as would be expected of a random series.
The Rank Test is performed as follows: Compare each value (call it X1) in the time series with all
values which occur later in time (call them X2). Assuming randomness, the probability of X2 > X1 is
1/2, and the expected number of comparisons in which X2 > X1 is
1
N ( N − 1)
4

WinSTAT User's Manual Basic Statistics • 33


The difference between the actual number and the expected number can be transformed to a Z-value
(the abscissa on the normal curve) to calculate the significance of the difference.
The rank test is particularly useful for detecting a linear trend in the data.
In the above example, none of the differences is significant, so we may claim that the variables are
organized randomly.

Basics/Outliers
Outliers are values which lie so far away from the mean that one may suspect that the case in question
is not representative of the population measured. In the following example, we are looking for outliers
in the variables Income, Price, and Horsepower.

For each variable chosen, all cases which are suspected to be outliers are listed. The distance from the
mean (in multiples of sigma, the standard deviation of the variable) is printed. Then, a P-value is
printed, indicating the probability of finding at least one value at this distance from the mean in a
normally-distributed sample. For a given n-sigma, the probability increases with sample size.
All cases are listed for which n-sigma is greater than 4 or for which the P-value is less than 0.05. By
clicking within these fields with the mouse, you can change the parameters and redo the calculation.
Now note the pushbuttons below the results. These are used to manipulate the outliers within the
database. The first button will hide all of the outliers listed (which you can confirm by returning to the
database worksheet). Any calculations will then ignore these cases. The second button will hide only
those outliers which you specifically select. For example, click on the row containing "Income, 29",
then click on the second pushbutton to have only this case hidden in the database. To reinstate any
outliers which you have previously hidden, use the third pushbutton.

34 • Basic Statistics WinSTAT User's Manual


Basics/Test of normal distribution
Use this command to compare the distributions of any number of variables with a normal distribution.
Since many statistical methods require that the variables be distributed normally, you will want to
support that claim by using this test on the variables involved. For a more detailed discussion, see the
section on The Normal Distribution. A dialog box appears, allowing you to specify which variables are
to be tested.

Kolmogorov-Smirnov test
This test is valid for continuous distributions. The cumulative frequency distribution should display no
obvious steps.

D is the maximum distance measured between the curve of the actual distribution and the best-fit
normal distribution. P is the probability that the given D-value could arise by random fluctuation in a
sample taken from a normally distributed population. Thus, a non-significant (high) p-value allows us
to assume that the variable is distributed normally. In the example, we may assume that Income is
distributed normally and that Horspower is not.
Note that the distribution curve of a variable may be examined graphically using the
GRAPHICS/CUMULATIVE command.

Chi-square test
This test is valid for noncontinuous (discrete) distributions. Because such a distribution displays
visible steps in the cumulative frequency diagramm, the Kolmogorov-Smirnov test measures an
exaggerated distance to the normal curve, so the Chi-square test yields better results. Again, P is the
probability that the given value could arise by random fluctuation in a sample taken from a normally
distributed population. Thus, a non-significant (high) p-value allows us to assume that the variable is
distributed normally.

WinSTAT User's Manual Basic Statistics • 35


Basics/Chi-square
This command compares measured frequencies with expected frequencies for a single variable, and is
best explained with an example. Suppose we throw a die a number of times and record the number of
occurrences for each possible outcome, 1 to 6. The expected frequencies would be equal numbers for
each possibility. Fill in a worksheet with the observed frequencies and the expected frequencies:

A chi-square Worksheet

Look at the entries carefully. The first column, Outcome, could be ommitted, since its data is
irrelevant to the chi-square test, but it helps to keep track while entering the other values. The die was
thrown 65 times, and the observed frequencies are recorded in column Observed. Note that the
Expected values need not add up to 65. It is only necessary that their proportions to one another be as
expected. WinSTAT automatically calculates the true expected frequencies based on the total number
of Observed frequencies. In this case, all outcomes 1 to 6 are equally expected. Fill in the dialog box
as follows:

Chi-square dialog box

The results are displayed as:

36 • Basic Statistics WinSTAT User's Manual


The resulting p-value (not significant) indicates that there is no reason to suspect that the die being
used is unfair.

Basics/Crosstabs
This command produces a crosstabulation (bivariate frequency distribution) table for a pair of
variables. It is used to look for dependencies between nominal-level variables.

Crosstabulation dialog box

WinSTAT User's Manual Basic Statistics • 37


Crosstabs Normal
In the following example, we crosstabulate the variables Type and New or Used, displaying all
available statistics:

The output is a matrix of cells, each containing a set of numbers. The contents of each cell are
determined by the headers to the left of the output, which in turn depend on the statistics chosen in the
dialog box. Outside the matrix are row and column totals for the frequencies.
The cell chi-square is a measure of the significance of the difference between actual cell frequency and
expected frequency (that frequency which one would expect to observe if the two variables had no
effect on each other). The sum of all cell chi-squares yields the total chi-square, in this case 16.0417.
D.F. is the number of degrees of freedom. P is the probability that the observed chi-square could be a
result of random fluctuations in unrelated variables rather than of a true dependency. The number of
cells with an expected frequency less than 5 is also displayed, since the chi-square evaluation is
considered invalid if this number is greater than 20% (as in the example).
If the table is 2 by 2 (i.e. both variables are dichotomous), Fisher's Exact Test is computed and the p-
value displayed. This statistic is preferred over the chi-square statistic, especially for small sample
sizes.
The Contingency Coefficient and Cramer's V are measures of the dependency between the two
variables, and can range between 0 (no dependency) and 1 (total dependency).

38 • Basic Statistics WinSTAT User's Manual


In the example, the highest cell chi-squares come from the columns station wagon and sports car.
Investigating the observed and expected frequencies, we could conclude that station wagons tend to be
bought used, while sports cars tend to be bought new.

Crosstabs Direct
Usually, as in the example above, the cell frequencies are calculated by the program using the case
data from the file. Sometimes, however, you may already have the frequencies in tabular form, rather
than the individual raw data. Consider the following worksheet:

Crosstabs direct worksheet

The worksheet already contains the crosstabulation frequencies. For instance, teacher B was rated
good by 11 students and bad by 5 students. Now fill in the dialog box as follows:

Crosstabulation dialog box (Direct)

WinSTAT User's Manual Basic Statistics • 39


As you can see, you need only select the Direct button. It is not necessary to specify any variables,
since the program will analyse all available data in the worksheet. The results are:

40 • Basic Statistics WinSTAT User's Manual


Compare 2 Groups

Overview
This next category of statistics functions includes commands which allow you to look for significant
differences between two sets of data. The two data sets can either be the values of two different vari-
ables, or they can be the values of a single variable divided into two groups by the values of a second
variable. Two of the tests (the t-tests) assume that the data are distributed normally, and two of the
tests (U-test, Wilcoxon test) assume only an ordinal (non-parametric) distribution of the data.

Compare 2 Groups/Indepenent t-test


This command performs a t-test for two independent (or uncorrelated) samples. The values for a given
dependent variable (measured at the interval level and assumed to have a normal distribution) are
placed into two groups by a second, independent variable. The arithmetic mean will generally be
different for the two groups. The t-test determines to what extent the difference is significant.
Note: If the dependent variable is not distributed normally, use the U-test instead of the t-test.
The following dialog box appears:

T-test (independent) dialog box

WinSTAT User's Manual Compare 2 Groups • 41


Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Price as the dependent variable and New or Used as the independent variable.
Sometimes, however, you may have already sorted the dependent data itself into two columns. To go
along with the example, you would have the prices of new cars in one variable and the prices of used
cars in the second variable. In this case, use the right half of the dialog box to select the variables. Here
are the results:

For each group, the mean and standard deviation of the dependent variable is calculated. In this
example, we see that the mean price for new cars is higher than for used cars. F is a measure of the
difference in variance of the two groups. The first P indicates the two-tailed significance of this
difference in variance. If the variances are significantly different, the homogeneous (pooled) t-test is
considered invalid. In this case, one relies on the heterogeneous (separate) t-test, which takes into
account the different group variances.
P indicates the two-tailed significance of the t-value. It is the probability that the observed difference
in means could be the result of random fluctuations in the dependent variable rather than of a true
dependency. In the example, the difference in price is significant.

Compare 2 Groups/Dependent t-test


This command performs a t-test for dependent (or correlated) samples. A pair of variables is analyzed.
The variables must be measured at the interval level. Each must belong to “the same object” and must
measure “the same thing” under different circumstances. For example, you might want to investigate a
subject's weight before a diet and after the diet. For each case, the difference of the two variables is
computed. These differences are assumed to have a normal distribution. Then, the mean and standard

42 • Compare 2 Groups WinSTAT User's Manual


deviation of these differences over the entire sample are computed. The t-test determines to what
extent the differences are significant.
Use the dialog box to select the two variables:

T-test (dependent) dialog box

As an example, we have selected the price of a family's first car compared to the price of the family's
second car. It is important to note that we are making comparisons within one family, thus satisfying
the “same object” requirement mentioned above.

The results indicate that the mean difference in price between the first and second car is significant
(low p-value).

WinSTAT User's Manual Compare 2 Groups • 43


Compare 2 Groups/U-test (Mann-Whitney)
This test is similar to the t-test for independent samples. The dependent variable need only be
measured at the ordinal level, and this is the main difference between the t-test and the U-test.
All cases of the dependent variable are first sorted into ascending order. Each value is then replaced by
its position after sorting (its rank). Thus, the lowest value becomes 1, the next lowest becomes 2, etc.
These rank values are then separated into two groups according to the independent variable. Each
group has its own mean, this being the mean rank of the original values, and the U-test determines to
what extent the difference in mean rank is significant.
The following dialog box appears:

U-Test dialog box

Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Horsepower as the dependent variable and New or Used as the independent variable.
Sometimes, however, you may have already sorted the dependent data itself into two columns. To go
along with the example, you would have the horsepowers of new cars in one variable and the
horsepowers of used cars in the second variable. In this case, use the right half of the dialog box to
select the variables. Here are the results:

44 • Compare 2 Groups WinSTAT User's Manual


The output includes the mean rank and the number of cases for each group. The U-value is a statistical
measure of the difference in mean rank. This value is then translated to a Z-value, which is the
abscissa of the equivalent point on the Gaussian normal curve. The Z-value is used to compute the
significances. P indicates the significance of the U-value, being the probability that the observed
difference in mean ranks could be the result of random fluctuations in the dependent variable rather
than of a true dependency.
In the example, we see that new cars tend to be bought with more horsepower than used cars, and that
the results are significant (low p-value).

Compare 2 Groups/Wilcoxon test


This test is similar to the t-test for dependent samples in that, for each test, a pair of variables is
analyzed. Each variable must belong to “the same object” and must measure “the same thing” under
different circumstances. The variables need only be measured at the ordinal level, and this is the main
difference between the t-test and the Wilcoxon test.
For each case, the difference between the two variables is computed. The absolute values of these
differences are then sorted to get a rank order. Finally, the mean rank of the negative differences is
compared to the mean rank of the positive differences. The Wilcoxon test determines to what extent
the difference in mean rank is significant.
Use the dialog box to select the two variables:

Wilcoxon test dialog box

WinSTAT User's Manual Compare 2 Groups • 45


As an example, we have selected the horsepower of a family's first car compared to the horsepower of
the family's second car. It is important to note that we are making comparisons within one family, thus
satisfying the “same object” requirement mentioned above.

The number of negative and positive differences is displayed, as well as the corresponding sum of
ranks. This value is then translated to a Z-value, which is the abscissa of the equivalent point on the
Gaussian normal curve. The Z-value is used to compute the significances. P indicates the significance
of the difference in mean ranks, being the probability that the observed difference could be the result
of random fluctuations in the variables rather than of a true dependency.
In the example, we see that a family's second car has, on average, less horsepower than its first car,
and that this result is highly significant.

46 • Compare 2 Groups WinSTAT User's Manual


Compare 2 Groups/McNemar test
The McNemar test is a comparison of two dependent variables which measure "the same thing" on the
same object under different conditions. The variables must be dichotomous, that is, allow only two
answers (such as a yes/no question). Example: Headache patients are given two different treatments.
Each patient undergoes both treatments at different times, and responds "effective" or "not effective"
to each one:

McNemar worksheet

Fill in the dialog box:

McNemar-test dialog box

WinSTAT User's Manual Compare 2 Groups • 47


The results are:

The McNemar test calculates a chi-square statistic for the differences in response. The P-value reflects
the significance of the difference.
The Kappa index is useful if comparing the judgement of two people pertaining to the same objects.
Suppose, for example, two doctors examine the same set of patients. Each doctor assigns a "yes" or
"no" to each patient, depending on whether he diagnoses a certain symptom or not. The Kappa value
can range from 0 to 1, with high values indicating good agreement between the two variables.

48 • Compare 2 Groups WinSTAT User's Manual


Compare N Groups

Overview
In the previous section, you were introduced to functions which compare exactly two data sets for
significant differences. This next category includes commands to compare more than two data sets
among each other. The general name for such comparisons is Analysis of Variance (ANOVA), and in-
cludes the pure analysis of variance for variables with a normal distribution as well as non-parametric
methods.

WinSTAT User's Manual Compare N Groups • 49


Compare N Groups/Analysis of Variance
This command is used to perform a univariate one-factor or two-factor analysis of variance. This
means that a single variable (measured at the interval level and assumed to have a normal distribution)
is grouped according to the values of one or two independent variables. The arithmetic mean of the
dependent variable will generally be different for the different groups. The analysis of variance
determines to what extent the difference is significant. The analysis of variance is valid only if the
groups can be assumed to have equal variances. A Bartlett test is automatically performed to test this
assumption.
Use the dialog box to select the variables:

Analysis of variance dialog box

Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Price as the dependent variable and Type as the independent variable. Sometimes,
however, you may have already sorted the dependent data itself into several columns. To go along
with the example, you would have the prices of sports cars in one variable, the prices of station
wagons in a second variable, and so on. In this case, use the right half of the dialog box to select the
variables.

50 • Compare N Groups WinSTAT User's Manual


One-factor ANOVA
Here, we have chosen Price as the dependent variable and Type of car as the single independent
variable. Thus, we will be performing a one-factor ANOVA. Here is the first section of the results:

D.F. is the number of degrees of freedom. The Mean Sum of Squares is the quotient of the sum of
squares and D.F. The Between groups F-value is the quotient of the mean S.S. (between) and the mean
S.S. (within) and is a measure of the differences in means among the various groups. P indicates the
significance of this difference, it being the probability that the observed difference could be a result of
random fluctuations in the dependent variable rather than of a true dependency. Under the Bartlett test,
the chi-square value is a measure of the differences in variances among the various groups. P again
indicates the significance. In the example, there is a significant indication that the price of a car
depends on its type.
It should be noted, however, that the significant results indicated in the above example are invalidated
by the Bartlett test, since it indicates a significant difference in the variances. In this case it would be
advisable instead to use the Kruskal-Wallis analysis, described later, since it is not dependent on
homogeneous variances.

Multiple comparisons
If the one-factor ANOVA determines that differences in the means exist, it is often of interest to ask
which of the group means differ significantly from which other group means. Or put another way,
which subsets of groups can be built whose members have means which do not differ significantly
from each other. The answer to this question is provided by a set of methods called “multiple compari-
sons,” also known as “a posteriori” or “post hoc” tests.
All of the methods listed follow the same general pattern: A range statistic is calculated for each
possible subset size, and if the greatest difference in means within a subset is less than the calculated
range, then the subset is considered to contain members which do not differ significantly from each
other. Various theories regarding the calculation of the ranges are reflected in the different methods.

WinSTAT User's Manual Compare N Groups • 51


LSD (Least Significant Difference): The critical range for each pair of groups is calculated using a t-
distribution, regardless of the subset size.
A weakness of the LSD method lies in the fact that, with the large number of comparisons being made,
the probability of falsely labeling at least one of the differences as significant increases. This is called
the experimentwise error rate, and it can be reduced by adjusting the significance level to result in
greater critical range values.
Bonferroni: The adjusted significance level pa is calculated according to the formula:

p
pa =
1
g ( g − 1)
2
where g is the total number of groups. The critical range for each pair of groups is then calculated
using the adjusted level in a t-distribution.
Scheffé: The range for any pair of groups is calculated using an F-distribution. This method is the
most conservative in that it produces the largest ranges of all methods, thus requiring a greater
separation of means before the difference is considered significant.
Tukey: The critical range for each pair of groups is calculated using the studentized range distribution,
which takes into account the total number of groups present.
B-Tukey: Also known as modified Tukey, the critical range for each comparison is simply the average
between the standard Tukey value and the value calculated by S-N-K.
S-N-K: Named after the authors Student, Newman, and Keuls, this statistic also calculates the
studentized range, but uses the number of groups in each subset being investigated rather than the total
number of groups.
Duncan: As in S-N-K, this method calculates the studentized range using the number of groups in the
subset. In addition, it adjusts the significance level to account for the fact that many comparisons are
being made. The formula is:

p n = 1 − (1 − p ) n
where n is the number of groups in the subset being investigated.
REGW: Named after the authors Ryan, Einot, Gabriel, and Welsch, this method is similar to Duncan
in that the significance level is adjusted before the studentized range is calculated. The formula is:
n
p n = 1 − (1 − p ) g

where n is the number of groups in the subset being investigated and g is the total number of groups.

52 • Compare N Groups WinSTAT User's Manual


The multiple comparison results appear in the second half of the table. The results for our example,
using S-N-K, are:

Note that you can click within the worksheet itself to change the method and p-value.
The question as to which types of cars are responsible for the significant difference in price is
answered by the S-N-K multiple comparison test. Note that the table is sorted according to price, that
is, from station wagon (least expensive) to sports car (most expensive). To explain the table, let's
look at one example. In the first row (station wagon), we see the number 2781.80 under sedan. Read
this to mean: within the subset containing all car types from station wagon to sedan (and including
pickup, since its price lies between the two), we may call the differences significant only if the means
vary by more than 2781.80. Since the mean of station wagon is 4562.56 and the mean of sedan is
6591.88, and since the difference between the two is less than 2781.80, there are no significant
differences among these three types. Accordingly, there is a no in the corresponding lower-left
position (in the sedan row under station wagon).
By analyzing the table, WinSTAT can calculate subsets of the grouping variable which are
significantly different from each other. In our example, one subset consists of the sports cars and the
other subset consists of all other car types.

Two-factor ANOVA
The goal of a two-factor (or two-way) ANOVA is to look at the effect of one variable after controlling
for the effect of another variable. For instance, we may want to compare the average mileage of the
cars from different manufacturers after controlling for the number of cylinders.

WinSTAT User's Manual Compare N Groups • 53


If the first factor groups the dependent variable into m groups and the second factor groups the
dependent variable into n groups, then an m x n table is created. Each cell contains all cases with a
given value-pair for the two factors, and the arithmetic mean is calculated for all cases of the
dependent variable within each cell. If each cell contains the same number of cases, then the
experiment is called balanced. This, of course, will usually be the case only if the experiment was
designed with this in mind (e.g. randomized block design).
WinSTAT uses the Method of Unweighted Means, which assumes that the data are fairly well
balanced. To be more precise, the method is only valid if there is not more than a twofold variation in
the number of observations in the row-column combinations.
Important: If the twofold variation is exceeded, the results are not valid. We are then dealing with an
unbalanced two-factor design. There are ways of analyzing such data, requiring you to create
“dummy” variables and then perform a multiple regression. Refer, for example, to:
J.D. Jobson, Applied Multivariate Data Analysis
Volume I, Springer-Verlag, 1991
for a discussion on the use of dummy variables. Such variables can be created and then analyzed with
the multiple regression techniques offered by WinSTAT.
In the sample file, [Link], we could try to look for a significant difference in mileage according
to the type of car, after controlling for whether the car is new or used:

Two-factor ANOVA

54 • Compare N Groups WinSTAT User's Manual


A warning is issued to inform you that the data is overly unbalanced and that therefore the analysis is
invalid.
Sometimes, an experiment is designed such that there is only one measurement per cell in the table. In
this case, it is not possible to calculate the effect of the Interaction between the two variables. Such a
design, then, should only be approached if you are convinced that the interaction is negligible. In this
case, uncheck the Interaction box, and the calculation will be omitted from the analysis.

WinSTAT User's Manual Compare N Groups • 55


Compare N Groups/Repeated measures
This command performs a repeated measures analysis of variance for dependent samples. The analysis
examines a given set of variables. These variables must be measured at the interval level and have a
normal distribution. Each variable must belong to “the same object” and must measure “the same
thing” under different circumstances. For example, the cholesterol levels of patients at several different
time periods after a given treatment. The similarity to the t-test for dependent samples is evident, this
being an extension to the n-variable case. The repeated measures analysis of variance determines to
what extent the differences among the variables are significant.
A dialog box appears, allowing you to select the variables of interest. Here are the results for Price and
Price(2):

D.F. is the number of degrees of freedom. The Mean S.S. is the quotient of the sum of squares and
D.F. F is the quotient of Mean S.S. and the Error Mean S.S. The Between Groups F-value is a
measure of the differences in means among the various variables and is the statistic of interest. P
indicates the significance of this difference, it being the probability that the observed difference could
be a result of random fluctuations in the variables rather than of a true dependency.
The demonstration file has at most two variables that measure “the same thing,” in this example Price
and Price(2). With only two variables, the comparison can be carried out more efficiently with the
T-test (dependent) command. If you compare the results of the two methods, you will see that the p-
values are identical.

Split Plot Repeated Measures


The analysis can be extended to "split plot repeated measures", a form of two-way analysis of
variance. In this analysis, the cases are divided into groups by a separate grouping variable, and the
results include information about the difference between groups. Use the normal repeated measures
dialog box and select the additional grouping variable.

Compare N Groups/H-test (Kruskal-Wallis)


This test is similar to a one-factor analysis of variance, in that the values for a given dependent
variable are placed into groups according to a given independent variable. The dependent variable

56 • Compare N Groups WinSTAT User's Manual


need only be measured at the ordinal level, and this is the main difference between the Kruskal-Wallis
test and the one-factor analysis of variance.
All cases of the dependent variable are first sorted into ascending order. Each value is then replaced by
its position after sorting (its rank). Thus, the lowest value becomes 1, the next lowest becomes 2, etc.
These rank values are then separated into groups according to the independent variable. Each group
has its own mean, this being the mean rank of the original values, and the Kruskal-Wallis test
determines to what extent the difference in mean rank is significant.

H-test (Kruskal-Wallis) dialog box

Depending on how your data is organized, there are two ways of filling in the dialog box. Use the left
half of the dialog box if, as in the example shown, the grouping characteristic is a variable of its own.
Here we choose Horsepower as the dependent variable and Type as the independent variable.
Sometimes, however, you may have already sorted the dependent data itself into several columns. To
go along with the example, you would have the horsepowers of sports cars in one variable, the
horsepowers of station wagons in a second variable, and so on. In this case, use the right half of the
dialog box to select the variables.

WinSTAT User's Manual Compare N Groups • 57


The results are:

H is a measure of the difference in mean rank among the groups. D.F. is the number of degrees of
freedom, and P indicates the significance, it being the probability that the observed H-value could be
the result of random fluctuations in the dependent variable rather than of a true dependency.

Compare N Groups/Friedman
This test is similar to a repeated measures analysis of variance, in that several variables are examined,
each of which must belong to “the same object” and measure “the same thing” under different
circumstances. The variables need only be measured at the ordinal level, and this is the main
difference between repeated measures and the Friedman test. An example of such variables could be
the ratings of three similar products, each case representing the ratings of one expert.
For each case, the values of the given variables are inspected. The variable with the lowest value has
its value replaced by 1, the second-lowest variable becomes 2, etc. After all cases have been handled,
the mean of each variable is computed, this being its mean rank. The Friedman test determines to what
extent the differences in mean rank are significant.
Use the dialog box to select the desired variables. Choosing Horsepower and Horsepower(2) the
results are:

58 • Compare N Groups WinSTAT User's Manual


The chi-square value is a measure of the difference in mean rank among the variables. D.F. is the
number of degrees of freedom, and P indicates the significance, it being the probability that the
observed chi-square could be the result of random fluctuations in the variables rather than of a true
dependency.
The demonstration file has at most two variables that measure “the same thing,” in this example
Horsepower and Horsepower(2). With only two variables, the comparison can be carried out more
efficiently with the Wilcoxon test command, which will also yield a better p-value. You may confirm
this by comparing the results of the two methods.

Compare N Groups/Cochran Q-Test


The Cochran Q-test is a comparison of more than two dependent variables which measure "the same
thing" on the same object under different conditions. The variables must be dichotomous, that is, allow
only two answers (such as a yes/no question). Example: Headache patients are given three different
treatments. Each patient undergoes all three treatments at different times, and responds "effective" or
"not effective" to each one. This is an extension of the McNemar test, for which an example is given
above.
The Q-test calculates a chi-square statistic for the differences in response. The P-value reflects the
significance of the difference.

WinSTAT User's Manual Compare N Groups • 59


Correlation

Overview
This submenu gives you access to all correlation functions. Correlation is a measure of the degree of
dependency between two variables. It says nothing about the cause of the dependency, which could be
some third factor entirely. The following sections describe each command in detail.

The Correlation Dialog Box


The dialog box used for all correlation commands (except cross-correlation) looks like this:

Correlation dialog box

The results of a correlation are displayed as a table of rows and columns. Each cell contains
information about the correlation between the row variable and column variable. Use the dialog box to
select which variables should appear as rows and which as columns. If you wish, use the Copy
pushbutton to make the column variables the same as the row variables already selected.

WinSTAT User's Manual Correlation • 61


Correlation/Pearson
Use this command to compute the Pearson correlation coefficients among selected variables. The
Pearson correlation is a measure of the linear dependency between the variables, which must be
measured at the interval level.
The results are displayed as in this example:

The first entry in each cell is the correlation coefficient itself. The second entry is the number of cases
that entered into the calculation (those cases which have non-missing values for both variables in
question). The last entry is the significance of the correlation. The p-value decreases with increasing
correlation and with increasing case count, indicating increased significance.
Note that if the exact same variables for rows and columns are chosen, the two “consistency” measures
Cronbach’s Alpha and Scott’s Homogeneity Quotient are included in the analysis.

62 • Correlation WinSTAT User's Manual


Correlation/Spearman rank
Use this command to calculate the Spearman correlation coefficients among selected variables. The
variables need only be measured at the ordinal level.
The results are displayed as in this example:

Refer to the PEARSON command for a description of the table contents.

WinSTAT User's Manual Correlation • 63


Correlation/Kendall's tau
Use this command to calculate the Kendall's Tau among selected variables. As with the Spearman rank
correlation, the variables need only be measured at the ordinal level.
The results are displayed as in this example:

Refer to the PEARSON command for a description of the table contents. You may wish to compare this
example with the Spearman rank correlation and note the similarity of the results.

64 • Correlation WinSTAT User's Manual


Correlation/Partial
This command calculates the partial correlation between selected variables and produces a table of
these values. The partial correlation is a measure of the linear dependency between two variables,
where the influence of a third variable is “partialled out.” If you suspect that the correlation results of
two variables are being clouded because they are both related to a third variable, use this command to
find out. As with Pearson correlation, the variables must be measured at the interval level.
Within the Correlation dialog box, you will be asked to select the control variable. Here is an example,
with the variable Income partialled out:

Partail correlation dialog box

WinSTAT User's Manual Correlation • 65


Comparing this table to that of the plain PEARSON command, note for example that the correlation
between Children and Price has changed from -0.226 to -0.547 and has become much more sig-
nificant. We now have more reason to believe that, for a given income, an increasing number of
children has a negative effect on the amount of money spent on the family car. The effect wasn't as
obvious before partialling out Income, because family income itself plays a signicant role in the
number of children and the price of the car.

Correlation/Cross-correlation
This command is best described by an example. Suppose we take temperature readings every x
minutes in two neighboring rooms. One of the rooms is heated and cooled directly, the adjacent room
indirectly through its proximity to the first room. Obviously, the two temperature variables will be
highly correlated, but we would expect the temperature in the second room to lag behind the
temperature in the primary room by some number of time units. That is, the correlation would be even
higher if we could shift the variables in time with respect to one another.
This is exactly what a cross-correlation accomplishes. It shifts the variables one unit (case) at a time
forward and backward and calculates the resulting Pearson correlation coefficient, up to a prede-
termined number of lags.
Look at the sample file [Link]. The source of the data is
G.E.P. Box and G.M. Jenkins, Time Series Analysis: Forecasting and Control, Holden Day, 1976.

Leading indicator and sales

The variable Indicator represents a financial leading indicator which is supposed to predict the
volume of future sales. The variables Sales is the actual monthly sales figure. We expect a correlation
between the two variables, and would hope to find a maximum correlation at some time lag,
confirming that the indicator can be used as a predictor.

66 • Correlation WinSTAT User's Manual


Fill in the dialog box:

Cross-correlation dialog box

This results in the following graphical display (note that the CROSS-CORRELATION command always
produces a diagram as well as a table):

Cross-correlation Indicator - Sales


1.2

0.8
Correlation

0.6

0.4

0.2

0
-30 -20 -10 0 10 20 30
Lag

Cross-correlation Indicator vs. Sales

Rule: a high correlation at a negative lag means that the 1st variable can be used to predict the 2nd
variable. A high correlation at a positive lag means that the 2nd variable can be used to predict the 1st
variable.
It is clear that the indicator correlates with actual sales many months in advance. However, an exact
interpretation is impossible, because we may assume that both the indicator and the sales figures

WinSTAT User's Manual Correlation • 67


exhibit a high autocorrelation (i.e. sales will be high if the previous month was high), thus clouding the
issue.

68 • Correlation WinSTAT User's Manual


Regression

Overview
The various commands on this submenu try to describe one variable (the dependent variable) as a
function of one or more other variables. The following sections describe each command in detail.

Regression/Simple
Use this command to look for simple equations relating a given dependent variable to a single
independent variable. Various classes of equations may be taken into consideration. Fill in the dialog
box appropriately:

Simple regression dialog box

In this example, we will model Mileage as a function of Horsepower, looking for the best fit among
all of the equation classes listed.

WinSTAT User's Manual Regression • 69


Note the checkbox Write residues in. The residuals are the differences on a case by case basis
between the actual Y values and the values calculated by the best-fit equation (the values which lie on
the regression curve). By checking the box and choosing a variable name, we cause WinSTAT to write
the residuals to the worksheet under the variable of the given name. This variable must already exist in
the database, and its previoius contents will be overwritten.
The results of the analysis appear as follows:

R is the correlation coefficient between the actual Y values and the values calculated by the given
equation. High values of R (maximum = 1.0) indicate a good fit. The square of R appears in the next
column and, in statistical terms, is the percentage of variance in the dependent variable which can be
explained by the given equation.

70 • Regression WinSTAT User's Manual


The equation yielding the highest R is underlined. It is this equation which is used to create the
residuals, if specified. WinSTAT automatically produce a scatterplot of the two variables, with the
best-fit curve (the underlined one in the table) superimposed:

Data Y = 10.0919 + 396.081/X

35

30

25

20
Mileage

15

10

0
0 50 100 150 200 250 300
Horsepow er

Best-fit simple regression curve

Residual Analysis
Having created the new variable Regression residuals, we can perform some tests which may lead to
more information about the quality of our best-fit function.
Normality: the residuals should be normally distributed. This can be tested with the
GRAPHICS/CUMULATIVE FREQUENCY command or with the BASICS/TEST OF RANDOMNESS command.
In our example, the Kolmogorov-Smirnov p-value of 0.97925 indicates normality.
Randomness: the residuals should exhibit randomness. Even if the input variables show trend or
periodicity, this fact should not influence the residuals. If it does, then we have missed something
important in modelling the data.

WinSTAT User's Manual Regression • 71


Plot of X vs. Residuals: If we produce a scatterplot of the independent variable vs. the residuals, we
should get a set of points distributed approximately equally to both sides of the X-axis, for the entire
range of X-values.

12
10
8
Regression Residuals

6
4
2
0
-2 0 50 100 150 200 250 300
-4
-6
-8
-10
Horsepow er

Scatterplot of X vs. residuals

In the above plot, it appears that the variance of the residuals decreases as X (Horsepower) increases.
However, since there are many fewer data points for the higher X-values, this is probably not
significant, especially since the normality of the residuals has already been demonstrated.

Prediction of unknown Y-values


In the original dialog box, note the checkbox Recalculate cursive rows and overwrite Y-values. If
this box is checked, WinSTAT will first calculate the best-fit equation as always, but ignoring any data
contained in rows whose data are in a cursive font. Then it will take all X-values within these cursive
rows and, applying the best-fit equation, calculate a value for the Y-variable. Use this feature to
predict values for cases in which only the independent variable is known.
Use this feature with care, since any previous (cursive) Y-values will be overwritten.

72 • Regression WinSTAT User's Manual


Regression/Polynomial
This command looks for the best polynomial equation relating a given dependent variable to a single
independent variable. The degree of the polynomial to be taken into consideration may be freely
chosen.

Polynomial regression dialog box

See REGRESSION/SIMPLE for a discussion of the Write residues in and Recalculate cursive rows and
overwrite Y-values checkboxes. By checking Constant = 0, the regression equation is forced to pass
through the origin.

WinSTAT User's Manual Regression • 73


Weighted Regression
By the nature of the measuring process, we sometimes know that some measurements are more precise
than others, and would like to have this fact considered in the regression calculation. More precise
measurements should be give more weight than imprecise measurements. In the case of Mileage, it is
clear that the higher the mileage is, the more precise the measurement, simply because we will be
reading more significant places off the odometer. This is an inverse rule, which is also typical for
regression when applied to the calibration of laboratory instruments. In the Excel worksheet, we can
insert a new column to the right of Mileage and name it Inverse Mileage. Then, using the function
definition and fill-down capabilites of Excel, we can define the new variable to be the inverse of
Mileage. Selecting this variable as the weighting variable, we get the following results:

R and R square have the same meaning as in simple regression. The corrected values (in parentheses)
include a “punishment factor” for each increased number of terms. Obviously, a better fit (higher R)
can always be found by increasing the order of the polynomial, but at the expense of simplicity and
ability to interpret the curve. Maximizing the corrected value of R is a method of deciding when to
stop.
Std. error is the standard error of the regression curve (compared to the observed values), measured in
units of the dependent variable.
For each term, the calculated coefficient is displayed as well as a confidence interval for the
coefficient. One may be confident that the true value lies within the indicated bounds. Note that the
percent value (here 95%) may be changed by clicking within the field. T is the quotient of the
coefficient and its standard error. P indicates the significance of the given term in the equation.

74 • Regression WinSTAT User's Manual


WinSTAT automatically produce a scatterplot of the two variables, with the best-fit curve
superimposed:

Data Y = 31.3976 - 0.34087*X + 2.00323E-03*X^2 - 3.79134E-06*X^3

35

30

25

20
Mileage

15

10

0
0 50 100 150 200 250 300
Horsepow er

3rd-degree polynomial regression

WinSTAT User's Manual Regression • 75


Regression/Multiple
This command looks for the best linear equation relating a given dependent variable to any number of
independent variables. It is also possible to choose among several methods of stepwise regression, in
which the program decides on its own, according to certain criteria, which of the independent variables
are to be included in the regression equation.
Fill in the dialog box:

Multiple regression dialog box

See REGRESSION/POLYNOMIAL for a discussion of the Write residues in, Recalculate cursive rows and
overwrite Y-values, Y-weightings and Constant = 0 checkboxes.
The methods available are:

Direct
All of the independent variables chosen are included in the regression equation. This is standard
multiple regression.

Stepwise
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. After each addition, all
variables now in the equation are reinvestigated. The variable with the least significance (highest p-
value) is then removed from the equation, assuming its p-value is higher than p-out from the dialog

76 • Regression WinSTAT User's Manual


box. The process continues until no variables can be added or removed or until the maximum number
of variables (as indicated by Maximum number of variables in the dialog box) has been reached.

Forward
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. The process continues until
no more variables can be added or until the maximum number of variables (as indicated by Maximum
number of variables in the dialog box) has been reached.

Backward
Starting with all independent variables in the regression equation, the program investigates each one.
The variable with the least significance (highest p-value) is then removed from the equation, assuming
its p-value is higher than p-out from the dialog box. The process continues until no more variables can
be removed.

Maximum R-Square
This method demands significantly more calculation time than the others, but it is guaranteed to find
the “best” (in the R-Square sense) one-variable solution, then the best two-variable solution, etc. At
each step it may remove and add variables as it finds better solutions.

Output
As an example, we choose Horsepower as the dependent variable, and Income and Price as the
independent variables. The Stepwise method is selected, with a maximum of two variables. The
results are:

WinSTAT User's Manual Regression • 77


At each step, an independent variable may be added to (+) or removed from (-) the equation. In the
example, it was never necessary to remove a variable once it had been added. P is the p-value for the
given variable, always less than p-in if the variable is being added or greater than p-out if the variable
is being removed.
R square indicates the extent to which the dependent variable can be determined by the independent
variables at each step, and of course increases as new variables are added. The Corrected value for R-
square (in parentheses) includes a “punishment factor” for each increased number of variables.
Obviously, a better fit (higher R) can always be found by increasing the number of explanatory
variables, but at the expense of simplicity and ability to interpret the results.
In the example, the stepwise regression stops right after the inclusion of the first variable, meaning that
none of the other variables under consideration satisfies the p-in requirement. When the process stops,
complete information about the regression equation is displayed.
Std. error is the standard error of the regression curve (compared to the observed values), measured in
units of the dependent variable.
For each variable in the equation, the calculated coefficient is displayed as well as a confidence
interval for the coefficient. One may be confident that the true value lies within the indicated bounds.
Note that the percent value (here 95%) may be changed by clicking within the field. T is the quotient
of the coefficient and its standard error. P indicates the significance of the given term in the equation.
Finally, the analysis of variance for the entire regression is displayed.
Why is it that the p-value for a given variable changes from step to step? In general, the independent
variables not only correlate with the dependent variable, but also with each other. The addition of a
new variable can thus reduce the importance of a variable already in the equation. This is why, in a
stepwise regression, a variable can at some point be excluded from the equation after it has been
accepted, even if p-in and p-out are equal.

78 • Regression WinSTAT User's Manual


Discriminant Analysis

Overview
In Discriminant Analysis, it is assumed that the data at hand are grouped according to a given nominal-
level variable. Then, any number of further variables are examined to see to what extent they can be
used to “predict” the value of the grouping variable. The similarity to multiple regression is apparent,
the difference being that the variable to be “predicted” in multiple regression is interval-level.
Once the analysis has been performed, the derived equations can be used to classify cases which have
not yet been assigned to groups. A classic example is the bond market. First, discriminant analysis is
performed using bonds of known quality (AAA, AA, ..., C), relating these ratings to various financial
indicators of the respective companies. Following this, the financial indicators of previously unrated
companies can be fed into the equations to get bond ratings for those companies.

Discriminant Analysis
Fill in the dialog box:

Discriminant analysis dialog box

The Prior Probabilities are only of meaning when using the derived functions for classification
purposes. If all equal, the group closest to the value calculated by the equations will be chosen for the
classification. If proportional to present frequencies, a bias is installed which makes classification

WinSTAT User's Manual Discriminant Analysis • 79


more likely into those groups which represent more cases in the data at hand. For example, if the equa-
tions yield a value exactly between two groups, that group will be chosen which is “more likely”
according to the present data.
If the Recalculate cursive rows and overwrite Y-values box is checked, WinSTAT will first cal-
culate the best-fit equation as always, but ignoring any data contained in rows whose data is in a
cursive font. Then it will take all X-values within these cursive rows and, applying the best-fit
equation, calculate a value for the Y-variable. Use this feature to predict values for cases in which only
the independent variables are known. Use this feature with care, since any previous (cursive) Y-values
will be overwritten.
Choosing Type as the grouping variable, and the variables Size of Household, Income, and Price as
the independent variables, and, furthermore, selecting prior probabilities equal, the results of the dis-
criminant analysis appear as follows:

80 • Discriminant Analysis WinSTAT User's Manual


WinSTAT User's Manual Discriminant Analysis • 81
The results may be interpreted as follows:
Eigenvalues: In discriminant analysis, several functions are calculated which describe the relationship
between the grouping variable and the independent variables. These are known as the “canonical
discriminant functions.” The number of functions is equal to the number of groups minus one or the
number of independent variables, whichever is smaller. The different functions tend to contrast various
groups of variables which help “predict” the grouping variable. Some functions are more effective
than others, and the effectiveness of a function is reflected in its eigenvalue. The higher the eigenvalue,
the greater the percentage of the grouping variable's variance that is explained by the function. The
functions are always sorted according to the highest eigenvalue.
Canonical Correlation: This value can be directly derived from the eigenvalue, C.C. = sqrt(
eigenvalue / (1+eigenvalue) ), and is a measure of the overall dependency of the grouping variable on
the given function.
Wilks' Lambda: This value is obtained by subtracting the squared C.C. from 1, then multiplying this
value by the Wilks' Lambdas of all following functions. It may be transformed to an equivalent Chi-
Square statistic, which is used for calculating P.
P: This is the probability that, if there were, in fact, no dependency between the grouping variable and
the independent variables in the general population, the given eigenvalue (and those following it)
could reach the calculated values in the sample by chance alone. In the example, the eigenvalue of the
first function differs significantly from 0, that of the second function less so, and the eigenvalue of the
third function is not significantly different from 0.
Standardized coefficients of the discriminant functions: This is a table of the calculated coefficents
for all functions. “Standardized” means that the independent variables have been transformed to vari-
ables with mean = 0 and variance = 1 before the coefficients are applied. The signs of the coefficients
may be used to group the independent variables. That is, the variables with positive coefficients tend to
“push” in one direction, and those with negative coefficients in another. The absolute value of a
coefficient is a measure of the importance of that variable within the given function, since a low
coefficient is not going to affect the outcome as much as a high coefficient. In the example, we would
interpret the coefficients of the first function as meaning that a large household “pushes” for one type
of car, and a high price for another. What those types might be is answered by the next table.
Values of the discriminant functions at group centroids: For each of the groups, the discriminant
functions can be calculated for all cases within the group and then averaged. The results are displayed

82 • Discriminant Analysis WinSTAT User's Manual


in this table. In the example, the extreme values for the first function are given by “station wagon” and
“sports car.” We conclude that a large household is an indicator for a station wagon and a high price is
an indicator for a sports car.
Mahalanobis distances between the groups: The Mahalanobis distance is a measure of the
difference between the average function values from group to group, with all functions taken together.
A large Mahalanobis distance between two groups indicates that the two groups are well separated by
the functions, and it is unlikely that a case belonging to one group would be falsely classified by the
functions as belonging to the other group. A small distance between two groups indicates that the
functions cannot discriminate well between them. The P-values measure the probability that the given
distance could occur by chance in the sample, when in fact there is no difference in the general
population.
Classification results: In this table (also known as a “confusion table”), the discriminant functions are
used to classify all cases in the file, and the resulting values are compared to the actual group mem-
bership. For example, of the 10 sports car owners, 2 were falsely classified as station wagon owners
and 8 were correctly classified as sport car owners.

WinSTAT User's Manual Discriminant Analysis • 83


Cluster Analysis

Overview
Cluster analysis refers to methods which attempt to group cases in such a manner that the members of
each group are, in some sense, “close” to one another. Several variables may be chosen for the analy-
sis, and the differences in these variables between two cases determines the “distance” between the
two cases.
The most common method of calculating the distance is, for the two cases in question, to sum the
squared differences of each of the variables in question. This is known as the squared Euclidian
distance. Before the distances are calculated, the values of all variables in question are standardized
(that is, transformed to variables with mean = 0 and variance = 1), so that the variables with greater
values and variances do not dominate the equation. Note that WinSTAT automatically standardizes the
variables before calculating the Euclidian distances.
Once the distances have been obtained, there are several methods for collecting the cases into groups.
Because the groups grow, as cases (and groups) are combined, and because, once in a group, a case is
never removed, these methods are known as agglomeration methods. The various agglomeration
methods are described in detail below.

WinSTAT User's Manual Cluster Analysis • 85


Agglomeration
The following dialog box appears:

Cluster analysis (agglomeration) dialog box

Select the variables which are to determine the distances, as well as one of the agglomeration methods.
Also, if one of the file's variables contains labels which identify each case, you may indicate this using
the checkbox. This results in the case labels being used for all output instead of case numbers.

Single linkage
This method (also known as nearest neighbor) says that the distance between two groups is simply the
distance between the two closest members of the groups.

Complete linkage
This method (also known as farthest neighbor) says that the distance between two groups is equal to
the distance between the two farthest members of the groups.

Average linkage (UPGMA)


This method says that the distance between two groups is equal to the average distance determined
over all possible pairs of members.

Centroid method
This method says that the distance between two groups is equal to the distance between the two group
centroids. A group centroid is its center of gravity measured over all variables.

Ward's method
This method (also known as incremental sums of squares) says that the distance between two groups is
proportional to the change in the within group sum of squares (see Analysis of Variance) which results
when the two groups are combined.

86 • Cluster Analysis WinSTAT User's Manual


Agglomeration results
Once the desired method has been selected, groups are automatically agglomerated until all cases are
in one single group. This process takes as many steps as there are cases present in the file. The first ten
steps for Ward's method with the variables Income, Price, and Horsepower are:

At each step, a combination of two clusters occurs, and so the total number of clusters is reduced by
one. The two columns joining Cluster 1 and with Cluster 2 indicate the two clusters which are being
combined at this step. The number displayed is the lowest case number of all cases in the cluster. If
case names were indicated in the dialog box, then the labels are displayed instead of numbers. The
number in parentheses indicates the number of cases presently in each of the clusters. Finally, the last
column displays the distance between the two clusters as calculated by the method chosen.

WinSTAT User's Manual Cluster Analysis • 87


The clustering process is also displayed graphically in what is known as a dendogram (because of its
tree-like form):

Case num ber

0
10
20
30
40
Distance

50
60
70
80
90
100

Dendogram

A main purpose of the dendogram is as a tool in deciding how many clusters one wants to retain.
Looking at the complete dendogram, we see that the reduction of three groups to two groups (at
distance 38) was only possible with a considerable increase in distance compared to previous
combinations. That is, two clusters were combined which were much farther apart than the clusters
which had been combined up to that point. It would appear logical, therefore, to retain three clusters.
Note that cluster separation is not available for the neighbor joining method
Once the number of clusters has been decided upon, click on the Cluster separation button to assign
cases to respective groups.

Cluster separation button

88 • Cluster Analysis WinSTAT User's Manual


Cluster separation
Fill in the dialog box:

Cluster separation dialog box

You must indicate the number of clusters and the name of a variable which will contain the cluster
number of each case. This variable name must previously have been inserted into the worksheet
database. In this example, a new variable Ward's clusters has been added. Each case in cluster 1 is
given the value 1, each case in cluster 2 the value 2, etc. This variable may be used to group cases in
further analyses, such as discriminant analysis.
Use the GRAPHICS/SCATTERPLOT command, using the new variable as Grouping variable, to get a
plot of the data as grouped.

Ward's Clusters

70000
1
60000 2
3
50000

40000
Income

30000

20000

10000

0
0 50 100 150 200 250 300
Horsepow er

Scatterplot with clusters

WinSTAT User's Manual Cluster Analysis • 89


Factor Analysis

Overview
Factor Analysis is used to look for basic, independent dimensions underlying the existing variables.
One tries to find the smallest set of such dimensions which nevertheless explain the variables to a
sufficient degree. There are many options to consider when preparing a factor analysis, and often a
satisfactory solution will only be found by repeating the analysis a number of times with different
parameters. You should have a general idea ahead of time which variables are to be explained and the
approximate number of factors which may lie behind these variables.

Factor Analysis
Fill in the dialog box:

Factor analysis dialog box

WinSTAT User's Manual Factor Analysis • 91


First, choose the variables which are to included in the analysis. Then select among the indicated
options:
Initial Communalities: The communality of a variable is a measure of the extent to which that
variable can be determined by the calculated factors. This, of course, is not known in advance, since
the factors have yet to be calculated, but there are two popular methods for making a first estimate:
1. Assume communalities of 1.0 (principal components method).
2. Assume a communality for each variable equal to the highest correlation coefficient
involving that variable (principal axis method).
Calculated Communalities: The factor analysis, when complete, delivers values for the actual
communalities as yielded by the computed factors. If the computed communalities do not correspond
well with the estimated communalities, you may wish to repeat the analysis with the revised values,
since the initial assumption was apparently weak. This process can be continued until a given
correspondence level is reached. A standard acceptance level is a difference no greater than 0.05
between the estimated values and the computed values, but there is no guarantee that this level can be
reached with a restricted number of factors.
Factor Extraction: A decision must be made as to how many factors are to be extracted. One standard
method is to have the program continue extracting factors as long as the calculated eigenvalue of the
factor is greater than 1.0. This is equivalent to saying that each factor must contain at least as much
information as any single variable. Of course, any other minimum eigenvalue may be specified. You
may limit the maximum number of factors extracted in this process by filling in the appropriate field.
Alternatively, a predetermined number of factors may be extracted. In this case, set the minimum
eigenvalue to 0 and the maximum number of factors to the predetermined number. In general, the
desired number of factors will also depend on how well the calculated factor loadings can be
interpreted. Thus, several analyses can be performed, each with a different number of factors, and a
subjective decision reached as to which is best.
Rotation: One goal of factor analysis is to be able to interpret the factors as they relate to groups of
variables. Since the initial calculation of the factors is in a more or less random orientation with
respect to the variables, it is desirable to manipulate the factor axes in such a way that the majority of
variables lie close to one of the axes. This process is called factor rotation, and the two most popular
methods are Varimax and Quartimax. Before rotation, the Kaiser normalization may be applied to
the factor loadings. This procedure assures that the variables with large communalities are given
greater weight during the rotation.

92 • Factor Analysis WinSTAT User's Manual


In the output of eigenvalues, the factors which have been accepted according to the options chosen
appear in bold. In this case, they are the factors with an eigenvalue of at least 1.0.

WinSTAT User's Manual Factor Analysis • 93


Note that the factor analysis had to be repeated two times before the residuals of the communalities
became acceptable. Two factors were extracted, and the rotated factor loadings yield the following
groups of variables:
Factor 1 (Family Situation):
Size of Household
Children
Income
Factor 2 (Car):
Horsepower
Price
Mileage (negative)

In all, 76 percent of the variance of all selected variables can be explained by these two factors.
WinSTAT also produces a graphical display of the factor loadings. Here is the plot of the factor
loadings from the example:

0.5
Size of Household
Children
Factor 2

0 Income
Horsepow er
Price
-0.5 Mileage

-1
-1 -0.5 0 0.5 1
Factor 1

Plot of factor loadings

By viewing the plot, it is easy to ascertain which variables “belong” to only one factor (those that lie
close to an axis), and which variables are determined by information from both factors.

94 • Factor Analysis WinSTAT User's Manual


Factor Scores
If this option is selected, a new variable is automatically created for each of the extracted factors. On a
case-by-case basis, values are assigned to these new variables reflecting the factor values for that case.
These scores can be used instead of the original variables to perform further statistical calculations,
such as cluster analysis or regression.

WinSTAT User's Manual Factor Analysis • 95


Survival Analysis

Overview
WinSTAT offers two survival analysis features: Kaplan-Meier and Cox Regression.
Kaplan-Meier calculates and displays survival tables according to the data at hand and can also check
if a single grouping variable has a significant influence on the survival probability curves.
Cox Regression allows the user to supply several independent variables which may or may not have a
significant influence on the survival probability curves. A regression equation for the influence of
these variables in then calculated. As with multiple linear regression, there are several options which
cause WinSTAT to search for those variables which are significant. Another feature allows the user to
plot a custom survival curve for a single row of data representing one patient or other item.

Kaplan-Meier
The sample file [Link] contains typical data for a survival analysis:

Survival analysis worksheet

WinSTAT User's Manual Survival Analysis • 97


Fill in the dialog box:

Survival analysis dialog box

A minimum of two variables is necessary for any survival analysis:


1. The first variable contains the survival time data. This is usually an integer value representing
a number of days, months, or years. The variable is often defined as the difference between
two dates. The first date is the time when the case (e.g. patient) entered the survey, and the
second date is when the case left the survey.
2. The second necessary variable is called the event variable. This variable indicates whether the
survey ended for the given case because the pre-defined event (e.g. death) occurred, or
whether the case dropped out of the survey for another reason. If the event variable is not
missing and non-zero, it indicates that the event occurred. If missing or zero, it indicates that
the case was censored, that is, dropped out of the survey.
In the example, the variable survival time is calculated using the function:
"last follow-up" - "1st diagnosis"
to get the number of days survived.
In addition, a third variable, the Grouping variable, may be specified. These values of this variable
may represent, for example, different treatments, different geographic locations, or different time
periods.

98 • Survival Analysis WinSTAT User's Manual


Looking first at the sample data, without grouping, the results are printed in the following Kaplan-
Meier life table:

The entries are automatically sorted according to survival time. The column At Risk indicates how
many cases were still in the survey at the given time, before the event or censoring occurred. The next
two columns indicate how many events occurred and how many cases were censored at the given time.
The column Survival probability indicates, for the data at hand, the probability of surviving up to the
given time. The final column gives the standard error for the calculated probabilities.

WinSTAT User's Manual Survival Analysis • 99


WinSTAT also produces a plot of the survival probabilities:

1.2

0.8
Probability

Censored
0.6
Probability

0.4

0.2

0
0 100 200 300 400
Survival Tim e

Survival probabilities

100 • Survival Analysis WinSTAT User's Manual


If you select a grouping variable (in the sample data, this is the variable Treatment), the Kaplan-
Meier life table is printed for each of the two groups. In addition, a log-rank test is performed to check
the significance of the difference between the two tables:

For each group, the expected number of events is calculated under the assumption that there is no
difference between the two groups. Comparing these values with the actual number of events observed
yields a chi-square statistic. P is the significance of this statistic.

WinSTAT User's Manual Survival Analysis • 101


Again, WinSTAT produces a graphical representation of the survival probabilites:

1.2

0.8
Probability

Censored
0.6 Drug A
Placebo
0.4

0.2

0
0 100 200 300 400
Survival Tim e

Survival probabilities of the groups

102 • Survival Analysis WinSTAT User's Manual


Cox Regression
Cox Regression allows the user to supply several independent variables which may or may not have a
significant influence on the survival probability curves. A regression equation for the influence of
these variables in then calculated. As with multiple linear regression, there are several options which
cause WinSTAT to search for those variables which are significant. Another feature allows the user to
plot a custom survival curve for a single row of data representing one patient or other item.
The sample file [Link] contains typical data for a Cox regression:

Cox regression worksheet

Here we see the same data as in the Kaplan-Meier survival analysis, but we also see two new columns
to the right, Treatment numeric and Age. The first of these, Treatment numeric, is nothing other
than a copy of Treatment, but using the numeric coding 0 and 1 instead of the text values Drug A and
Placebo. This is necessary because the regression calculation depends on the mathematical use of
numeric variables. All independent variables to be used in a Cox regression must be numeric.

WinSTAT User's Manual Survival Analysis • 103


What we are interested in is the effect of treatment and age on the life expectancy of the patients
involved. Fill in the dialog box:

Cox regression dialog box

The following variables are necessary for any Cox regression analysis:
1. The first variable contains the survival time data. This is usually an integer value representing
a number of days, months, or years. The variable is often defined as the difference between
two dates. The first date is the time when the case (e.g. patient) entered the survey, and the
second date is when the case left the survey.
2. The second necessary variable is called the event variable. This variable indicates whether the
survey ended for the given case because the pre-defined event (e.g. death) occurred, or
whether the case dropped out of the survey for another reason. If the event variable is not
missing and non-zero, it indicates that the event occurred. If missing or zero, it indicates that
the case was censored, that is, dropped out of the survey.
3. Finally, one or more independent variables which may influence the survival time must be
selected.
As in REGRESSION/MULTIPLE there are several options which determine how WinSTAT decides which
of the independent variables are significant enough to be used in the regression equation:

Direct
All of the independent variables chosen are included in the regression equation. This is standard Cox
regression.

104 • Survival Analysis WinSTAT User's Manual


Stepwise
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. After each addition, all
variables now in the equation are reinvestigated. The variable with the least significance (highest p-
value) is then removed from the equation, assuming its p-value is higher than p-out from the dialog
box. The process continues until no variables can be added or removed or until the maximum number
of variables (as indicated by Maximum number of variables in the dialog box) has been reached.

Forward
Starting with no variables in the regression equation, the program investigates all independent
variables not yet in the equation. The one with the greatest significance (lowest p-value) is then added
to the equation, assuming its p-value is less than p-in from the dialog box. The process continues until
no more variables can be added or until the maximum number of variables (as indicated by Maximum
number of variables in the dialog box) has been reached.

Backward
Starting with all independent variables in the regression equation, the program investigates each one.
The variable with the least significance (highest p-value) is then removed from the equation, assuming
its p-value is higher than p-out from the dialog box. The process continues until no more variables can
be removed.

WinSTAT User's Manual Survival Analysis • 105


Output
The results are printed as follows:

For each variable in the equation, the calculated coefficient is displayed as well as a confidence
interval for the coefficient. One may be confident that the true value lies within the indicated bounds.
Note that the percent value (here 95%) may be changed by clicking within the field. P indicates the
significance of the given term in the equation. Hazard is e to the power of the coefficient and
indicates to what extent the variable increases the risk of the object under examination.
In the example, one can say that receiving a placebo instead of the drug increases the hazard, as does
an increased age. The effect of age is less relevant. It should be noted, however, that neither of these
variables is particularly significant.

106 • Survival Analysis WinSTAT User's Manual


The second table is a representation of the survival curve. At covariate mean is the curve for an
individual who represents the mean of all independent variables. This curve is also displayed:

Survival probabilities for the average individual

Examining a specific new case


The power of Cox regression lies in the ability, using the equation found, to calculate the specific
survival curve for an individual with known values for the independent variables. To see this is action,
add the following data to the original table and, using Excel formatting, make the entire new row
cursive:

A new individual

Thus, we are answering the question: What survival curve can be expected for a 50 year old patient
treated with the placebo?

WinSTAT User's Manual Survival Analysis • 107


Before the analysis may be performed, the WinSTAT Database range must be extended to include the
new data row. This is done easily with DATA/SET DATABASE RANGE, then selecting the option All used
cells of the active worksheet. The analysis may now be repeated with the option Calculate survival
curve for 1st cursive row:

The option “cursive row”

108 • Survival Analysis WinSTAT User's Manual


The new results are:

The column Survival Probability, Cursive case is now filled with the specific probabilities for the
new individual. The probability curve is also displayed, along with the original curve for comparison
purposes:

Probability curve for the new individual

WinSTAT User's Manual Survival Analysis • 109


Process Capability

Process Capability
The process capability ratio compares the allowed variation in a product's specification with the
variation actually found in a series of measurements on the actual product. It is used to estimate the
fraction of defective products which the process will produce.
The command opens the following dialog box (using the file [Link]):

Process capability dialog box

WinSTAT User's Manual Process Capability • 111


Enter the variable containing the series of measurements, as well as the lower and upper specification
limits for the characteristic being measured. The results are:

The program calculates two process capability ratios: Cp and Cpk.


Cp is the ratio of the specification range to six times the standard deviation of the sample series. Since
Cp does not take into account the mean of the sample series, it indicates the potential process
capability which could be reached if the process were centered on the specification mean.
Cpk is the process capability for an off-center process or for a one-sided process. This value will tend
towards Cp as the process mean approaches the specification mean.
Defects is the expected number of defective products (in ppm) which the process (as measured by the
sample series) will produce.
One-sided specifications
Example: If we are measuring the bursting strength of glass bottles, the lower specification limit might
be 200psi. An upper specification would not be provided, since any high strength is allowed. Simply
enter a very high number (one which would not reasonable ever occur) as the upper specification limit,
for instance 2000psi. The resulting Cpk and Defects are then correct.

112 • Process Capability WinSTAT User's Manual


The GRAPHICS Menu

Overview
This menu gives you access to WinSTAT's graphics capabilities. Some graphics, which relate directly
to a given statistical method, are available directly through that statistic (for example, the plot of a
regression curve). Those plots are described along with the corresponding statistics function and not
repeated here.

Graphics/Histogram
The GRAPHICS/HISTOGRAM command is exactly the same as the BASICS/FREQUENCIES command, so
we simply refer you to that section here.

Graphics/Means plot
The GRAPHICS/MEANS command is exactly the same as the BASICS/MEANS command, so we simply
refer you to that section here.

WinSTAT User's Manual The Graphics Menu • 113


Graphics/Box & Whisker
This command functions much like the GRAPHICS/MEANS command, but it yields information
regarding percentiles instead of means. A box & whisker plot of Income and Price appears like this:

70000

60000

50000

40000

30000

20000

10000

0
Income Price

Box & whisker plot

The short line within the box represents the median of the given variable. The bottom and top edges
represent the 25th and 75th percentiles. In other words, 50% of the data fall within the box, and 25%
each above and below. The “whiskers” extend to the 5th and 95th percentiles. Finally, the minimum
and maximum values in the sample are indicated with a '+' sign.

114 • The Graphics Menu WinSTAT User's Manual


Graphics/Scatterplot
This command provides you with a graph in which the observations of two given variables are plotted
case for case against each other. In addition, you may group the cases according to the values of an
additional, grouping variable. Fill in the dialog box:

Scatterplot dialog box

Here is a sample scatterplot, plotting Income and Price against each other and grouping according to
New or Used:

New or Used

30000
new

25000 used

20000
Price

15000

10000

5000

0
0 20000 40000 60000 80000
Incom e

Scatterplot with groups

WinSTAT User's Manual The Graphics Menu • 115


Graphics/Cumulative frequency
This command provides you with a display of the cumulative distribution curve (in percent) for a
single variable. In addition, it superimposes the best-fit normal distribution curve, so that you can
make a visual comparison. After selecting the desired variable, the plot appears as follows:

120.00

100.00
Percent (cumulative)

80.00

60.00

40.00

20.00

0.00
0 50 100 150 200 250 300

Horsepow er Normal

Cumulative plot with normal distribution

116 • The Graphics Menu WinSTAT User's Manual


To the right of the plot, WinSTAT displays an option button which allows you to rescale the Y-axis
such that a probability transformation occurs:

10.00

9.00

8.00
7.00

6.00
Probit

5.00

4.00

3.00
2.00
1.00

0.00
0 50 100 150 200 250 300

Horsepow er Normal

Probability scale

On a probability scale, the Y axis is transformed in such a manner that a normal distribution plots as a
straight line.

WinSTAT User's Manual The Graphics Menu • 117


Graphics/Quality control
WinSTAT offers several types of charts for use in quality control. Each type of chart takes data from
one variable, which are then subdivided into groups of equal size and evaluated. The nature of the
data, group size, and evaluation differ from one chart type to the next. You will find a detailed
description for each chart type below.
The data in the various examples are taken from the book
Douglas C. Montgomery, Introduction to Statistical Quality Control, John Wiley & Sons, 1991
When the QUALITY CONTROL command is selected, the following dialog box appears:

Quality control dialog box

Variable
Select the variable of interest. Note that the examples below refer to several different sample files and
to specific variables within them.

Sample Size
The idea of a sample size is common to all chart types. It is assumed that, at regular intervals, a
constant number of items (the sample) is taken from the production process and measured or tested.
For instance (see X-bar, below), we could measure the inside diameter of 5 piston rings taken from the
production line every 10 minutes. If this is repeated 20 times, the variable will contain 100
measurements (cases) in all, and the sample size is 5. The average measurement of the 5 items (thus
X-bar) will be plotted at each interval, so there will be 20 points on the chart.

118 • The Graphics Menu WinSTAT User's Manual


Control limits
This value determines the limits beyond which we consider the process to be no longer under control.
A certain variance (thus the term sigma = σ = standard deviation) is expected in the items being
measured, but if the measured variable is suddenly “way off,” i.e. too far away from the expected
value, we assume something is wrong. The value specified in the n sigma box determines the limit.
In all examples, we will use 3-sigma limits, which are widely used and have given good results in
practice.

Warning limits
This optional value determines the limits beyond which we consider the process to be in need of
examination. A value which is widely used is 2-sigma.

Specification
In addition to control and warning limits, the actual specification limits for the article being
manufactured may be included in the diagram. This is only applicable when an X-bar chart (see
below) is being drawn.

Using a subset of measurements to determine the control limits


Sometimes it is desired to calculate the control and warning limits using only the data during a specific
interval, when the process is known to be under control. If this is the case, the checkbox Use only the
first must be used.

WinSTAT User's Manual The Graphics Menu • 119


X-bar Chart
The term X-bar refers to the symbol x , meaning a measurement averaged over the sample size. The
file [Link] contains 125 measurements of the inside diameter of piston rings. Assuming a
sample size of 5, this means there were 25 samples taken in all. Selecting Inside diameter as the
variable and setting a sample size of 5, we get the following results:

74.02

74.015

74.01

74.005 Inside diameter


Mean

Mean
74
Control limits

73.995

73.99

73.985
0 10 20
Sam ple, N = 5

X-bar chart

There are 25 data points, each one showing the average diameter of the 5 rings in the corresponding
sample.

120 • The Graphics Menu WinSTAT User's Manual


R Chart
The R refers to range, and the chart is used to display information about the variance within each
sample. The range is simply the difference between the smallest and largest measurement within each
sample. It is generally recommended that the R chart (or S chart, see below) be analyzed before the X-
bar chart, since an averaging of data is only meaningful if the within-sample variation is under control.
Using the same data and sample size as above, we get the following results:

0.06

0.05

0.04

Inside diameter
Range

0.03 Mean
Control limits
0.02

0.01

0
0 10 20
Sam ple, N = 5

R chart

WinSTAT User's Manual The Graphics Menu • 121


S Chart
The S refers to sigma (standard deviation) and is another measure of the variance within the sample. It
is more accurate than R, since R is used to approximate the standard deviation, whereas here it is
actually calculated (R became popular for doing charts by hand, since the calculation is trivial). Don't
confuse this sigma with the sigma in the dialog box, which, for an S chart, is the sigma of sigma, that
is, the standard deviation of the within-sample standard deviation.
Using the same data and sample size as above, we get the following results:

0.025

0.02
Standard deviation

0.015
Inside diameter
Mean
0.01 Control limits

0.005

0
0 10 20
Sam ple, N = 5

S chart

The similarity to the R chart results is apparent.

122 • The Graphics Menu WinSTAT User's Manual


p Chart
This chart is also known as the fraction nonconforming chart. It assumes that products are being tested
on a pass/fail basis. That is, each product either passes the test or is rejected. At intervals, a constant
number of items (the sample) are taken from the production line, and the number of rejections
(nonconforming items) are counted.
In contrast to the data in X-bar, R, and S charts, there is only one number (the number of rejections)
entered in the worksheet for each sample. The p chart first transforms the absolute numbers to
fractions of sample size and then plots the results.
In the file [Link], 30 samples of 50 cans each were tested for leaks. The number of cans failing
the test in each sample is recorded in the variable bad cans. Selecting bad cans as the variable and
setting a sample size of 50, we get the following results:

0.6

0.5
Fraction nonconforming

0.4

Bad cans
0.3 Mean
Control limits
0.2

0.1

0
0 10 20 30
Sam ple, N = 50

p chart

In the above example, it is apparent that something has caused the process to be out of control at
samples 15 and 23.

WinSTAT User's Manual The Graphics Menu • 123


np Chart
This chart, also known as the number nonconforming chart, displays the same information as a p chart,
but charts the rejections as absolute numbers instead of as fractions of sample size. The above example
now looks like this:

30

25
Number nonconforming

20

Bad cans
15 Mean
Control limits
10

0
0 10 20 30
Sam ple, N = 50

np chart

Note that the only difference is in the scale of the Y axis.

124 • The Graphics Menu WinSTAT User's Manual


c Chart
This chart is also known as the count of nonconformities or number of errors chart. Instead of a
pass/fail test as in p charts and np charts, it is assumed that each product unit can contain a certain
number of errors (nonconformities). The product as a whole may or may not still be acceptable. A
constant number of products (the sample) is examined at intervals, and the total number of errors is
recorded. There is only one number (the number of errors) entered in the worksheet per sample.
The file [Link] contains the test results of circuit boards. This is a typical example of a
product which can have minor production faults and still function satisfactorily. 26 samples of 100
board each were taken. Each entry in the variable errors is the number of nonconformities found in all
100 boards of the corresponding sample.
Selecting errors as the variable and setting the sample size to 100, we get the following results:

45

40

35

30
Count of defects

25 Errors
Mean
20
Control limits
15

10

0
0 10 20
Sam ple, N = 100

c chart

WinSTAT User's Manual The Graphics Menu • 125


u Chart
This chart, also known as the count of nonconformites per unit or number of errors per unit chart,
contains the same information as the c chart, but displays the count of errors on a per-unit basis. The
above example now looks like this:

0.45

0.4

0.35

0.3
Defects per unit

0.25 Errors
Mean
0.2
Control limits
0.15

0.1

0.05

0
0 10 20
Sam ple, N = 100

u chart

126 • The Graphics Menu WinSTAT User's Manual


Graphics/Pareto Chart
The Pareto chart is a means of graphically viewing the information on a check sheet for production defects. Organize
the worksheet as follows:
Define a variable for each possible defect, e.g. "Paint out of limits" or "Misaligned weld". Under each
variable, enter the number of times the given defect occurred. These frequencies may be spread over
any number of lines (cases) to reflect organizational information. For example, each line may
represent the sum of defects for one week. In any case, the program will sum each column in its
entirety to calculate the Pareto chart.
Here is the sample file [Link]:

Worksheet for Pareto chart

In the dialog box, specify all of the variables to be included in the chart:

Pareto chart dialog box

WinSTAT User's Manual The Graphics Menu • 127


The output is a chart of frequencies, sorted in descending order. The cumulative frequency over all
variables is also displayed.

Frequency cumulative (%)


80 100

70
80
60

50
60
40
40
30

20
20
10

0 0
Size outside Bad parts Insufficient Misaligned Paint out of
specs glue w eld limits

Pareto Chart

128 • The Graphics Menu WinSTAT User's Manual


contingency coefficient 38
Control limits 119
Correlation 61

Index
Cox Regression 97, 103
Cramer's V 38
Cronbach’s alpha 62
Cross-correlation 66
Crosstabs 37
cumulative frequency plot 17
Cumulative frequency plot 116

D
A Database 4
a posteriori 51 dendogram 88
Agglomeration 86 Descriptive statistics 21
Analysis of Variance 50 Discriminant Analysis 79
arithmetic mean 22 dummy variables 54
average linkage 86 Duncan 52
Dynamic Link between Data and
Results 11
B
balanced experiment 54 E
Bartlett test 50, 51
Basic statistics 21 eigenvalue 92
Bonferroni 52 eigenvalues 82
Box & Whisker 114 Euclidian distance 85
Box-Cox transformation 20 event 98, 104
breakdown of means 30 expected frequencies 36
expected frequency 38
extraction of factors 92
C
c chart 125 F
canonical correlation 82
canonical discriminant functions 82 Factor Analysis 91
censor 98, 104 factor loadings, graph of 94
centroid method 86 Factor Scores 95
chi-square 38, 51 Fisher's exact test 38
Chi-square 36 Frequencies 25
chi-square (Friedman test) 59 frequency 38
Chi-square test 18, 35 Friedman 58
Class Definition 8 F-value 51
Cluster Analysis 85
Cluster separation 89 G
Cochran Q-Test 59
communality 92 Graphics 113
Compare 2 Groups 41
Compare N Groups 49 H
complete linkage 86
confidence interval 30, 74, 78, 106 histogram 16
confusion table 83 Histogram 25, 113
consistency measures 62 H-test 56

WinSTAT User's Manual Index • 129


I principal axis method 92
principal components method 92
Installing WinSTAT 1 prior probabilities 79
probability plot 117
K probability scale 117
Process Capability 111
Kaplan-Meier 97
Kappa index 48
Kendall's tau 64 Q
Kolmogorov-Smirnov test 18, 35 Quality control plots 118
Kruskal-Wallis 56 quartiles 24
kurtosis 24 quartimax 92

L R
Locking Results 12 R chart 121
LSD 52 randomness 33
range 24
M rank sum 46
Regression 69
Mahalanobis distance 83 REGW 52
maximum 24 Repeated measures 56
maximum R-Square regression 77 residual analysis 71
McNemar test 47 residuals 70, 73, 76
mean 22 rotation of factors 92
mean rank 44, 45, 57, 58
Means 29
median 24 S
minimum 24 S chart 122
multiple comparisons 51 sample size 118
Multiple regression 76 Scatterplot 115
Scheffé 52
N Scott’s homogeneity quotient 62
sigma 119
normal curve 26, 116 Simple regression 69
normal distribution 15 single linkage 86
np chart 124 skewness 19, 23
S-N-K 52
O Spearman rank correlation 63
Split Plot Repeated Measures 56
observed frequencies 36 standard deviation 23
one-factor ANOVA 51 standard error of mean 22
Outliers 34 stepwise regression 76
sum 24
Survival Analysis 97
P
p chart 123 T
Pareto Chart 127
Partial correlation 65 Templates 13
Pearson correlation 62 Test of normal distribution 35
percentiles 24 Test of randomness 33
Polynomial regression 73 Tests of Normal Distribution 18
post hoc 51 transforming non-normal data 19

130 • Index WinSTAT User's Manual


T-test (dependent) 42
T-test (independent) 41
Tukey 52
two-factor ANOVA 53

U
u chart 126
U-test (Mann-Whitney) 44

V
variance 23
Variation coefficient 23
varimax 92

W
Ward's method 86
Warning limits 119
Weighted Regression 74
Wilcoxon test 45
Wilks' lambda 82

X
X-bar chart 120

WinSTAT User's Manual Index • 131

You might also like