Statistics for Quantitative Research
Statistics for Quantitative Research
Quantitative
Researchers
School of Mathematics, Physics and Computing
Study Book
ii
Published by
[Link]
Copyrighted materials reproduced herein are used under the provisions of the Copyright Act
1968 as amended, or as a result of application to the copyright owner.
iii
iv Table of Contents
Contents
1.1 What is/are Statistics? . . . . . . . . . . . . . . . . . . . . . . . . 1.4
1.1.1 Population and Sample Datasets . . . . . . . . . . . . . . . . . . . 1.5
1.2 About Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.6
1.3 Categorical Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.8
1.4 Quantitative Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.9
1.5 Working with spss – data entry . . . . . . . . . . . . . . . . . . . 1.10
1.5.1 Working with spss – raw data . . . . . . . . . . . . . . . . . . . . 1.12
1.6 Describing Distributions with Numbers . . . . . . . . . . . . . . 1.13
1.7 Comparing distributions . . . . . . . . . . . . . . . . . . . . . . . 1.16
1.8 Using Technology . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.18
1.8.1 A Note about Graphs . . . . . . . . . . . . . . . . . . . . . . . . . 1.19
1.9 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.20
1.10 Tutorial Module 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.21
1.11 Module 1 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 1.23
1.3
Module Objectives
On successful completion of this module students should be able to:
use spss to enter data, open a data set and label variables and values;
use spss to construct a bar chart and pie chart display of a categorical variable;
use spss to calculate the five number summary of a data set, including the
median, lower and upper quartiles, and interquartile range (iqr) and interpret
these measures;
use spss to calculate the mean and standard deviation of a set of data; and
recognise the impact of outliers and skewness on measures of centre and spread.
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
1.4 Module 1. Exploring and Understanding Data
Introduction
This first module of the Statistics for Quantitative Researchers Course covers material
in Chapters 1, 2, and 4 of De Veaux, Velleman & Bock, (5th edition)). It also
introduces the statistical software spss. Note that material in Chapter 3 is covered
in Module 8.
In general terms the module covers the relevance and importance of Statistics as
a discipline and describes types of data. The basic ideas and definitions in this
module are used throughout the course. The module also covers the description of
distributions of variables by showing how data can be summarised using graphs and
numbers.
This course will help you develop the skills to become proficient at interpreting and un-
derstanding data; and more importantly communicating this information and knowl-
edge to others. De Veaux, Velleman & Bock (5th edition, Chapter 1) gives some
examples of how Statistics is important in everyday life.
Throughout this Study Book there are references to readings and exercises from the
prescribed textbook (De Veaux, Velleman & Bock, (5th edition)). These provide
additional explanations of content and practice exercises.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 1 (Section 1)
After each prescribed reading throughout the Study Book, some questions are asked
related to the reading and are designed to highlight some of the important points in
the reading. Even if you don’t formally write down the answers to these questions,
at least check the answers off in your mind and read back through the material if
necessary. We suggest the points raised might be used as the basis for your own notes
about the subject. Answers to these questions are not provided but, where necessary,
can be requested by posting to the appropriate forum on the Course StudyDesk.
What is Statistics?
Estimation: Often statistics from the sample are used to estimate the pop-
ulation parameters. For example, the value of a sample mean y (is a known
number) and is an estimate of the unknown population mean µ. In real life,
most of the datasets are from samples.
1.6 Module 1. Exploring and Understanding Data
Cases – in answering the Who question when collecting data we are defining
the ‘cases’ of the data set – the individuals or objects that the data describes.
If information is being collected on people then these people are the cases; if the
information is recorded for trees then the trees are the cases. Often the cases
that we describe in a dataset are a subset (sample) from some larger group of
interest (population).
The word distribution recurs throughout the course. Every time we see the word
we should recall it as a description of all the values a variable can take and how often
those values occur.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 1 (Sections
2 & 3).
Remember that a variable takes on values which are variable; i.e., they can vary from
one case to the next. Don’t confuse the variable with its values. ‘Program of Study’
is a variable; ‘business’ and ‘science’ are two of its values. When asked to describe
a variable, make sure your definition is precise. For example, ‘ice-cream’ is not a
variable, ‘flavour of ice-cream’ is.
Exercise 1.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 1.11, 1.13.
Note: 1.11 means De Veaux, Velleman & Bock, (5th edition) Chapter 1,
Exercise 11 at the end of the chapter.
Answers to the odd-numbered exercises are in the back of the textbook. Teaching
staff are able to check answers to even-numbered exercises where necessary. Post
questions of this type to an appropriate forum on the StudyDesk.
1.8 Module 1. Exploring and Understanding Data
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2 (Section 1).
Data for a single categorical variable can be represented by a graph such as a bar
chart or pie chart.
The above count data can be represented in the following contingency table:
In this module we have described categorical data and briefly mentioned contingency
tables. Module 8 covers contingency tables and the association between two categor-
ical variables extensively, exploring marginal and conditional distributions, and the
test of independence.
1.4. Quantitative Data 1.9
Exercise 1.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.1, 2.3, 2.5,
2.35, 2.41, 2.43.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2 (Sections 2
& 3).
What is the essential difference in appearance between a bar graph and a his-
togram?
What is an outlier?
What is meant by ‘skewed to the left’ ? Give an example of a skewed to the left
distribution.
An important idea from this reading is that of a visual description of the distribution
of a quantitative variable. Shape, centre and spread are the key elements of this
description. Shape can in turn be broken down into number of modes, degree of
symmetry, and outliers or unusual features. A description based on these elements is
intended to provide sufficient information for a reader to be able to roughly sketch
the histogram of the distribution, including a scale on the horizontal axis.
With this in mind, if a visual description is required, it should be concise and stick to
the main features. However, don’t over-interpret and rely only on a visual inspection
of a histogram. In Section 1.6 more mathematically precise numerical measures of
centre and spread are considered. As regards a visual inspection though, the measures
we discuss in Section 1.6 are not relevant here.
Similar comments apply when comparing two or more distributions visually (see
De Veaux, Velleman & Bock, (5th edition), Chapter 4, Section 1). However, we
prefer to use another type of graph for this procedure as we shall see in Section 1.7
of this Study Book.
there are no gaps between the bars of a histogram (unless there is a class of
zero frequency);
the bin width (or number of bins) is chosen to most clearly reveal the overall
shape of the distribution.
Exercise 1.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.33, 2.45, 2.47.
defaults to two decimal places even for data values that are whole numbers. This can
be adjusted in the ‘Variable View’.
Now click on ‘Variable View’ and enter the information as shown in Figure 1.2. Not
all the information you need is visible on this figure. The values of the variable
‘gender’ are 1 for male and 2 for female; for ‘faculty’, the values are 1 = Business,
2 = Sciences, 3 = Other; and for ‘mode’, 1 = On campus, 2 = Online. It may help
to refer to the spss video Inputting data in SPSS on the Course StudyDesk when
entering this information.
Notice the distinction made in spss between Variable Name, Variable Label, and
Variable Values. Also notice spss should be told the type of each variable, whether it
is ‘scale’ (i.e., quantitative), ‘nominal’ (i.e., categorical) or ‘ordinal’. After completing
the ‘Variable View’, click again on the ‘Data View’ and check that it looks the same
as in Figure 1.3.
These data are part of a dataset collected in a past survey of Statistics students. The
complete dataset is made use of in the spss videos which can be found in the spss
Resources link on the Course StudyDesk.
Notice that all the data in this example are numeric, even though three of the five
variables (gender, faculty, and mode) are categorical. The values of the categorical
variables have been coded. In other words, categories such as ‘male’ have been
1.12 Module 1. Exploring and Understanding Data
given numeric labels. This is a common and strongly recommended practice. Coding
increases the speed and accuracy of data entry and can simplify the manipulation of
data once it is entered into the computer. Certainly, coding makes the data harder to
read but spss allows us to toggle between the coded and uncoded versions at the click
of the icon called ‘Value Labels’, 2nd active icon from the right in the top toolbar.
Coding is made use of extensively in the data files used in this course.
The data can be saved as an spss data file using the ‘File/Save As’ command. The
extension on the file name that spss creates is ‘sav’. Try this on the current example—
call the file ‘[Link]’ say. If spss is now closed and the [Link] name double-
clicked in Windows Explorer, spss will automatically open and load [Link], and
all the variable information that was inserted will still be there.
Exercise 1.4
Watch the following spss videos. After watching each video try to
replicate each one.
Bar Graph
Pie Graph
Contingency Tables
Weight Cases
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2, Sections 4
& 5.
After this reading you should be able to write down the answers to the following
questions:
What is a quartile?
1.14 Module 1. Exploring and Understanding Data
What does iqr stand for? What is the formula for the iqr?
What is the range of a set of numbers? Why is range not commonly used as a
measure of spread?
What is the formula for the mean? Interpret this formula in words.
If the units of measurement of the observations is metres, what are the units of
measurement of each of the mean, median, standard deviation and variance?
Which descriptions of centre and spread are best for a skewed distribution?
Which measures are best for a reasonably symmetric distribution?
This reading concerns the calculation, interpretation and usage of statistics (plural);
that is, numbers calculated from data. Commonly the types of statistics described in
this chapter are called summary statistics in that they provide a concise numerical
summary of the distribution of a quantitative variable.
P
y
mean = y = (1.1)
n
(y − y)2
P
2
variance = s = (1.2)
n−1
√
rP
(y − y)2
standard deviation = s = s2 = (1.3)
n−1
Notice we have used the notation Q1 and Q3 for lower and upper quartiles. Sometimes
these are called instead the first and third quartiles—then the median is the same as
the second quartile.
The text mostly uses y to describe the mean of a sample. There’s no reason why x
couldn’t be used, or t or u for that matter. Using y or x are just conventions and it’s
best to stick to them unless there’s a good reason not to. In Module 6, for example,
we deal with differences between pairs of numbers and it makes sense to call the mean
of the differences d. Just make sure you make it clear to the marker in any of your
assignment answers what your notation represents.
Statistics (singular) is about variation. In many ways then measures of spread such
as iqr (which measures the range of the central 50% of a set of data) and standard
deviation (which has no easy interpretation except as the square root of variance
which in turn is approximately the mean squared deviation) are more important
than measures of centre such as median and mean. Certainly it should be clear by
now that it would be quite misleading to report the centre of a distribution only,
without mention of its spread.
The word ‘average’ is in common usage of course. From a statistical point of view
however it’s preferable to be more specific and use ‘mean’ or ‘median’. Otherwise
we could be accused of being misleading knowing that the mean and median can be
considerably different because of skewness or outliers.
Exercise 1.5
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.15, 2.17, 2.21,
2.49, 2.51, 2.55, 2.77, 2.83 (use spss for calculations)
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 4, Sections 1
& 2. Note Sections 3 & 4 are optional.
A boxplot is a quick and dirty way of displaying the distribution of a quantitative vari-
able compared to the more refined method described earlier; namely the histogram.
1.7. Comparing distributions 1.17
Example 1.1
An experiment was carried out in Tasmania, Australia between 1964 and 1971
to ascertain the effectiveness of cloud-seeding to promote rain (Miller, A.J. et
al (1979) ‘Analyzing the results of a cloud-seeding experiment in Tasmania’,
Communications in Statistics—Theory & Methods A8(10),1017–1047). Some
of the data from this experiment are displayed in Figure 1.4. Compare in fewer
than 75 words words, the distributions of rainfall from the seeded and unseeded
clouds.
Solution
Both distributions are right skewed. The seeded rainfall distribution has a lower
centre but a larger spread than the unseeded rainfall distribution. The median
of the seeded rainfalls is about 3.5 cm and of the unseeded rainfalls about
4.5 cm. The interquartile range is about 6 cm for the seeded and about 4 cm
for the unseeded rainfalls. Both distributions have one or two high outliers.
1.18 Module 1. Exploring and Understanding Data
Notice that what we’re doing here in general terms is describing a relationship between
a categorical variable and a quantitative variable. In this example the categorical
variable is ‘whether or not seeding occurred’ and the quantitative variable is ‘amount
of rainfall’.
Check out also the spss video on boxplots (on the StudyDesk – ‘Graphing in spss’
then boxplot).
Exercise 1.6
Do De Veaux, Velleman & Bock, (5th edition), exercises 4.15, 4.25, 4.41
(use spss), 4.49.
Calculators and computers may give answers with 8 or 10 significant figures. That’s
fine part way through a calculation, but make note of the rule that says you should
not round numbers until the final answer is reached. Common sense dictates the
degree of rounding of the final answer. Statistics is not an exact science. We are
often estimating values based on limited or approximate information or using an
approximate model to describe the real world. As such, statistics like means and
standard deviations can only be expected to be slightly more precise than the original
data and answers based on the normal distribution are approximations to the truth.
The larger the quantity of data, the relatively more precise statistics can sensibly be.
Exact rules about this exist but are not needed in this course because a fair amount
of latitude is given. Just demonstrate commonsense and round numbers correctly.
spss is available at all times during the course. It should be used to draw graphs,
produce tables, and calculate statistics. Before getting carried away with including
computer output in assignment answers however, read the instructions given in the
assignments carefully. spss output often contains more than we need to answer a
particular question. In assignments you need to demonstrate that you can select
relevent information from the spss output to answer the question that was posed.
Determination of relevant summary statistics is quite straightforward in spss and is
discussed in the spss video Quantitative Descriptive Statistics located in the spss
Resources link on the StudyDesk. Note that a major aim of this course is to gain
familiarity with statistical software.
As regards the last item, the point is that sample size gives the reader some idea of how
reliable the results shown in the graph are. For example, if relative frequencies rather
than frequencies are displayed and the reader sees a figure such as 50%, knowing that
50% is based on 200 individuals is far more compelling than knowing it’s based on
just 10.
1.20 Module 1. Exploring and Understanding Data
Exercise 1.7
Watch the following spss videos. After watching each video try to
replicate each one.
Selecting Cases
Boxplot
Histogram
Make sure that you keep moving forward and try to keep ahead. There’s a lot of
material in this course. It pays to keep up even at the expense of things which may
still be a mystery. They will make more sense after doing later work.
The secret of success in mastering the course is problem solving. It’s not enough
just to read the material through and look at examples. Ultimately you need to be
problem solvers.
A minimum set of problems from the text have been prescribed. Do as many as you
are able in the time available, concentrating on areas which you feel less comfortable
with. Another source of practice questions is the tutorial worksheets. These have an
added benefit in that the worked solutions to these worksheets (which will be made
available on the StudyDesk) will give you guidance on how to express your answers
in assignments. There is no substitute for tackling problems!
The terms, concepts and graphs covered in this module recur throughout the course.
Again, getting familiar with them will therefore put us in good stead for later work.
1.10. Tutorial Module 1 1.21
Question 1:
The following represent the results (out of 100) for a class of 35 statistics students on
their final exams and the symbol (B or S) next to each student result indicates the
student’s program of study (B – Business Studies, S – Applied Sciences):
85(B), 38(B), 92(S), 93(S), 100(B), 88(S), 78(B), 79(S), 96(B), 81(B), 62(S), 69(S),
77(B), 78(S), 83(S), 91(S), 90(S), 85(S), 99(S), 97(S), 43(B), 100(S), 90(S), 73(B),
84(B), 85(B), 93(S), 63(B), 94(S), 94(S), 85(B), 74(B), 92(S), 86(B), 90(S)
(a) Enter the data into spss and create a histogram of results for the whole class.
(b) Referring to the graph only, describe the distribution in terms of shape, centre
and spread (and outliers).
(c) What measures of centre and spread would you use to describe this distribution?
Why?
Question 2:
Refer to the relevant data in Question 1 to complete the following questions:
(a) Using spss, create an appropriate graph to compare the distribution of class
results for Business Studies students with that of Applied Sciences students.
(b) Referring to your graph only, comment on the differences in these two distribu-
tions in relation to shape, centre, spread (and outliers).
1.22 Module 1. Exploring and Understanding Data
(c) What measures of centre and spread would you use to describe each of these
distributions? Why?
Question 3:
The statistics class was also asked about their usual mode of transport to campus.
From this class of 35 students, 12 indicated that they walked to campus, 5 rode a
bicycle, 14 came by car and 4 came on the bus.
(a) Create a bar graph for this data using spss (note that you will need to use Weight
Cases because the data is already grouped).
(c) Use another type of graph to display this data using spss.
(d) Describe what this graph tells you that is different from the previous graph.
Question 4:
For the histogram of body mass index (BMI) in Figure 1.5 answer the following
questions:
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
recognize the impact of outliers and skewness on measures of centre and spread;
Contents
2.1 Ethics and Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3
2.2 The Normal Model . . . . . . . . . . . . . . . . . . . . . . . . . . 2.6
2.3 Using the Normal Model . . . . . . . . . . . . . . . . . . . . . . . 2.10
2.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.13
2.5 Tutorial Module 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.14
2.6 Module 2 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 2.16
2.1. Ethics and Statistics 2.3
Module Objectives
On successful completion of this module students should be able to:
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
This module introduces guidelines around ethical practice in statistics. These ethical
considerations are designed to help guard against the propogation of misinformation.
Discussion around ethical use of data and/or research often focuses on important
aspects such as:
Intellectual property.
Quality of data, where the data may have been compromised by a conflict
of interest held by those collecting, analysing, reporting or commissioning the
reporting of the data and/or resulting research outcomes. For example:
– Where known problems with the collection method that could distort in-
terpretation are intentionally not reported.
– Where details of data collection methodology are not adequately described
allowing for later misuse or misinterpretation of the true nature of the data.
This might not be done to intentionally misinform; however, simple lack
of clarity and incomplete disclosure can lead to ethical issues.
Use of statistical methods appropriate to the data and questions that hope to
be answered:
– Verification of assumptions;
– Appropriate use of statistical methods if more than one is available;
– Appropriate use of visualisation (more than one type of graph) if more
than one approach is available.
2.1. Ethics and Statistics 2.5
– Sample size and its effect on statistical significance (we will cover this in
more detail in Module 10);
– Reporting significance without contextualising the meaning and impor-
tance of the effect size (e.g. the size of the difference between means or the
strength of the relationships between variables. More details in Module
10);
– Applying multiple hypothesis tests to subsets of the data and selectively
reporting only the significant results (we will cover this in more detail in
Module 10);
– Over-interpreting results e.g. inferring cause-and-effect based on an obser-
vational study, discussing results in the context of a population for which
the sample is non-representative
Even if your future work does not require you to produce statistical analyses and
results, you may need to rely on statistics produced by others. As consumers of sta-
tistical information we rarely have access to the original data to check if the analyses
and claims are true. Ethical guidelines help us identify the information that should
be available to us in order to trust the claims made.
Reports the sources and assessed adequacy of the data, accounts for all data
considered in a study, and explains the sample(s) actually used.
Clearly and fully reports the steps taken to preserve data integrity and valid
results.
In publications and reports, conveys the findings in ways that are both honest
and meaningful to the user/reader. This includes tables, models, and graphics.
2.6 Module 2. Using the Normal Model
When reporting analyses of volunteer data or other data that may not be repre-
sentative of a defined population, includes appropriate disclaimers and, if used,
appropriate weighting.
To aid peer review and replication, shares the data used in the analyses when-
ever possible/allowable and exercises due caution to protect proprietary and
confidential data, including all data that might inappropriately reveal respon-
dent identities.
Strives to promptly correct any errors discovered while producing the final re-
port or after publication. As appropriate, disseminates the correction publicly
or to others relying on the results.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 5.
What impact does adding a constant to every data value have on measures of
centre? What about measures of spread?
What impact does multiplying or dividing every data value by a constant have
on measures of centre? What about measures of spread?
2.2. The Normal Model 2.7
Which two values are needed to specify a particular normal distribution? What
notation is used?
Why do we need new notation for the mean and new notation for the standard
deviation?
What is a parameter?
What is the mean of the standard normal distribution? What is the standard
deviation of the standard normal distribution?
In the normal model, what percentage of values fall within one standard devia-
tion of the mean? Within two standard deviations of the mean? Within three
standard deviations of the mean?
There are two formulas that we will need to use quite often. The first one is called
the standardising formula because it turns the value y of a quantitative variable into
a z-score or standard score:
2.8 Module 2. Using the Normal Model
y−µ
z= (2.1)
σ
The second is called the unstandardising formula because it turns a z-score back into
a y value:
y =µ+z×σ (2.2)
Standardising and the normal distribution are two different concepts. Keep them sep-
arate in your mind. Apart from being used as a means of comparison, standardising
is also useful as a way of enabling us to make use of the normal model for describing
many variables of interest. The z-score in this situation is really just a means to an
end, not an end in itself.
By convention, everyday Roman letters are used to describe the summary measures
of actual data and Greek letters are used to describe summary measures of model
distributions such as the normal distribution. For example, the model distribution
counterparts of y, s and s2 are µ, σ and σ 2 .
spss is not designed to easily find percentages associated with the normal model. It
can be done but is hardly worth the effort when Normal Tables such as Table Z in
the back of the text are readily available. Thoroughly check through the step-by-step
examples of using the normal model to solve problems in De Veaux, Velleman & Bock
(5th edition, Chapter 5, Section 4 – Working with Normal Percentiles).
It’s always a good idea to sketch the normal distribution if a normal model
is being used to help answer a problem. By shading in an area under the curve,
remembering the total area under the curve is one (or 100%), we can estimate what
the answer should be before doing the calculation. A reasonably accurate sketch is
needed though. Figure 2.1 shows how to produce one. Assuming a mean µ = 60 and
standard deviation σ = 15, we’ve marked off points on a horizontal axis at 45, 60 and
75 representing µ − σ, µ and µ + σ. Three crosses are located as shown directly above
the three points and a convex upwards curve put through these crosses as shown with
convex downwards curves extending out each tail.
2.2. The Normal Model 2.9
Many aspects of using Statistics involve judgement. We have already been asked,
for example, to assess symmetry, recognise outliers, choose appropriate measures of
centre and spread, and round-off answers appropriately. In this section we are asked
to assess whether data can be modelled using the normal distribution. In other words,
given a set of values of a variable, does it appear reasonable to assume that if we had
access to all possible values of that variable, the distribution of all those values would
have a normal distribution. Strictly speaking we could safely say no because we might
argue that nothing in real-life has a perfectly normal distribution. But insisting on
exactness isn’t going to help us solve real-life problems. We need to be prepared to
accept that a model is an idealised or theoretical description of reality. If it comes
reasonably close to the reality it will be useful in providing approximate answers to
relevant questions. The Normal Probability Plot described in De Veaux, Velleman &
Bock (5th edition, Chapter 5, Section 5), available in spss as a P-P Plot, provides
a suitable visual means of checking the reasonableness of using a normal model in a
real situation if some data is available. spss also has a Q-Q plot which essentially
does the same job as the P-P plot. Sometimes, of course, no data is available, and
we need to appeal to other information or our’s or other people’s experience to make
a judgement.
The decision as to whether data are normally distributed or not can have important
implications as to how we proceed with further analysis.
To assess normality we would initially construct a histogram and look for features
which would suggest data are not normally distributed, before proceeding to a normal
probability plot. Those features in the histogram would include things like skewness
or lack of symmetry, outliers and absence of a central peak.
Remember, what we are assessing is normality of the population of data from which
our sample of data is selected. A histogram of the sample itself will never look
perfectly normal even when the population from which it is drawn is normal. Natural
variation will cause anomalies in the shape of the graph. When a sample is small (say
less than 40), the sample distribution may be quite misshaped but still be consistent
with a normal population. For larger samples we would of course expect a shape
closer to normality.
There is considerable subjectivity involved in assessing normality using just the ap-
pearance of a histogram. While more sophisticated methods do exist, there is no
way to know if the population is normally distributed for sure. For small samples
especially, the most we can often say is that the sample is consistent with a normal
population (but what is unsaid is that the sample is consistent with other non-normal
population distributions also!).
On a more positive note, as we shall see in Module 5 and beyond, the issue of normality
in choosing a method of analysis generally decreases in importance as sample size
increases. Often, even with samples as small as 15, only approximate normality is
2.10 Module 2. Using the Normal Model
required and beyond 30, only gross departures from normality are worth worrying
about.
This information is not of much use for very small samples. Often in practice we know
something more about the distribution of the data under consideration than just the
data itself either from experience in similar contexts or from theoretical consideration
of the mechanism that causes the data to be variable.
Two important categories of data which are known to be nearly normal are firstly,
measurements of physical stature (height, weight and so on) of many organisms (an-
imals, plants, etc), and secondly, errors in making measurements. For example, the
heights of human adult males are well approximated by a normal distribution. Also
the errors in these measurements of height will tend to be nearly normal. Both these
statements can be justified from both empirical evidence (experience based on data)
and from theoretical considerations.
Example 2.1
Assuming the height of adult humans is normally distributed with mean of
172 cm and standard deviation of 5 cm, what proportion of adults are between
166 cm and 180 cm tall?
Sketch a normal curve and locate the points 166 and 180 on the horizontal axis
(see Figure 2.2). The area shaded under the curve between these two points
equals the proportion required.
2.3. Using the Normal Model 2.11
Figure 2.2: Area under the normal curve for Example 2.1.
Construct a z axis as shown below the x (height) axis. Note the mean of
the distribution 172 corresponds to z = 0 and the points 172 − 5 = 167 and
172 + 5 = 177 correspond to z = −1 and z = 1 respectively.
Using formula (2.1), the z-scores for the endpoints of interest, 166 and 180, are
respectively (166 − 172)/5 = −1.2 and (180 − 172)/5 = 1.6.
From Table Z, the area to the left of z = −1.2 is 0.1151 and the area to the left
of z = 1.6 is 0.9452. The shaded area is therefore 0.9452 − 0.1151 = 0.8301.
Hence the proportion of adults between 166 cm and 180 cm tall is approximately
83%.
(Note: these probabilities can also be checked using spss; see the spss video
Normal Probabilities under the spss Resources link on the StudyDesk).
Example 2.2
Using the information as given in Example 2.1, what would be the height of a
doorway that requires 5% of adults to stoop in order to pass through it?
Since only the tallest 5% of adults are forced to stoop, the value (height)
required, x, is at the upper end of the distribution as shown in Figure 2.3. The
area to the left of x is 1 − 0.05 = 0.95.
2.12 Module 2. Using the Normal Model
Figure 2.3: Area under the normal curve for Example 2.2.
From Table Z, the entry closest to 0.95 in the body of the table is 0.9505. This
is the entry corresponding to z = 1.65. (Yes, z = 1.64 would be equally valid.)
This value is positive and greater than one as expected for a height at the
upper end of the distribution.
Using formula (2.2) for unstandardising, we have x = 172 + (1.65)(5) = 180.2.
An adult has to be at least 180 cm tall before being forced to stoop.
Exercise 2.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 5.1, 5.3, 5.11,
5.17, 5.25, 5.43 (use spss to construct a histogram and a P-P plot), 5.53,
5.55.
Exercise 2.2
Watch the following spss videos. After watching each video try to
replicate each one.
Normal Probabilities
P-P plot
2.4. Closing Comments 2.13
Note that a Quick Review and additional exercises for the first two modules are given
after Chapter 5 in the text book (after the exercises).
2.14 Module 2. Using the Normal Model
Now use your diagram of the Normal distribution and Table Z to answer the following
questions:
(b) In a standard Normal model, what value (s) of z cut(s) off the region described?
(b) Below what height will the shortest 10% of the male population of this country
be?
(c) What percentage of males from this country are between 175 and 185 cm tall?
2.5. Tutorial Module 2 2.15
(d) Calculate the interquartile range of heights of males from this country.
(e) Above what height are the tallest 5% of the males in this country?
(a) If boxes of this cereal are labelled as containing 700 grams, what proportion of
customers are going to be dissatisfied with their purchase?
(b) Below what weight will the lightest 5% of boxes of this cereal be?
(c) What percentage of boxes of this cereal are between 700 grams and 750 grams?
(d) Above what weight are the heaviest 10% of boxes of this cereal?
The local school gives one prize for Social Studies at the end of the school year. Two
subjects are studied in the Social Studies discipline – history and geography. George
studied history and obtained a result of 86 marks out of 100 overall for the year.
Ringo studied geography and obtained a result of 89 marks out of 100 overall for the
year.
(a) Before Speech Night some students in the class thought that Ringo was a certainty
to get the Social Studies prize.
Explain the students’ reasoning. Do you agree?
(b) The Social Studies HOD (Head of Department) studied statistics at university
and thought that she had a fair solution. She knew that the distribution of marks
in the history class had a mean of 63 with a standard deviation of 12 and the
distribution of marks in the geography class had a mean of 60 with a standard
deviation of 16.
Explain how she could use this information to arrive at a fair solution.
(c) Applying the teacher’s procedure, determine who got the Social Studies Prize.
2.16 Module 2. Using the Normal Model
(d) What is the probability that anyone would have scored better than the Social Sci-
ences prize winner (if the results on the examinations followed a Normal model)?
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
Contents
3.1 Scatterplots, Association, and Correlation . . . . . . . . . . . . 3.4
3.2 Simple Linear Regression . . . . . . . . . . . . . . . . . . . . . . . 3.6
3.3 Summary of Correlation and Regression . . . . . . . . . . . . . 3.8
3.4 Cautions and Precautions . . . . . . . . . . . . . . . . . . . . . . 3.11
3.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.13
3.6 Tutorial Module 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.14
3.7 Module 3 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 3.18
3.3
Module Objectives
On successful completion of this module students should be able to:
distinguish between the explanatory and the response variable where appropri-
ate;
list and check the conditions required for using a regression line;
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
3.4 Module 3. Exploring Relationships Between Quantitative Variables
Introduction
Modules 1 and 2 mainly concentrated on describing data for just one variable. We’re
seldom however interested in just one variable in isolation. Life is not one-dimensional.
Life is multidimensional. Many variables interact with each other in complex ways. To
help understand these interactions requires exploration of the relationships amongst
variables.
Much data is multivariate; that is the values of two or more variables are collected
for each of a number of people, countries, industries, or whatever. The discipline of
Statistics largely concerns methods of exploring, describing and measuring the type
and strength of relationships amongst variables. Module 1 dealt with methods of
graphing or calculating summary statistics of individual variables (as appropriate)
and introduced some useful ways to summarise the location, spread and shape of the
distributions of such variables. However, the reason for measuring or observing more
than one variable in a given situation is usually because a relationship or association is
suspected to exist amongst some or all of them. Our interest in this module is in using
appropriate means to detect and measure that relationship. We restrict attention to
pairs of variables not only for simplicity but also because graphs and tables don’t
easily lend themselves to dealing with more than two variables simultaneously. Later
modules in this course extend some of these concepts to more than two variables.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 6; the sec-
tions on Measuring Trend: Kendall’s Tau, Nonparametric Association:
Spearman’s Rho and Straightening Scatterplots is optional.
Sorry about the American examples. We need an Australian version of this text!
After this reading write down answers to the following questions:
3.1. Scatterplots, Association, and Correlation 3.5
In describing scatterplots, what three features should be looked at? What else
might also be looked for?
What are the generic names commonly used for the horizontal and vertical
axes?
Which variable should be graphed on the horizontal axis? Which variable should
be graphed on the vertical axis?
What happens to r if the units in which the variables are measured are changed?
Why is this?
What happens to r if the variables on the two axes are swapped around?
The full title of the correlation discussed here is Pearson’s correlation coefficient, also
called Pearson’s product moment.
Notice that the eXplanatory variable goes on the X-axis. Also that X comes before
Y in the alphabet and similarly we assume the explanatory x comes before or drives
the response y.
3.6 Module 3. Exploring Relationships Between Quantitative Variables
It is not always clear in a problem which is the explanatory and which is the response
variable. For example, in dealing with weights and heights of human beings, the
role of the two variables is unclear without more information being provided. We
know that r is the same regardless of which variable is put on the x-axis so if all
we’re interested in is determining correlation it doesn’t matter. However, in the next
section on regression we take things further and look at predicting one variable from
the other. We find then that the variable being predicted is always the y or response
variable. Consequently, if a problem involves correlation, look further at the questions
and ascertain if regression is involved. If so, information will exist to determine the
explanatory and response variables.
Exercise 3.1
Watch the following spss videos. After watching each video try to
replicate each one.
Scatterplot
Exercise 3.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 6.3, 6.7, 6.11,
6.13, 6.19, 6.43 (a) (Reminder: use spss for graphs and calculations).
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 7. (Read
through the Math Box for understanding ONLY. You are not required
to reproduce this).
What symbol is used for the intercept? What about the slope?
What is the formula for the standard deviation of the residuals? What symbol
is used for this standard deviation?
Remember the first thing to do before doing any correlation or regression calculations
is to produce a scatterplot. Correlation and simple linear regression are only relevant
if the form of the scatterplot is linear.
The word ‘simple’ is not used in the text but is used in the heading of this section. It
indicates that we are interested in one explanatory variable and how it relates to one
response variable. In a later module we will deal with multiple regression, meaning
more than one explanatory or x variable is involved.
Computers are ideal for coping with the amount of calculation required in the methods
described in this module. Our task is not to get bogged down with formulae or
calculations. Those things are best left to a computer. We need to be able to judge
the suitability of the techniques here, know when to apply them, know how to get
spss to perform them and be able to interpret the results.
Correlation Coefficient
The direction and strength of the linear relationship between two quantitative vari-
ables are measured by Pearson’s correlation coefficient. r, the sample correlation
coefficient, is a number between −1 and 1. The value of r is 1 if the associa-
tion/relationship between the two quantitative variables is a perfect positive rela-
tionship. The value of r is −1 if the association/relationship between the two quan-
titative variables is a perfect negative relationship. r = 0 implies that there is no
linear relationship between the two variables. The slope, b1 , of the regression line
and the correlation coefficient, r, share the same sign/direction (both positive or
both negative).
Coefficient of Determination, R2
The square of r, often denoted by R2 , is called the coefficient of determination which
indicates the proportion of variation in the response variable that is explained by
its relationship with the explanatory variable. Usually, it is expressed as a percent-
age, and R2 close to 100% indicates that X is a good predictor of Y , provided the
scatterplot indicates a linear relationship.
The above regression equation is found by using the ordinary least squares method
that yields the estimate of slope parameter β1 to be
sy
b1 = r ×
sx
and the estimate of the intercept parameter β0 as
b0 = y − b1 x,
The above regression equation is used to predict a value of Y , the response variable,
for a give value of X = x, the explanatory variable.
Residuals
The difference between the observed response y and predicted response yb (for a given
value of x) is called a residual. So, the residual is defined as
e = y − yb.
P
The
P 2sum of residuals is always zero, that is, e = 0. But the sum of squared residuals,
e , is not zero, and is minimised to find the least squares estimators of the slope
and intercept.
3.10 Module 3. Exploring Relationships Between Quantitative Variables
A residual plot (residual versus predicted response) is used to check the validity of
the assumptions of linear regression model.
Exercise 3.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 7.5, 7.27, 7.29,
7.33, 7.63, 7.69.
3.4. Cautions and Precautions 3.11
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 8.
What is extrapolation?
There is not much in the way of formulas or calculations in this reading but the
message is strong. This chapter of the textbook brings us back to reality after reading
about the elegance and power of regression in the previous chapter. It reveals traps for
the unwary (or should that be tools for the unscrupulous) and gives us a perspective
about Statistics as a science involving a large dose of commonsense, good judgement,
and a lot of work looking beyond just the application of some formula or blackbox of
some statistical software. Notice the importance of exploring the data and residuals
of the data using plots—just the sort of thing computers are good at producing.
The caution that the existence of an association between two variables (whether they
be quantitative variables as in this module or categorical variables to be covered in
Module 8) does not necessarily imply that changing one variables causes the other
3.12 Module 3. Exploring Relationships Between Quantitative Variables
one to change is so important it deserves mentioning again (and again). In the next
module we delve into what we might do about it. Here however we need to be clear
that if we are asked to explain an association on the basis of a lurking variable,
we should not just describe some potential lurking variable but also give a plausible
reason why it might be associated with both the explanatory variable and the response
variable. It’s not enough to postulate an association with just one of these variables
and not the other.
Example 3.1
Does driving with the car’s headlights on during the day reduce traffic ac-
cidents? Explain clearly (in fewer than 100 words) why data showing that
drivers who leave their lights on do have fewer accidents is not necessarily good
evidence that turning on lights causes fewer accidents.
Solution
Those drivers who turn on their lights during the day (in the belief that it
makes them more visible to other traffic on the road) may well have fewer
accidents than those who don’t because drivers who turn on their lights may
tend to be more cautious and it is the cautiousness rather than any increase
in visibility that results in fewer accidents. In other words, fewer accidents
and lights-on driving may be a common response to the overall caution of the
driver.
Exercise 3.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 8.15, 8.25, 8.29,
8.35, 8.45.
Exercise 3.5
Watch the following spss videos. After watching each video try to
replicate each one.
Scatterplot
Residual Plot
3.5. Closing Comments 3.13
We’ve looked at how to display the relationship between a categorical and quantitative
variable in Module 1 (see boxplots in Section 1.7). The relationship between two
categorical variables is covered in Module 8 (see contingency tables in Section 8.1).
In this module we’ve seen that scatterplots are appropriate to display the relationship
between two quantitative variables. A summary of these results is shown in Table 3.1.
Categorical Quantitative
Just a comment about using the correct language in dealing with relationships. The
word ‘correlation’ is used all over the place in real-life. Strictly it should only be
used to describe a linear relationship between two quantitative variables. The word
‘association’ is more general and can be used to describe a relationship where one
exists between any two variables, categorical or quantitative.
Note that a Quick Review and additional exercises for elements of Module 3 are given
after Chapter 9 in the text book (after the exercises).
3.14 Module 3. Exploring Relationships Between Quantitative Variables
The following spss output was obtained for a sample of 16 chocolate bars (Figure 3.2
and Figure 3.3):
Figure 3.2: Scatterplot of Energy (in calories) versus Fat Content (in grams) for
Chocolate Bars.
(a) Identify the variables, the type of each variable and the units of measure of each
variable.
(b) Describe the scatterplot from the SPSS output in Figure 3.2 in terms of form,
direction and scatter.
(d) From the spss output (Figure 3.3), what is the correlation coefficient for this
data, and explain what it means in this context.
3.6. Tutorial Module 3 3.15
Figure 3.3: SPSS output of Energy (in calories) versus Fat Content (in grams) for
Chocolate Bars.
(e) From the spss output, what is the value of the slope (b1 ) of the regression line
for this data?
(f) Give an interpretation for what the slope means in this context.
(g) From the spss output, what is the value of the intercept (b0 ) of the regression
line for this data?
(h) Give an interpretation for what the intercept means in this context.
(j) Write down the regression line for predicting Energy from Fat Content (be sure
to define the explanatory variable and the response variable in your equation).
(k) I have a chocolate bar with 40 grams of fat, what would be the predicted energy
rating of this chocolate bar? (i.e., make a prediction of ‘y’ when x = 40).
(m) Using the residual plot (Figure 3.4), comment on the suitability of using a linear
regression model for this chocolate data.
3.16 Module 3. Exploring Relationships Between Quantitative Variables
To answer this research question data was collected on 15 cars (see Figure 3.5).
Figure 3.5: SPSS data on Age of Corolla (years) and Price Advertised ($)
3.6. Tutorial Module 3 3.17
(a) Identify the variables, the type of each variable and the units of measure of each
variable.
(b) Enter the data in Figure 3.5 into spss and create a scatterplot to represent the
relationship between the two variables.
(c) Describe the scatterplot from your spss output in terms of form, direction, scatter
and outliers.
(e) Using spss, calculate the correlation coefficient for this data, and explain what
it means in this context.
(h) From your spss output, state the value of the slope (b1 ) of the regression line for
this data?
(i) Give an interpretation for what the slope means in this context.
(j) From your spss output, state the value of the intercept (b0 )?
(l) Write down the equation of the regression line for predicting advertised price
from age (be sure to define any variables you use). Use spss to draw this line
onto your scatterplot.
(m) I have a Corolla that is 10 years old, how much do you think I would advertise
this for if I were to sell it? (i.e., make a prediction of ‘y’ when x = 10).
(n) Using spss, create a residual plot of these data and use it to comment on the
suitability of using a linear regression model for this Corolla data.
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
3.18 Module 3. Exploring Relationships Between Quantitative Variables
calculate residuals, and plot them against x and against time order; and recog-
nise unusual patterns.
Module 4
Gathering Data
4.2 Module 4. Gathering Data
Contents
4.1 Understanding Randomness . . . . . . . . . . . . . . . . . . . . . 4.4
4.2 Sample Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.4
4.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.6
4.4 Using Random Numbers to Select a Sample . . . . . . . . . . . 4.8
4.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.12
4.6 Tutorial Module 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.14
4.7 Module 4 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 4.19
4.3
Module Objectives
On successful completion of this module students should be able to:
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
4.4 Module 4. Gathering Data
Introduction
We now take a step backwards. It’s all very well to describe data and relationships
between variables like we’ve been doing but where did the data come from in the first
place. Somebody, somewhere has collected it. How should it be gathered? What do
we need to be aware of before running a survey or an experiment designed to produce
data? Statistics has a lot to say about these things as we see in this module.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 10.
What symbols are used for mean, standard deviation, proportion, correlation,
intercept, and slope, when calculated from a sample?
We prefer to think of a census, not as a sample that consists of the entire population
(as per the text definition), but as an attempt to sample the entire population. If we
stick to the text definition, the census run by the Australian Bureau of Statistics every
4.6 Module 4. Gathering Data
five years would not really be a census because some people are invariably missed.
An attempt is made however to ‘catch’ everyone.
Exercise 4.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 10.17, 10.23,
10.35, 10.39, 10.41, 10.45.
4.3 Experiments
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 11.
What is an experiment?
What is a subject?
What is a treatment?
What is a factor?
What is a level?
What is a placebo?
What is matching?
There’s a lot of jargon here. For students of psychology and those in the physical sci-
ences, especially biology, the ideas and language of experimental design are especially
important. For business students, the language and ideas of sample surveys are more
important. For all of us however, a key point is to recognise that an experimental
study, provided the subjects or experimental units have been randomly assigned, of-
ten allows us to draw a conclusion about cause and effect whereas an observational
study does not.
For example, if tomato plants grown under fertiliser A produce significantly more
tomatoes than plants grown under fertiliser B, and the plants were randomly assigned
to the two fertilisers, we can legitimately conclude (everything else being equal) that
the difference in yield is due to the difference between the fertilisers. If however
one farmer is observed to use fertiliser A and another fertliser B on their respective
tomato plants, a difference in yield cannot necessarily be ascribed to a difference in
the fertilisers because lurking variables might be the cause. For example, other factors
such as different levels of care given by the farmers, or a difference in soil fertility or
climatic conditions might confound the results making it invalid to conclude that the
fertilisers caused the difference in yield (even if they did!).
4.8 Module 4. Gathering Data
Notice the use of the terms lurking variables, factors and confounding in this example.
These terms can be used in the context of any study, observational or experimental.
If we call a lurking variable a confounding variable or a confounding factor, nobody
will complain.
Incidentally, we also don’t need to be as rigid as the textbook concerning the definition
of an experiment. The text suggests an experiment involves manipulation of factor
levels, random assignment of experimental units to treatments (or, equivalently, of
treatments to experimental units) and comparison of responses across treatments. A
more generally accepted definition is simply a study in which an experimenter has in-
tervened in assigning experimental units to treatments, whether or not randomisation
has occurred.
Exercise 4.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 11.21, 11.23,
11.31, 11.47.
alphabetically ordered, but it doesn’t need to be. It just needs to be complete. The
members of this list are then numbered. The usual numbering system is such that
all members of the sampling frame have a unique number with an equal number of
digits. For example, if there are 28825 members in the sampling frame, we might
number the members consecutively from 00001 to 28825. Or we might number them
from say 10000 to 38824.
To obtain the chosen 500 we just generate 500 random numbers from the numbers
assigned to the sampling frame and these random numbers identify the 500 individuals
in the simple random sample. In principle this can be done using a table of random
numbers. In practice, with all but small random samples, a computer package such as
Excel or spss would be used making use of their built-in random number generators.
To generate a small random sample let’s see how to use the random number table
in the back of De Veaux, Velleman & Bock, (5th edition). The following example
follows the principles described above.
Example 4.1
Use the table of random digits to select a simple random sample of size five
from the following sampling frame:
Bock, D
Carmichael, C
De Veaux, R
Dunn, P
Fahey, P
McDonald, C
Moore, D
Norusis, M
Plank, A
Shi, M
Velleman, P
Khan, S
Number the members of the sampling frame from 01 (for Bock) to 12 (for
Khan). Notice that because there are 12 members in the sampling frame we
need to use two digits to identify each member. Go to the table of random digits
on page 1009 of De Veaux, Velleman & Bock, (5th edition). Pick a starting
point at random on this page (e.g. by letting something fall on the page). Pick
a forwards or reverse direction at random (e.g. by tossing a coin). Suppose we
start at the beginning of line 27 going forwards. Because our sampling frame
numbering uses two digit numbers, we mark off pairs of digits starting with 31,
08, 74, etc. The first pair 31 is ignored because no members of the sampling
4.10 Module 4. Gathering Data
frame have this number. The next pair 08 extracts Norusis, M. as the first
member of the sample. Continuing on, 74 is ignored as are a long sequence
of numbers until 05 occurs, identifying Fahey, P. as the second member of the
sample. Eventually we obtain the sequence of relevant numbers 08, 05, 05, 04,
08, 04, 04, 02, 11 giving the sample of numbers 08, 05, 04 02, 11 after ignoring
the repeats.
The required srs is therefore Norusis, Fahey, Dunn, De Veaux and Velleman.
Example 4.2
An experiment compares three treatments A, B and C. Twelve subjects are
available and we allocate four to each group at random. Suppose the subjects
are the same ones as in the sampling frame in the previous example. As before
they are numbered from 01 to 12. The idea is to take a srs of size 4 and
put them in treatment group A, then take a random sample of size 4 from
the 8 remaining subjects and put them in treatment group B, and then the
remaining 4 subjects are put in treatment group C. (This is not the only way
of randomising subjects to treatment groups—you might think of another way
of doing it. The important thing though is that each subject must have an
equal chance of being in each group.)
Suppose we start reading the random table part way along row 16 and decide
to read backwards; i.e., the digits start as 7009754052. . . . Check for yourself
that after discarding repeats, the first eight relevant random number pairs are
09, 06, 11, 10, 08, 12, 02, 04. Therefore 09, 06, 11, and 10 go into group A, 08,
12, 02, and 04 into group B and the rest into group C. Hence the subjects are
allocated to the treatment groups as follows:
Solutions to the following problems are given at the end of the module.
Exercise 4.3
The six people listed below are enrolled in a statistics course taught over
the Web. Use the list of random digits:
27102 56027 55892 33063 41842 81868 71035 09001 43367 49497 54580 81506
Exercise 4.4
Exercise 4.5
The basic principles of good data collection design have been talked about, but in-
sufficient time is available in this course for putting them into practice. This is
unfortunate because, more than any other topic in this course, learning by doing is
invaluable here. Some follow up statistics courses such as Experimental Design in-
clude project work involving data collection. Also, research students will more than
likely be exposed to contexts which will involve data collection and implementation of
the principles discussed here. Anyway, we can at least call ourselves armchair experts
on data collection after this module, if nothing else.
As far as this course is concerned, we are now in a position to turn our attention to
how data, collected according to the principles of good design, can be turned into
useful information.
Note that a Quick Review and additional exercises for elements of Module 4 are given
after Chapter 11 in the text book (after the exercises).
Answers to Exercises
Exercise 4.3
De Veaux, Velleman, Jones. (The first three distinct digits between 1 and 6, reading
from the left are 2, 1, 5 representing these three people.)
Exercise 4.4
The sampling frame (i.e. list of voters on the electoral roll) could be numbered from
00001 to 86314. A sequence of random digits might then be obtained or generated.
For example, many calculators will generate random numbers (e.g. 0.348, 0.019,
0.272, 0.951, 0.877, . . . ) which can then be grouped into lots of five (e.g. 34801,
92729, 51877, . . . ). These are then used to pick out 1500 members of the sampling
frame to include in the srs ignoring any five digit numbers greater than 86314 or
4.5. Closing Comments 4.13
any that have previously occurred (e.g. 34801, 51877, . . . ). This is very laborious in
practice—fortunately there are computer packages that automate this process.
Exercise 4.5
Assuming the subjects are numbered consecutively from 01 to 20, we obtain the
following sequence of relevant random numbers 03, 07, 12, 15, 04, 06, 11, 09, 16, 17,
20, 05, 10, 01, 02. Assuming the subjects are firstly assigned to group 1, then to
group 2, group 3 and finally group 4, we obtain the following allocation:
(a) Using a random number generator, select a simple random sample (SRS) of eight
people from the sampling frame defined above.
(b) How many business students (b) and science students (s) are there in your sample?
(c) Using a random number generator, select a second simple random sample.
(d) How many business students and science students are there in this sample?
(e) Compare your two samples: Did you expect the result you got? Explain your
answer with reference to the sampling process.
(f) Using a random number generator, select a stratified random sample, stratified
according to discipline studied so as to contain equal numbers of business and
science students.
(g) Describe, in detail, the steps you took to produce your stratified sample.
(h) Why might you decide that a stratified sample is more representative?
4.6. Tutorial Module 4 4.15
(Note: [Link] is
one possible random number generator or the function RANDBETWEEN() in Excel).
Firstly, we need to determine the type of study being reported to be able to appreciate
its strengths and weaknesses. Does the study include the features of a good design?
Answer the following questions about the given reports, giving reasons for your com-
ments where appropriate. Due to copyright restrictions actual newspaper articles
reporting on real research studies were not able to be used for this tutorial; the
following activities refer to fictitious studies reported in a newspaper style.
Article 1
62 per cent of runners had a protein drink and/or protein bar after
a race
18 per cent of runners just drank water or soft drink
10 per cent of runners consumed fruit, such as a banana, orange or
watermelon
8 per cent of runners enjoyed a beer, wine or vodka cruisers post-race
2 per cent of runners had nothing
Institute President Steve Coe said they were also asked about how soon af-
terwards they returned to training or racing. Mr Coe said a staggering 95
per cent of runners who consumed protein drinks or bars after races were
back on the track or road within 48 hours. ‘The vast majority reported
rapid recovery and were ready to run again very quickly,’ he said. ‘While
some of it might be mind over matter, there is without doubt proof in the
pudding that an early drink or meal of protein aids muscle recovery and
4.16 Module 4. Gathering Data
The poll, perhaps not surprisingly, revealed that those who drank alcohol
or consumed nothing post-race were likely to not run again in the following
week. ‘Many reported feeling lethargic, tired, even hung over. They said
they did not usually train or race again for at least 3-4 days,’ Mr Coe said.
The Institute had long advocated consuming protein-rich food or drink in
the 30 to 45-minute window immediately after hard exercise. ‘We don’t
specify any particular product, but anything is better than a beer post-
race ... well maybe not for some,’ he said.
Runners’ Worldly, August, 2015
Article 2
Research indicated that 398 SMEs were still trading today, while 402 were
no longer registered. CICG Director Jedidiah Jones said of those 402, al-
most 300 had closed the door before January 2015, while the remaining
100 shut it down in the following 6-9 months. ‘It is tough being in busi-
ness, particularly starting out,’ Mr Jones said. ‘Many start-ups open with
insufficient capital or backing. They hit the doldrums after three months
and simply fail to hang on.’ Mr Jones said, that, conversely, the 50 per
cent or so who make it beyond the first year and are still trading are seeing
a brighter light on the horizon. ‘They have become largely self-sufficient.
The bottom line is in the black, not the red. And they have got through
the tough times and are ready to grow.’
CICG provided support to start-ups to help them get going and hopefully
make it to that all-important second year. But there was little support
from state or federal governments. ‘Politicians need to go out on a limb
and help our struggling start-ups and entrepreneurs. These are the people
4.6. Tutorial Module 4 4.17
with the next great ideas, the Google, Facebook and Twitter creators of
this world,’ Mr Jones said. ‘However, they need financial help to get
through the first year. A 50 per cent fail rate is not good for a modern,
progressive country like Australia.’
The Monthly, CICG, January, 2016
Article 3
Flu shot:
Control:
Question 3: For each of these three articles, answer the questions related
to that type of article.
(a) Identify:
(c) Are there any lurking variables that might explain the observed association(s)?
(e) Draw a diagram that represents the experimental design used in the study.
4.7. Module 4 Checklist 4.19
(a) Identify:
(c) Are there any lurking variables that might explain the observed association(s)?
A study to test the effects of marijuana recruited Australian adults (aged 18-35 years)
who have used marijuana. 20 adults were randomly assigned to smoke marijuana
cigarettes, while 25 adults were given placebo cigarettes. What are the flaws in this
study?
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
recognise the existence and potential effect of lurking variables on both obser-
vational and experimental studies;
use random numbers from a table or computer software to select a simple ran-
dom sample (SRS) and a stratified random sample;
Contents
5.1 Sampling Distribution of a Sample Mean . . . . . . . . . . . . . 5.6
5.1.1 Modelling the Distribution of a Sample Mean . . . . . . . . . . . . 5.6
5.1.2 The Central Limit Theorem (CLT) . . . . . . . . . . . . . . . . . . 5.7
5.1.3 How Large is Large? . . . . . . . . . . . . . . . . . . . . . . . . . . 5.8
5.1.4 z-Score Formulas . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.10
5.1.5 Reporting Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . 5.11
5.1.6 Jargon and Other Things . . . . . . . . . . . . . . . . . . . . . . . 5.13
5.1.7 In summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.15
5.2 Statistical Inference for a Mean . . . . . . . . . . . . . . . . . . . 5.16
5.2.1 Estimating with Confidence for One Mean . . . . . . . . . . . . . . 5.16
5.2.2 Hypothesis Testing for One Mean . . . . . . . . . . . . . . . . . . . 5.19
5.2.3 Sample size determination . . . . . . . . . . . . . . . . . . . . . . . 5.26
5.2.4 More about Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.27
5.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.29
5.4 Tutorial Module 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.30
5.5 Module 5 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 5.32
5.3
Module Objectives
On successful completion of this module students should be able to:
recognise that a statistic will take different values when sampling is repeated;
recognise that statistics based on large samples are less variable than statistics
based on small samples;
explain when a t-model rather than the standard normal model is appropriate;
state the assumptions for and check the appropriateness of a one sample t
procedure;
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
So far the course has focused on describing and producing data with a bit about
probability thrown in. The rest of the course concentrates on drawing conclusions
from data about the world at large in a range of different scenarios.
Collectively this is called Statistical Inference. Estimation (in the form of a confi-
dence interval)—estimating a population parameter using a sample statistic is one
aspect of statistical inference. Hypothesis testing—making decisions about popula-
tions (including the values of population parameters) based on sample information is
the other.
Example 5.1
One role of a university student counsellor is to advise students on techniques
to improve their study skills. Good sleep is believed to be part of any good
study routine and caffeine is thought to be a possible issue in maintaining good
sleep patterns.
Our interest is in ascertaining, on the basis of information obtained from 151
randomly-chosen students, whether there is an association between the amount
of coffee consumed and the number of hours of sleep students manage to have
per night. What we are doing then is asking a question about a population (all
students of the university) and trying to answer it on the basis of a represen-
tative sample (our 151 randomly-chosen students) from that population.
This is the idea of hypothesis testing. That description suggests we are testing the
plausibility of some conjecture and that’s essentially what hypothesis testing amounts
5.5
to. We take some statement (a hypothesis) about a population and test the consis-
tency of data with that statement. We don’t go as far as suggesting we’re trying
to establish the absolute truth or otherwise of some statement about a population.
We are seldom in a position on the basis of the limited information available in our
sample to be able to do that. What we hope for is that we can give some measure of
how consistent our data is with the hypothesis and draw a conclusion from that.
Let’s look at how that works with this student data. The hypothesis we are testing
might be stated as ‘There is an association between the amount of coffee consumed
and the number of hours of sleep students manage to have per night’. This is called
the ‘research’ or ‘alternative hypothesis’ (Ha ). The starting point of the test
always is to be sceptical; that is, we assume initially that the alternative hypothesis
does not hold. In other words we assume that ‘There is no association between the
amount of coffee consumed and the number of hours of sleep students manage to have
per night’. This is called the ‘null hypothesis’ (H0 ). The null hypothesis then is a
statement of the absence of whatever it is that the alternative hypothesis is stating.
The word ‘null’ is used to indicate this absence.
The test proceeds by assuming the null hypothesis to be true and testing how
consistent the data is with this assumption. If we decide that the data is inconsis-
tent with the null hypothesis then we would be inclined to believe the alternative
hypothesis rather than the null hypothesis described the population correctly.
How do we decide on how consistent the data is with the null hypothesis?
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 17 (Sec-
tion 1).
What shape has the sampling distribution of the sample mean if the population
is normal?
What shape has the sampling distribution of the sample mean in general?
What is the standard deviation of the sampling distribution of the sample mean?
Is the variability of a sample mean greater than, equal to or less than that of
an individual observation?
Try not to lose sight of the basic concept in this section. The idea of averaging,
whether it’s to work out a mean as a way of estimating something about the world
at large, is best done with larger rather than smaller samples because the larger the
sample the closer we expect the mean to be to the true value. That’s not particularly
brilliant! It’s really just commonsense (or an application of the Law of Large Numbers
if you want to impress!).
This section provides us with a formula that allows us to apply this basic idea to real-
world problems. Notice the presence of the square root of n in the standard error
formula for a mean. This is no accident. It’s what the law of diminishing returns
is getting at—one of the golden rules of statistics is that as n gets bigger things
only improve by the square root of n. For example, how much better off would
we
√ be with a sample that is ten times as large. The answer is that we would be only
10 = 3.2 times as well off. Or putting this another way. If we wanted to be ten
times as well of, we would need 102 = 100 times as much data. Sad, but true—that’s
the nature of the world.
How do we know how well off we are? Well, that’s the idea of the standard error. The
smaller the standard error the better off we are. Recall the standard deviation is like
a ruler measuring how far we are away from the true mean (as described in Module 1
and Chapter 5 of De Veaux, Velleman & Bock, (5th edition)). When we calculate
a mean from some data then we’d like to know how far away the value we’ve got is
from the true mean.
A standard error is just a standard deviation, but in this case it’s the standard
deviation of a mean. Actually, to be more exact, the term standard error is used to
describe our best estimate of the standard deviation of any statistic, whether it’s a
mean, proportion, correlation, intercept, slope, or even another standard deviation!
It follows that we wouldn’t expect a mean to be more than about three standard
errors away from the true mean according to the 68-95-99.7 rule provided the mean
is approximately normally distributed.
The CLT allows us to use the normal model to answer probability questions about
the sample mean even though we don’t know the population distribution.
Example 5.2
The ages of individuals in a srs of size 50 are recorded. Suppose the mean
and standard deviation of age of the population from which the sample was
drawn are 44.5 years and 13.2 years respectively. What is the probability that
the mean age in years of our sample is
Solution
By the clt, the sampling distribution of the sample mean Y is approximately
normal with mean 44.5 years and standard deviation 1.8668 years (the standard
error of Y ). This distribution is displayed in Figure 5.1.
(a) The required probability is the area to the right of y = 50. Converting to
a z-score we have z = (50 − 44.5)/1.8668 = 2.95. From Table Z this gives
an area below y = 50 of 0.9984. Hence the probability that the mean age
in years of our sample is greater than 50 is 1 − 0.9984 = 0.0016 ≈ 0.2%.
(b) ‘Within one year of the population mean’ means that Y is between
43.5 years and 45.5 years. The respective z scores for these ages are
−0.54 and 0.54, for which Table Z gives 0.2946 and 0.7054. Hence the
probability that the mean age of our sample is within one year of the
population mean is 0.7054 − 0.2946 = 0.4108 ≈ 41.1%.
Exercise 5.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 17.27, 17.51.
Statistics texts talk about ‘large samples’ and ‘small samples’. By ‘large samples’
what is usually meant is, large enough so that we can make use of the clt! Conversely,
by small samples is meant, not large enough to safely use the clt! Nobody can tell
us exactly when n changes from being small to large. Many factors are involved not
the least of which is how approximate is approximate. Nonetheless a few guidelines
exist based on experience and theory.
For srs’s of size 25 or more, there is seldom any problem in using the clt to
argue that the sampling distribution of the mean is approximately normal.
We should always plot the data (histogram) from which the mean is being
calculated and make some judgement about it depending on the amount of
skewness and the presence or otherwise of outliers. If the population appears
likely to have just a single peak and unlikely to be severely skewed based on the
sample plot, and the sample contains no outliers, applying the clt to a srs of
5.10 Module 5. Statistical Inference for One Mean
size 15 or more is likely to give reasonable results. Plots of samples of size less
than 15 can suffer a lot from sampling error and are therefore not very reliable.
If no other evidence is available about the ‘true’ distribution apart from such a
plot, it will be necessary to make it clear up front that normality is assumed,
or make use of a method that does not rely on normality (see Section 6.3).
Outliers also can make a big difference to an analysis if the sample is small and
it pays to check out just how big that difference is by doing the analysis with
and without the outlier(s).
A z-score is just a measure of how many standard deviations the value of interest is
above the mean. In terms of a formula when dealing with sample means we have
y−µ
z= (5.1)
√σ
n
The similarity between (5.1) and (5.2) often causes confusion so is clarified here.
Notice that if we put n = 1 in (5.1), the formula reduces to (5.2) apart from the bar
over the y. Obviously working out the mean of a sample of size one is pretty boring
because there’s only one number involved, but the mean of one number is really just
that same number, right? So we could drop the bar over the y in this case and formula
(5.1) would be exactly the same as (5.2). All that is saying is that (5.1) works for
any size sample n including n = 1.
How does this help us? Well, the formula (5.1) is the general formula and (5.2) is
a special case of it. Thinking in these terms, problems such as 17.51 in De Veaux,
Velleman & Bock, (5th edition) should be less confusing. In part (a) of this problem
we are effectively asked about one pregnancy (n = 1). OK, it looks like we’re asked
5.1. Sampling Distribution of a Sample Mean 5.11
about all pregnancies but think of this question as asking what is the probability
of one pregnancy chosen at random being between 270 and 280 days and we get
the same answer as the original problem. Then in parts (c) and (d) the questions
involve n = 60 pregnancies. We can therefore use formula (5.1) in parts (a), (c) or
(d) provided we put in the right value of n! While we’re at it, part (b) is found using
the unstandardising formula
y =µ+z×σ (5.3)
which we’ve seen before as (5.2) rearranged. Better still, by rearranging (5.1) we get
σ
y =µ+z× √ (5.4)
n
Usually in practice, we don’t know what the mean of the model is so we don’t know
what the error is. But we can still figure out the standard error! That’s very nice.
So, for the pregnancy example, even though we may not know that the model or true
mean is 266 days, or that the true standard deviation is 16 days, we can still calculate
the standard error from the data we have. Say, for example, the n = 10 pregnancies
had durations (in days) as follows:
269, 279, 245, 257, 259, 246, 270, 261, 279, 275
Then the sample mean is y = 264.0 days and the sample standard deviation is s =
12.47 days. The standard error (of the mean) is then
s 12.47
SE(Y ) = √ = √ = 3.94 days
n 10
and this tells us a lot about how far apart the sample mean y = 264.0 days and
(unknown) true mean are likely to be.
5.12 Module 5. Statistical Inference for One Mean
Figure 5.2: Mean rainfalls (with standard errors) for data in Example 1.1.
It’s common practice in reporting statistics to include the SE and n where appropriate.
For example, ‘The mean gestation period was 264.0 (±3.9) days (n = 10)’.
Included with this method of reporting should be a clear statement explaining that
the bracketed value is the standard error. A footnote at the first occurrence is often
used (for example, ‘statistics are reported (±SE)’, or an indication as appropriate in
the column headers of a table containing summary statistics).
Failing to give some idea of the precision of important summary statistics can be quite
misleading. If for example a report stated that 54.36% of respondents were against
the proposed change, the reader would be forgiven for believing that a majority of
the target population were against the change. The statement 54.4% (±7.5%) puts a
different perspective on matters. It suggests that a repeat of the study may well find
a minority of respondents against the change since 54.4 − 7.5% is below 50%.
It’s smart also to include standard errors in graphs representing statistics. An ex-
ample is shown in Figure 5.2. See the spss video Error Bars on the StudyDesk for
instructions on how to do this.
5.1. Sampling Distribution of a Sample Mean 5.13
Statistics are random variables. Parameters are not. Every sample of Australian
workers can provide a sample mean income, and these means will vary from one
sample to the next. The mean income of all Australian workers however is not a
variable but, at least at some point in time, is a fixed number that would remain
unknown unless we were to take a census at that time.
A statistic such as a sample mean has got a standard deviation. We know that the
√
standard deviation of the sample mean Y is SD(Y ) = σ/ n where n is the sample
size. This formula is not much use to us though because we need to have values for
the parameter σ to use it, and we seldom know what this value is. What we do in
practice is replace the σ with the sample standard deviation s. Then we talk about
√
the standard error of the sample mean Y as SE(Y ) = s/ n. In general terms
1
Remember that samples and statistics both start with ‘s’; parameters and populations both start
with ‘p’.
5.14 Module 5. Statistical Inference for One Mean
then, a standard error is an estimate of the standard deviation of a statistic. You will
notice as you keep using spss that standard errors get reported for many statistics,
not just for means. Just remember they are estimates of the standard deviation of
the distribution of the statistic.
Usually, upper-case letters such as X and Y are used to refer to random variables
and lower-case letters such as x and y to refer to observed values of random variables.
When we talk about a random variable or a statistic in a random or sampling situation
we should use the upper-case version of the letter, e.g. Y . The lower-case version, y,
refers to values that are actually observed in a random or sampling situation. For
example: ‘The value of y from the experiment was 12.4’. Similarly, on the graph of the
distribution of X we use the upper-case X when talking about the distribution itself
and the lower-case x to describe particular values of the variable on the horizontal
axis. In practice, in this course, you will not be penalised for inappropriate use of
lower and upper-case letters.
Write down answers to the following questions to check your understanding of this
material:
What is a population?
What is a sample?
Write down the symbols used to describe a sample mean and a population mean.
5.1.7 In summary
A number of concepts of importance to the rest of the course are discussed in this
section of the module.
Firstly, the sample we collect or the sample that we are presented with should be
thought of as only one of many possible samples that could have been obtained of
the same size from the population. Hence the sample mean or standard deviation
or whatever statistic we calculate from the data is really just a value of a variable.
Knowing about the distribution of the variable is the basis of the idea of generalising
from the data at hand to the world at large. We introduce this in more detail in the
next section.
As sample size n increases, the variability of these statistics decreases. This variability
is measured by the standard error. The standard error gets smaller in proportion to
one over the square root of n. This means that the rate at which the SE gets smaller
slows down as n gets larger and so a law of diminishing returns applies.
Because the normal model applies to statistics of data (at least when the sample is not
too small), the mean and standard deviation are appropriate measures of centre and
spread for sampling distributions. Hence, even though the median and iqr may better
describe the centre and spread than the mean and standard deviation of a population
which is skewed, when our main interest is in sampling with a view towards drawing
conclusions, the mean and standard deviation are more important parameters than
the median or iqr. For this reason in the rest of this course we focus on the mean and
standard deviation rather than median and iqr as measures of centre and spread.
The ideas in this section are used in one way or another in the rest of the course. If
you are still struggling with them, don’t worry and don’t spend too much time on this
material. The relevance of this material will become clearer once you have actually
applied these concepts in conducting a range of inferential methods.
5.16 Module 5. Statistical Inference for One Mean
Specifically, in this module, we look at a scenario involving one mean. From the Cen-
tral Limit Theorem, we know that the sampling distribution of the mean is approxi-
mately normal with mean µ and standard deviation √σn regardless of the distribution
of the population, if n is large. However, if the population standard deviation is not
known (which is generally the case), the next best thing is to use the sample standard
deviation s instead of σ. In doing so, the sampling distribution of the mean becomes
a t distribution provided the underlying population has an approximate normal dis-
tribution. This ‘nearly normal’ condition in relation to the population distribution is
important when the sample size is small, but not so critical for larger samples.
De Veaux, Velleman & Bock, (5th edition) covers the material discussed in this section
in parts of Chapters 17, 18 and 19. In addition to presenting inference for a mean
these chapters include aspects of inference for a proportion. This course does not
specifically cover inference for proportion although many of the concepts in these
chapters are common to both mean and proportion.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 17 (excluding
Section 4).
5.2. Statistical Inference for a Mean 5.17
What is the difference between the Student’s t-model and the normal model?
What is the formula for the standardised sample mean when σ is unknown?
What is the formula for the one-sample t-interval for the mean?
What is the formula for the degrees of freedom of the Student’s t-model in this
context?
The idea of a confidence interval (CI) is a natural extension of the idea of a standard
error. In Section 5.1.5 we suggested (actually, more like insisted) that we include
standard errors when describing sample means. Provided the sample size is large
enough so that the normal model applies, a straightforward application of the 68-95-
99.7 rule tells us that y ± SE is really a 68% confidence interval, y ± 2 × SE is a 95%
confidence interval and y ± 3 × SE is a 99.7% confidence interval for µ.
We don’t usually talk about 68% or 99.7% CI’s though, although 95% CIs are popular.
The common confidence levels are given at the bottom of Table T. When using formula
(5.5), the degrees of freedom used to find the critical value t∗ in Table T is given by
n − 1. The degrees of freedom in other situations when t is used are given by other
formulas (you will encounter this in later modules). Performing a one-sample t-test
and constructing a confidence interval for a mean using the t distribution are basic
5.18 Module 5. Statistical Inference for One Mean
skills in Statistics. You should be able to perform the analysis both by hand and, if
the data is available, by using spss.
Notice that t and z are very similar—z is used if σ is known; t is used if σ is not
known. If s is replaced by σ in the above formula then the t becomes a z and Table Z
rather than Table T is used for the calculations (however, in practice we rarely know
σ). Don’t be put off by the ∗ on the t and z. It just indicates that these values are
‘important’ values that can be found in Table T. If you leave them off the formulas,
nobody will complain.
Example 5.3
Beanies (a fruit-flavoured gummy-textured sweet in the shape of a broad bean
seed) are sold in 180 gram bags. A random sample of 24 bags of Beanies were
collected and contents were weighed by the members of a tutorial group of
students, yielding the following data:
185, 182, 177, 180, 181, 182, 171, 182, 176, 182, 182, 176,
185, 175, 175, 176, 175, 186, 170, 182, 172, 185, 176, 176
Find a 95% confidence interval for the weight of Beanies in bags labeled as
containing 180 grams.
Solution:
From the data, the sample mean (y) is calculated to be 178.71 grams and the
sample standard deviation (s) 4.686 grams.
Using Formula 5.5, with df = n − 1 = 24 − 1 = 23 and thus t∗ = 2.069 from
Table T,
s 4.686
y ± t∗n−1 √ = 178.71 ± 2.069 × √
n 24
= 178.71 ± 1.979
With 95% confidence we can say that the population mean weight of 180 gram
bags of Beanies is between 176.73 grams and 180.69 grams.
Checking the assumptions are satisfied for formulating this confidence interval:
(a) Randomisation: we are told that students took a random sample of bags
of Beanies.
(b) Nearly Normal : we have a sample size of 24 which is reasonably large so
the ‘nearly normal’ condition is not as critical, however we can check the
reasonableness of this assumption by plotting the data in a histogram to
check for outliers or skewness. Figure 5.3 shows no outliers and minimal,
if any, skewness.
5.2. Statistical Inference for a Mean 5.19
Having said all this, note that in the assignments you only need to discuss
assumptions if asked explicitly to do so, however, in practice, it is vital to check
the validity of using the procedure by checking that the assumptions/conditions
have been satisfied.
Exercise 5.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 17.11, 17.13,
17.21, 17.29, 17.39, 17.55 (a)-(e), 17.59, 17.61.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapters 18 and 19
(note that these chapters also refer to hypothesis testing for a proportion
which is not required for this course; however it is worthwhile reading all
sections of these chapters as the broad concepts discussed refer to both
means and proportions).
What distribution does the test of one mean make use of?
What is the form of the null hypothesis in the one mean t-test?
What is the form of the alternative hypothesis in the one mean t-test?
What is the test statistic formula for the one-sample-t-test for the mean?
What is a P -value?
When using formulas (5.5) and (5.6), the degrees of freedom used to find t∗ in Table T
is given by n − 1. The degrees of freedom in other situations when t is used are given
by other formulas (you may encounter this in other courses). Performing a one-sample
t-test and constructing a confidence interval for a mean using the t distribution are
basic skills in Statistics. You should be able to perform the analysis both by hand
and, if the data is available, by using spss.
On some occasions it’s not necessary that a definitive decision be made. Then it’s
sufficient to report the P -value itself with some covering comment.
As a guide as to how to interpret the P -value when making a conclusion, there are
two situations depending on whether or not a level of significance is given with the
problem. The level of significance is a pre-assigned threshold for deciding whether to
reject a null hypothesis.
In practice, if a decision one way or the other is required, we set a criterion on the size
of the P -value before doing the study so that a definitive conclusion can be reached
at the end. This size is called the significance level and is denoted by α. If a level
of significance is given then it’s a matter of simply comparing the P -value with α.
Common levels of significance would set α as 0.001 (0.1%), 0.01 (1%), or 0.05 (5%),
depending on the context. If the P -value < α then we conclude that H0 is rejected
in favour of Ha at the α level of significance. If, for example, it was decided that
a significance level of 0.10 or 10% was appropriate before a test was run, a P -value
of 0.06 (6%) would allow us to reject a null hypothesis in favour of the alternative.
However, if the level of significance had been set at 5%, a P -value of 0.06 (6%) would
5.22 Module 5. Statistical Inference for One Mean
not allow us to reject a null hypothesis in favour of the alternative. It is essential that
the level of significance be set before the data is collected if a definitive conclusion
is required! If a level of significance is given then it’s a matter of simply comparing
the P -value with α. Of course, we then put our conclusion into the context of the
problem.
On some occasions it’s not necessary that a definitive decision be made. Then it’s
sufficient to report the P -value itself with some covering comment. In this case, write
down the conclusion in words as follows according to the size of the P -value:
P -value > 0.10 (or 10%): state ‘there is insuffcient evidence to support Ha ’.
P -value between 0.05 (5%) and 0.10 (10%): state ‘there is slight evidence to
support Ha ’.
P -value between 0.01 (1%) and 0.05 (5%): state ‘there is moderate evidence to
support Ha ’.
P -value between 0.001 (0.1%) and 0.01 (1%): state ‘there is strong evidence to
support Ha ’.
P -value < 0.001 (0.1%): state ‘there is very strong evidence to support Ha ’.
Of course, these conclusions should be made in the context of the problem. In other
words, Ha should be replaced by a contextual statement of the alternative hypothesis
in non-technical language.
1. State the null and alternative hypotheses (Check all relevant assumptions of the
test have been met);
2. Assume the null hypothesis is true and calculate the value of the test statistic;
Example 5.4
Using the data in Example 5.3, students decide to test the hypothesis that,
on average, the manufacturers of Beanies are underfilling the 180 gram bags,
using a 5% level of significance (note that the level of significance should be
decided before the data is collected).
Solution:
Using the four steps to hypothesis testing:
5.2. Statistical Inference for a Mean 5.23
Step 1: State the null and alternative hypotheses (Check all relevant
assumptions of the test have been met).
H0 : µ = 180
Ha : µ < 180
where µ is the population mean weight of Beanies.
The conditions/assumptions for this question are the same as those for the
confidence interval and have been shown to be satisfied in Example 5.3.
Step 2: Assume the null hypothesis is true and calculate the value of the
test statistic (using Formula 5.6).
y−µ s
t= where SE(Y ) = √
SE(Y ) n
178.71 − 180 4.686
= where SE(Y ) = √
0.9565 24
= −1.349
Using SPSS
Input the data in one column. Figure 5.4 shows a snapshot of the first 7 data
values for the sample of n = 24 bags of Beanies.
5.24 Module 5. Statistical Inference for One Mean
Output from the One-Sample T Test. From Figure 5.7, the t-test statistic is
−1.350 which is slightly different to the −1.349 we obtained when doing the
hypothesis test by hand. This is due to round off error. We used values of the
mean and standard deviation rounded to 2 and 3 decimal places, respectively,
to calculate the t-test statistic, whereas spss uses many more decimal places
than this. The df, as expected, is the same. The ‘Sig.(2-tailed)’ is the P -value
for a two-tailed test. We have performed a one-tailed test (only interested in
underfilling), thus the P -value is 0.095, half of this Sig.(2-tailed) value of 0.190.
This is consistent with what we found before with 0.095 being between 0.05
and 0.10. To obtain the 95% confidence interval from this output we add the
null hypothesised value to each bound thus 180 − 3.27 gives a Lower Bound
of 176.73 and 180 + 0.69 gives an Upper Bound of 180.69 (in agreement with
what we obtained in Example 5.3).
Exercise 5.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 18.11, 18.45,
18.47, 18.49, 18.53, 18.55.
Exercise 5.4
Watch the following spss video.
We would expect to have t∗ not z ∗ in this formula, but we don’t know n (we are
trying to find n) to be able to find the degrees of freedom and thus t∗ . However,
when the sample size n is large, t∗ and z ∗ will be about the same value. In this course
we will only use the sample size formula (5.7) when the answer we expect for n is
large, because then we can replace t by z.
The answer we get from this formula should be rounded up to the next whole number.
Either s (as shown) or σ can be used in this formula.
Example 5.5
Refer back to information given in Example 5.3. The students now wish to
estimate the population mean weight of 180 gram bags of Beanies to within
a margin of error of 1.5 grams with 95% confidence. What is the minimum
sample size required? (Use the sample standard deviation from the previous
sample to apply Formula 5.7).
5.2. Statistical Inference for a Mean 5.27
Solution:
The minimum required sample size is
z ∗2 s2
n=
ME2
1.962 × 5.1462
=
1.52
= 45.21
Rounding up 45.21 to the next integer number, we get the required sample size
to be 46.
Note: A quick way to find z ∗ values is to use the bottom row of Table T. The common
confidence levels are given at the bottom of Table T; 95% is there with a critical value
for z in the row immediately above it, namely 1.960 (which is pretty close to the 2
indicated by the 68-95-99.7 rule). We use the values z = 1.960 (95%), 1.645 (90%),
and 2.576 (99%) so often that they are worth remembering to save time in looking
them up.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 19. The sec-
tion about P -values (Chapter 19, Section 1) is optional since it involves
notation (conditional probabilities) which was deliberately omitted from
the course. However, reading this section might give you a better under-
standing of P -values.
Apart from the size of the effect being estimated, what else affects the P -value
in a test of significance?
How might the practical significance rather than the statistical significance of
an effect be judged?
De Veaux, Velleman & Bock, (5th edition) put emphasis on the null hypothesis in
setting up a hypothesis test and making conclusions from it. We prefer to give the
emphasis to the alternative hypothesis, often describing it as the research hypothe-
sis to indicate that it represents the hunch or conjecture that a researcher believes
represents reality. As such the alternative is decided on first. The null hypothesis
then represents an absence of the effect or condition described in the alternative hy-
pothesis. The reason for the text emphasis on the null hypothesis is that the test
procedure focuses heavily on the null hypothesis. The null hypothesis is assumed to
be true when doing a hypothesis test. The aim is to find whether or not the data is
consistent with the null hypothesis. We also accept conclusions which focus on the
alternative hypothesis and whether or not there is sufficient evidence provided by the
sample to support it.
Exercise 5.5
Do De Veaux, Velleman & Bock, (5th edition), exercises 19.3, 19.11 (c),
19.12 (a), 19.13, 19.15, 19.27, 19.39.
Note that a Quick Review and additional exercises for elements of Module 5 are given
after Chapter 19 in the text book (after the exercises).
5.30 Module 5. Statistical Inference for One Mean
It is known from student records that exam scores aggregated across the entire uni-
versity follow an approximately normal distribution with mean of 70 marks and a
standard deviation of 8 marks.
(a) Use this information to estimate the percentage of exam scores that are 78 or
higher.
(b) One of the statistics tutorial groups of 35 students achieves a mean exam score
of over 78 marks. The tutor claims outstanding performance but the lecturer
believes this result is not unusual. Assuming that the 35 students in the tutorial
group represent a random sample of university students, calculate the probability
that the mean exam score of this random sample of 35 exams is greater than 78
marks.
I measured the weight of 10 Mars Bars and found the weights to be (all measured in
grams):
61.5, 62, 59, 60, 61, 63, 57, 62, 62, 60.5
Using this data calculate a 95% confidence interval for the population mean weight
of Mars Bars.
(d) What is the critical value for this level of confidence (also give the degrees of
freedom)?
(h) Give a meaningful statement of the estimate that you have calculated.
Using the data in Question 2, test if the mean weight of the Mars Bars is different
from the 60 grams stated by the manufacturers, at the 1% level of significance.
(d) Assuming the null hypothesis is true, calculate the test statistic.
(j) If the population standard deviation is 2 grams, find the minimum sample size
required in estimating the known population mean with 95% confidence within
0.25 margin of error.
(b) Using spss, calculate a 95% confidence interval for the population mean weight
of Mars Bars.
(c) Using spss, test if the mean weight of the Mars Bars is different from the 60
grams stated by the manufacturers, at the 1% level of significance.
Question 5: Errors
Discuss the meaning of Type I and Type II error in the context of the following
hypotheses:
H0 : A person is not guilty of a murder
H1 : A person is guilty of a murder
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
recognise that a statistic will take different values when sampling is repeated;
interpret the meaning of a standard error in general and for a sample mean;
carry out ,by hand and using spss, a hypothesis test about one mean using a t
distribution;
Contents
6.1 Inference for Paired Samples . . . . . . . . . . . . . . . . . . . . 6.4
6.2 Independent Groups . . . . . . . . . . . . . . . . . . . . . . . . . . 6.11
6.3 Parametric and Nonparametric Tests . . . . . . . . . . . . . . . 6.18
6.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.20
6.5 Tutorial Module 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.21
6.6 Module 6 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 6.23
6.3
Module Objectives
On successful completion of this module students should be able to:
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
This module is really just a continuation of the previous one. Now that we’ve seen
the idea of generalising from data to the world at large by using a confidence interval
or doing a hypothesis test, we are ready to apply these ideas in many situations.
Specifically in this module we look at cases involving two groups measured on the
same variable.
De Veaux, Velleman & Bock, (5th edition) covers this material in parts of Chapters 20
and 21.
6.4 Module 6. Statistical Inference for Comparing Means
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 21.
What symbol is used to denote – the mean difference in the data; the mean
difference in the population?
There is essentially very little new work in this reading. The idea of blocking and of
matched pairs was briefly discussed in Module 4.
From the analysis point of view, once we can perform a one-sample t-test or produce
a one-sample t confidence interval we can analyse matched pairs data. All we must do
as the first step is calculate the directed differences between the pairs of observations.
‘Directed differences’ means we subtract one member of each pair from the other
member in the same direction for all pairs.)
There are really only the two formulas we need to be able to use and they are just
rewrites of (5.5) and (5.6) as follows:
To find a confidence interval for the mean difference use the formula
sd
d ± t∗n−1 √ (6.1)
n
To perform a hypothesis test for the mean difference use the test statistic
d
t= √ (6.2)
sd / n
6.1. Inference for Paired Samples 6.5
Comparing (6.2) with the formula given in Chapter 21, Section 2 of De Veaux, Velle-
man & Bock, (5th edition) (The paired t-test), we see that we have put ∆0 = 0. This
will be the case for all the examples and problems we deal with in this course.
Some problems may arise even though the method is not really new. Firstly we need
to be able to recognise a paired samples experiment. Secondly, we need to take care in
stating the hypothesis or confidence interval of interest on the basis of the information
given.
The attribute used for matching may be age (in which case each pair of experimental
units is of a similar age), weight, height, ability, status, condition, and so on. Often
the same experimental units, whether they are human beings, animals, plants or
inanimate objects, are used in both samples and are therefore obviously matched.
Typically such experiments are of the before-after kind consisting of measurements
taken before and after the application of or exposure to some condition. For example,
the heart rates of subjects measured before and after performing a physical task, the
weights of calves measured before and after being put on a new diet, or the state of
paint coatings measured before and after exposure to a harsh environmental condition.
The expectation is that, as a result of this matching, the measurements in each pair
will be more alike than they would be if the experimental units in both samples were
chosen completely independently of each other. We would therefore expect the two
samples to be positively correlated or at least have some positive association. A
scatterplot of the measurements in sample 1 against the measurements in sample 2
should reveal such an association. Because of this, matched samples are sometimes
called correlated or dependent samples.
6.6 Module 6. Statistical Inference for Comparing Means
Using SPSS
Example 6.1
We hear that listening to Mozart improves student’s performance on tests.
Perhaps pleasant odours have a similar effect. To test this idea, 21 subjects
worked a paper-and-pencil maze while wearing a mask. The mask was either
unscented or carried a floral scent. The response variable is their average time
on three trials. Each subject worked the maze with both masks, in random
order. The randomisation is important because subjects tend to improve their
times as they work a maze repeatedly. Table 6.1 gives the subjects’ average
times with both masks.
Table 6.1: Floral scents and learning data for Example 6.1.
Unscented Scented
Subject (seconds) (seconds) Difference
1 30.60 37.97 −7.37
2 48.43 51.57 −3.14
3 60.77 56.67 4.10
4 36.07 40.47 −4.40
5 68.47 49.00 19.47
6 32.43 43.23 −10.80
7 43.70 44.57 −0.87
8 37.10 28.40 8.70
9 31.17 28.23 2.94
10 51.23 68.47 −17.24
11 65.40 51.10 14.30
12 58.93 83.50 −24.57
13 54.47 38.30 16.17
14 43.53 51.37 −7.84
15 37.93 29.33 8.60
16 43.50 54.27 −10.77
17 87.70 62.73 24.97
18 53.53 58.00 −4.47
19 64.30 52.40 11.90
20 47.37 53.63 −6.26
21 53.67 47.00 6.67
6.1. Inference for Paired Samples 6.7
Figure 6.1: Data for Example 6.1 (i) as entered into spss and (ii) after computing
differences.
is confirmed by the spss output in Figure 6.5. The two-sided P -value is 0.730.
Because Ha is one-sided we need to halve the two-sided P -value to give the one-
sided P -value of 0.365 or 37%. We conclude that there is insufficient evidence
to support the claim that floral scents improve performance, i.e., the sample
mean difference in times of 0.957 sec is very likely due to sampling variation.
The mean improvement in performance is small, only about 1 second over the
50 or so seconds that subjects took wearing the unscented mask. This small
improvement is not statistically significant at even a 25% level of significance.
Doing this test by hand we use the test statistic (6.2). This gives
d
t= √
sd / n
0.957
= √
12.548/ 21
= 0.35
with df = 21. Table T shows that 0.35 is less than the critical t value with one-
sided probability of 0.10. The P -value is therefore greater than 0.10, leading
to the same conclusion as given above.
6.1. Inference for Paired Samples 6.9
Figure 6.4: Dialog box for the Paired Samples procedure in spss.
Exercise 6.1
Watch the following spss video.
Exercise 6.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 21.23 (a) & (b)
(use spss), 21.29 (use spss), 21.31 (by hand and using spss).
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 20 (Sec-
tions 4 & 5).
What type of plot is recommended to display the data from the two independent
samples?
What assumptions are needed to use Student’s t-models with two independent
samples?
Write down the formula for the two-independent-samples t-test for the difference
in means.
How many degrees of freedom are used in Student’s t-models with two indepen-
dent samples following the conservative approach?
6.12 Module 6. Statistical Inference for Comparing Means
Finding a confidence interval for the difference µ1 −µ2 between two means using
the formula s
s21 s2
(y 1 − y 2 ) ± t∗df + 2 (6.3)
n1 n2
with df given by one less than the smaller of n1 and n2 , that is, smaller of n1 − 1
and n2 − 1.
Performing a hypothesis test for the difference between two means using the
test statistic
y −y
t = q 12 2 2 (6.4)
s1 s2
n1 + n2
with df given by one less than the smaller of n1 and n2 , that is, smaller of n1 − 1
and n2 − 1. (This formula for the df is what the book calls the ‘easy rule’.)
Now that we are dealing with two samples, which we have denoted by the subscripts
1 and 2, we should make sure we state which sample is which in a problem. So, for
example, we might define µ1 as the mean height of males and µ2 as the mean height
of females (or, better still, use µm for males and µf for females). Also, if we are asked
for a confidence interval for the difference in means, we must state whether it’s for
µ1 − µ2 or for µ2 − µ1 , using (6.3) appropriately of course.
The randomness assumption usually relies on us being told or knowing that each
sample was either a srs (or could be treated as such) if the study was observational,
or involved randomisation if the study was an experiment. In practice we would
rely on information concerning the data collection process, remembering that simple
random sampling in an observational study and randomisation in an experiment are
descriptions of the protocol of the study, not the data itself.
Similarly, independence of two samples must often be taken on faith. More about this
in Section 6.3 though, where studies in which the two samples are not independent
are considered. The 10% condition is seldom an issue.
6.2. Independent Groups 6.13
Checking normality requires judgement, and opinions may vary. The discussion on
assumptions and conditions in Section 4 (Chapter 20) of De Veaux, Velleman & Bock,
(5th edition), should be noted. The basic gist is that for t procedures, the importance
of normality diminishes as sample size increases and for n larger than 40 (often 25
is enough as suggested in Section 5.1.3), the assumption can be pretty well ignored.
Although looking at boxplots and normal probability plots (P-P plots) is helpful,
with small samples it is usually not possible to decide from the data alone whether
or not normality is acceptable. We must rely on previous knowledge concerning the
distribution of the variable under consideration and should include a statement to
this effect in the conclusion, or, if considerable doubt exists about the validity of the
normality assumption, seek alternative analyses which do not rely on normality. Such
methods are called nonparametric procedures (see Section 6.3 for more details).
Having said all this, it’s worth repeating that in the assignments you only need to
discuss assumptions if asked explicitly to do so. For example, if a two-independent-
samples t-interval is required and (6.3) is appropriately applied, no marks will be
lost for not mentioning or checking assumptions unless specifically requested. In the
practice of Statistics however, whether it be in the workplace, in other
university studies or elsewhere, it is unwise and potentially dangerous to
make use of any statistical procedure without checking the assumptions.
Using SPSS
Here spss is applied to an exercise from the text. The data for this exercise can be
found by looking at 20.66 (Chapter 20, Question 66).
Example 6.2
Read Question 66 in Chapter 20 of De Veaux, Velleman & Bock, (5th edition).
The data as provided is in two columns, part of which is shown in the left-hand
panel of Figure 6.6. The first step is to reformat the Data View so all the skull
measurements are in just one column. A simple cut and paste achieves that.
Now a second column is inserted with a coding to represent the date of each
of the skull measurements. In the right-hand panel of Figure 6.6 part of the
reformatted data is shown with 1 representing 4000 B.C.E and 2 representing
200 B.C.E. Notice how the data is now in standard format with each of the
60 rows representing an individual skull and the two columns representing the
variables ‘breadth’ and ‘date’.
where the easy rule for df has been used giving df = 29 and t∗29 = 2.045
from Table T. This is a little wider than what spss produces because the
easy rule for df is conservative.
6.2. Independent Groups 6.15
Figure 6.6: spss data from 22.66 before and after reformatting.
which is the same as that given by spss. We can estimate the P -value
from Table T as follows. We have df = 29 and a two-sided test. From
Table T we see that 3.579 is larger than the largest entry (2.756) for 29 df
and therefore, because 2.756 corresponds to a (two-sided) P -value of 0.01,
3.58 corresponds to a (two-sided) P -value less than 0.01. Our conclusion
is the same as from the spss output.
spss needs the data to perform a procedure. This means that problems in which only
summary statistics (means, standard deviations, sample size, etc) are given cannot
be done using spss and must be done by hand using a calculator.
Figure 6.8: Dialog box for the ‘Independent-Samples T Test’ procedure in spss.
Exercise 6.3
Watch the following spss videos. After watching each video try to
replicate each one.
Exercise 6.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 21.9, 20.57,
20.59, 20.65 (use spss), 20.67 (by hand and using spss), 20.71, 20.73.
Sometimes the best thing to do would be to discard the data altogether and do no
analysis at all! This would be the case if we don’t believe the data is representative
of the population or model we’re wanting to describe. It might be that a survey has
been run with voluntary respondents for example—like those ones TV channels are
keen on promoting. Unless there’s no reason to believe a relationship exists between
the variable(s) of interest and whether or not a person is likely to respond to such
a survey, the data from these types of surveys are of no value (except of course
financially to the TV and telephone companies!).
For example, it might be possible to argue that the height of participants is unlikely
to be associated with whether or not a person takes part, so these heights might be
representative of a wider audience. But height is hardly going to be the focus of
attention in such a survey! Interest is likely to be on attitudes, opinions, perceptions,
popularity, etc., and these are exactly the variables that will be associated with the
likelihood of a person responding. Typically, for example, people who feel strongly
about an issue respond, whereas those who are indifferent do not. Also, of course,
there’s an undercoverage issue in that those who are most likely to be exposed to the
survey (i.e., watching that TV channel at that time) may not be representative of the
population of interest. The upshot is that the statistics reported from these surveys
are essentially worthless as a gauge on public opinion or the like.
6.3. Parametric and Nonparametric Tests 6.19
The point is that if an assumption like normality for a t procedure is looking a bit
dodgy, other procedures are available. There are lots of them. We still assume
a representative or random sample from the population of interest and the data
points are independent of those for any other in the population. However, we are
not restricted by the assumption that the values in the population have a normal
distribution.
As a rule, t procedures which rely on normality are the most powerful available for the
types of questions they answer. For quantitative data there is a strong compulsion
to use procedures relying on normality. Sometimes, as a result they get used when
they shouldn’t, when normality fails or when the data is not even quantitative.
How do we know when the normality assumption fails? We usually don’t have access
to information about the whole population, so we can only use our sample data to
check the normality assumption, by plotting the data. t procedures are fairly robust
when sample sizes are large, i.e., when the sample size is large, the sampling distri-
bution of mean differences will be approximately normal, even if the distribution of
differences from the sample is not quite normal. Provided the distribution of the data
is not too skewed nor contains outliers, the t procedure may be appropriate. However,
small sample size may increase vulnerability to violations of the normality assump-
tion because the sampling distribution of the mean differences may not necessarily
be normal in these cases. In summary, with quantitative data where a t procedure
would usually be preferred, we take the conservative/cautious approach that, if n is
small, it is safer to use a nonparametric test.
The word ‘nonparametric’ is used to describe inferential procedures which have weak
assumptions, none as strong as normality. These methods don’t usually involve means
and standard deviations but rely on counts or ranks, and so can work on data that is
not quantitative. Hence ordinal or categorical data is dealt with using nonparametric
methods.
6.20 Module 6. Statistical Inference for Comparing Means
Procedures which assume normality (or some other assumption about the exact shape
of a distribution) are called parametric. The t procedures are classified as being
parametric.
Obviously in the time available for this course we can only look at some of the many
techniques. Although practical guidelines are given for using t procedures in the
presence of outliers and non-normality, considerable judgement is needed in applying
them. It is possible that two analysts may disagree on the appropriateness of a t
procedure for a given set of data.
Do you get lower fuel consumption (i.e. fewer litres/100 km) using premium unleaded
petrol?
Many drivers of cars that can run on regular unleaded petrol actually buy premium
in the belief that they will use less fuel. To test that belief, 10 cars were tested in
which all the cars regularly run on regular unleaded petrol. Each car is filled first with
either regular or premium unleaded petrol, decided by a coin toss, and the number of
kilometres for that tank-full recorded. Then the fuel consumption in litres/100km is
recorded again for the same car for a tank-full of the other kind of unleaded petrol.
The drivers do not know about the experiment. The results are as follows (litres/100
km):
Car 1 2 3 4 5 6 7 8 9 10
Regular 14.7 11.8 11.2 10.7 10.2 10.7 8.7 9.4 8.7 8.4
Premium 12.4 10.7 9.8 9.8 9.4 9.4 9.1 9.1 8.4 7.4
Difference (R-P)
(a) Perform a hypothesis test to see if cars get on average lower fuel consumption
with premium unleaded petrol.
First calculate the differences between the two samples (define the way the
difference was obtained);
Second calculate the mean of the differences;
Third calculate the standard deviation of the differences;
Fourthly calculate the standard error of the mean of the differences;
Fifthly give the degrees of freedom, df;
Finally calculate the test statistic.
(b) Find a 95% confidence interval for the population mean of differences in fuel
consumption.
(c) Enter the data into spss and repeat the hypothesis test. Compare the test statis-
tic, degrees of freedom and P -value with those obtained in Part (a).
Independent random samples of Mars and Snickers Bars were weighed (in grams) to
test if the population mean weights of the two bars differ.
1 2 3 4 5 6 7 8 9 10
Mars 61 62 63 60 59 61 62 63 64 58
Snickers 57 58 59 61 60 59 62 59 58 57
(b) Find a 95% confidence interval for the difference of the two population means.
(c) Enter the data into spss and repeat the hypothesis test. Compare the test statis-
tic, degrees of freedom and P -value with those obtained in Part (a).
Question 3
“Absence of evidence is not the same as evidence of absence”. Discuss this statement
in relation to the concept of hypothesis testing.
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
6.6. Module 6 Checklist 6.23
denote the mean and standard deviation of the differences in a paired t proce-
dure;
perform a hypothesis test for the mean difference of paired data (by hand and
using spss);
Contents
7.1 One-way Analysis of Variance (ANOVA) . . . . . . . . . . . . . 7.3
7.2 One-way ANOVA and spss . . . . . . . . . . . . . . . . . . . . . . 7.9
7.3 Two-way Factorial ANOVA . . . . . . . . . . . . . . . . . . . . . 7.15
7.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.22
7.5 Tutorial Module 7 . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.24
7.6 Module 7 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 7.28
7.1. One-way Analysis of Variance (ANOVA) 7.3
Module Objectives
On successful completion of this module students should be able to:
read an ANOVA table, including identifying df, MSS and the F -statistic;
interpret the output for one-way and a two-way ANOVA and subsequent pair-
wise t-tests for significant effects.
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
In this module we look at situations where we want to compare the means of three
or more groups.
We cover Chapters 25 and 26 of De Veaux, Velleman & Bock, (5th edition) in this
module.
If we are comparing group means why is the test called Analysis of Variance?
Have a look at Figures 25.2 and 25.3 in the De Veaux, Velleman & Bock, (5th edi-
tion). Each boxplot represents a different group (sometimes called ‘treatments’). The
ANOVA compares the variances between the group means to the variances within the
groups.
In Figure 25.2 the over-lapping spread of the boxplots tells us that the variances
within the groups are larger than the variance among the group means. It would be
very difficult for us to tell whether any differences among the group means (31, 36,
38 and 31 respectively) are due to the groups representing different populations, or
whether the means vary just because of natural sampling variation.
In Figure 25.3 however, the variance within the groups is much smaller than the
variance among (or between) the group means. This could mean that the differences
among the means are more likely due to them representing different populations
rather than differences due to sampling.
In Figure 25.3 knowing the mean of each group would help us distinguish the obser-
vations in one group from the observations in another group.
In Figure 25.2 however, the mean of each group does not help us distinguish the
observations in one group from the observations in another group.
ANOVA is a significance test of a null hypothesis that states that all population
means are equal. This means the null hypothesis asserts that the groups represent
samples from the same overall population (or that we do not have enough evidence in
our samples to suggest otherwise). The alternative hypothesis is that the population
means are not all equal.
This alternative hypothesis does not necessarily mean that all the means are signifi-
cantly different to each other. It means only that at least one of the means is different,
indicating that at least one of the groups is sampled from a different underlying pop-
ulation. This also means that if we reject the ANOVA null hypothesis we will need
to perform additional analysis using t-tests, or similar pairwise tests, to determine
which groups are different to each other.
If we have k groups, then each group has a corresponding population of subjects, and
the means of the response variable for the k groups can be denoted by µ1 , µ2 , ..., µk .
Since we are trying to compare means of several groups, the null hypothesis, as always,
assumes that there is no difference between the groups we are investigating, that is:
H0 : µ1 = µ2 = ... = µk .
7.1. One-way Analysis of Variance (ANOVA) 7.5
Note that no matter what the specific context of the data we are analysing, we use
an ANOVA when we want to investigate and compare the group means and we use
the variances within and between the groups to determine whether we should accept
or reject the null hypothesis.
In order to statistically compare multiple means in the ANOVA we need to use a new
sampling distribution model, the F distribution model. The associated F -test allows
us to compare the variation (differences) between the means of the groups with the
variation within the groups. When the differences or variation between the means are
large compared with the variation within the groups, we reject the null hypothesis
and conclude that at least one mean is different to the others. Please note that in this
course, you will not be required to look up the F -tables. Instead, you will interpret
spss output.
It can be helpful to understand a bit more about how we calculate the within and
between variance components. When our groups have similar variance (σ 2 ) measures
we can pool all the within group variance estimates together to get an overall esti-
mate of the σ 2 . Because it is a pooled variance, it is denoted s2p . This quantity is
traditionally called the Error Mean Square, or the Within Mean Square and
is denoted by MSE . The error mean square is an estimate of the error or residual
variance and represents the amount of random variation within treatments/groups.
Remember though, that we also have a separate estimate of σ 2 from the variation
between the means of the treatments/groups we are investigating. This quantity is
called the Treatment Mean Square, or the Between Treatment Mean Square
and is denoted by MST .
If the null hypothesis is true, then the group means are equal and both MST and
MSE estimate σ 2 . Their ratio then, should be close to 1.0. If the null hypothesis is
false, then the MST will be larger because the group means are not equal. The MSE
is a pooled estimate in which the variation within each group is found around its
own group mean, so differences between means won’t inflate MSE . This makes the
ratio MST /MSE perfect for testing the null hypothesis. When the null hypothesis is
true, the ratio should be near 1. If the group means really ARE different, then the
numerator MST will tend to be larger than the denominator MSE and the ratio will
tend to be bigger than 1.
7.6 Module 7. Analysis of Variance (ANOVA)
The ratio of the treatment mean square MST to the error mean square MSE is
called the variance ratio or the F -statistic or the F -ratio. Under H0 , given all
assumptions are true, this variance ratio has an F-distribution.
Even when the null hypothesis is true, the variance ratio will rarely equal 1 exactly
due to natural sampling variability. So how can we tell when the ratio is big enough to
reject the null hypothesis? The answer lies with the F -distribution – the distribution
of MST /MSE (the F -statistic), and by comparing this statistic calculated from our
data, with the appropriate F -distribution we can obtain a P -value.
The resulting F -test is one-tailed because any differences in the means make the
F -statistic larger. Larger differences in the treatment effects lead to the means being
more variable, making MST bigger. This makes F increase, so the test is significant
if the F -statistic is large enough. In practice, larger F -statistic values have smaller
P -values, because there is a smaller probability of getting a large F -statistic if H0 is
true.
Independence Assumption: The groups must be independent of each other. The data
within each treatment must be independent as well, i.e. each case must be independent
of all others.
Randomisation Condition: Ensure the data was collected randomly within each group
or that appropriate randomisation was applied when allocating cases to treatments/
groups.
7.1. One-way Analysis of Variance (ANOVA) 7.7
Equal Variance Assumption: The ANOVA requires that the variances of the treat-
ments/groups be equal (or equal enough), i.e. homogeneity of variance.
Normal Population Assumption: Check Normality with a histogram. Check for out-
liers in the boxplots of the values for each group.
SST M ST
Between k−1 SST M ST =
k−1 M SE
SSE
Within N −k SSE M SE =
N −k
Total N −1 TSS
as long as we know the number of groups k and the total sample size N we
can calculate all df components. The Total df should also equal the sum of the
Between and Within df.
Example 7.1
An experiment to determine the effect of different post-surgery care plans
on the swelling in postsurgical patients recorded the hand volume changes for
7.8 Module 7. Analysis of Variance (ANOVA)
patients who had been randomly assigned to one of the care plans (treatments).
The ANOVA for the data is as follows:
We are not given a lot of information here about the experiment, but there is
quite a bit we can figure out from the ANOVA table just by knowing how each
part is calculated.
From the ANOVA table we can see that there must have been a total sample
size of 59 because N − 1 = 58 and there must have been 3 treatments because
k − 1 = 2 (but we don’t know what the treatments were). We can also see that
the TSS (Total Sum of Squares) of 3420.54 is the sum of the SST (Treatment
Sum of Squares) and SSE (Error Sum of Squares), i.e. 716.17 + 2704.38 =
3240.54. The ratio of the Mean Square values (358.08/48.29) equals the F -
statistic of 7.41.
What does the ANOVA say about the results of the experiment? Specifically,
what does it say about the null hypothesis?
The F -statistic of 7.41 has a P -value that is quite small (0.0014) and much
less than α = 0.05. We can reject the null hypothesis that the mean change in
hand volume is the same for all three treatments. In order to determine which
treatments were different to each other we would need to perform pairwise t-
tests or similar. [Note, you may see P -value, p-value or Sig. used in the text
book, spss output and these notes – they are all referring to the same thing].
If the ANOVA indicates the treatment effect is significant and we will need to
use t-tests to determine which treatments are different to each other, why do
we bother with the ANOVA? Why not just do the t-tests?
The ANOVA is a more statistically ‘powerful’ test. This means that it is less
likely to find a significant result by mistake. In any test we perform, we accept
that there is a certain probability (α) that we are rejecting H0 by mistake
(remember Type I error).
Using a single test (ANOVA) rather than multiple tests allows us to control
the probability of a Type I error, at least at the initial stages of our analysis.
With a significance level of α = 0.05 in the F -test, the probability of incorrectly
rejecting a true H0 is 0.05. When we do a separate t-test for each pair of means,
7.2. One-way ANOVA and spss 7.9
a Type I error probability applies for each comparison and for three treatments
that would need three t-tests to compare all pairs of treatments.
In the example above with 3 treatment groups there would be 3 pairwise com-
parison tests that we would need to perform (treatment A v B, A v C and B
v C). In that case, we are not controlling the overall Type I error rate for all
the comparisons.
If our ANOVA is not significant we should not proceed with applying t-tests or
similar. If we performed the t-tests without first checking the ANOVA there is
a chance that at least one of the pairwise t-tests would result in a Type I error.
Performing the ANOVA first protects us from these potential mistakes.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 25.
Following this reading, keep in mind these caveats in relation to this course:
you will not be required to look up the F -tables; instead you will be required
to interpret spss output.
you will not be required to calculate Sums of Squares by hand, however you
will need to understand how to calculate the df, MSS and F-statistic in the
ANOVA table based on other information provided (i.e. by understanding the
relationships between these values as shown above).
3. Transfer the dependent variable, percent damage, into the Dependent List:
box and the independent variable, Harvest system, into the Factor: box
using the appropriate right arrow buttons (or drag-and-drop the variables
into the boxes), as shown below:
7.2. One-way ANOVA and spss 7.11
4. Click on the Post hoc button. If the ANOVA is significant and we reject H0
then we will need to perform post-hoc pairwise tests to determine which
treatments are different to each other. There are a range of different tests
we can choose from. We will choose Tukey’s test as it includes a ‘correction
factor’ to account for the multiple tests it performs using the same group
data. Tick the Tukey checkbox as shown below:
SPSS Output
The Descriptives table (see below) provides some very useful descriptive
statistics, including the mean, standard deviation and 95% confidence inter-
vals for the dependent variable (percent damage) for each separate group (nil,
CS1, CS2 and new), as well as when all groups are combined (Total). These
values are useful when you need to describe your data.
The ANOVA table shows the output of the ANOVA analysis and whether
there is a statistically significant difference between our group means. The
hypotheses for this test would be:
H0 : µnil = µCS1 = µCS2 = µnew .
Ha : at least two of the population means are unequal.
where µ is the mean percentage of trees within two metres of the cut down tree
which have some form of damage.
The F -statistic is quite large at 27.920. Remember that if H0 is true and there
is no statistical difference between the group means then we would expect
the F -statistic to be close to 1 (because F is the ratio of the between group
and within group MSS and if they are the same then F =1) We can see that
the significance value (P -value) is 0.000 (i.e. p < 0.001), which is below 0.05
and, therefore, there is a statistically significant difference in the mean percent
damage to neighbouring trees resulting from the different harvesting systems.
This is great to know, but we do not know which of the specific groups differed.
7.2. One-way ANOVA and spss 7.13
We can find this out in the Multiple Comparisons table which contains the
results of the Tukey post hoc tests.
The table below, Multiple Comparisons, shows which groups differed from
each other. The Tukey post hoc test is generally the preferred test for conduct-
ing post hoc tests from a one-way ANOVA, but there are many others. The
Sig. column gives us the P -value.
The no guidance (nil) group compared to current system 1 (CS1), current
system 2 (CS2) and new system (new) has P -values of 0.002, 0.00, and 0.00
respectively. As each of these P -values is < 0.05 we would reject H0 for all
three tests and conclude that the mean percent damage using the nil treatment
is significantly different to all other treatments. There is also a significant
difference (i.e. p < 0.05) in percent damage between the CS1 treatment and
the new treatment where p = 0.003. However, CS2 is not significantly different
to CS1 or the new system (p > 0.05 in both comparisons).
By looking at the Mean Difference column we can also see that the significant
difference between new and nil (no guidance) of −56.02 is based on the cal-
culation of mean percent damage in the new group minus the mean percent
damage in the nil group. As this is a negative difference, the percent damage
in the new group must have been less than in the nil group.
Our conclusion, therefore, is not only that there was a significant difference
between the new and nil groups, but that there was significantly less damage
in the new system group. The difference between the new and CS1 groups was
also negative, but less so (difference = −27.44). Therefore, the new system
did perform significantly better than CS1 (significantly less damage using the
new system), however, there was not as much improvement as found when
compared to the nil group.
7.14 Module 7. Analysis of Variance (ANOVA)
When reporting the ANOVA results we need to provide the Between and
Within group df, the F-statistic and the P -value. When reporting the re-
sults of the post-hoc Tukey tests we should provide the P -values and also some
summary information describing each group – usually the mean and standard
deviation (you can find these in the Descriptives Table in the spss output
above).
Based on the results above, you could report the results of the study as follows:
There was a statistically significant difference between groups as determined by
one-way ANOVA (F (3, 16) = 37.92, p < 0.001). A Tukey post hoc test revealed
that the percent damage was significantly lower for the new system (30.01 ±
11.37 percent) compared to the no guidance system (86.03 ± 9.78 percent)
and the current system 1 (57.45 ± 10.04 percent). There was no statistically
significant difference (p = 0.095) between the new system and current system
2 (45.93 ± 8.62 percent).
3. Test statistic:
Between group variability Treatment Mean Square M ST
F = = =
Within group variability Error Mean Square M SE
F sampling distribution has df = number of groups −1 = k − 1 and N − k (total
sample size - number of groups).
4. P -value: Right tail probability of the above observed F -value.
5. Conclusion: Interpret in context. If decision is needed, reject H0 if P -value ≤
significance level (such as 0.05).
Exercise 7.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 25.6, 25.11,
25.13 - 25.22. Use spss where asked to perform an ANOVA.
7.3. Two-way Factorial ANOVA 7.15
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 26.
It is important to make sure we understand the terminology we are using as there are
often multiple words used to describe one part of the analysis.
Dependent variable or response variable – this is the thing that is being measured on
all cases regardless of treatment/group. It should be a quantitative or scale variable
so that we can sensibly calculate a mean, variance, etc. You should only have one
dependent variable in both One-way ANOVA and Factorial ANOVA. In Example 7.1
the response variable was ‘hand volume’. In Example 7.2 the response variable was
‘percentage of neighbouring trees damaged’.
In a One-way ANOVA the groups in the factor are often called treatments or
levels. In Example 7.1 there were 3 groups/levels/treatments in the post surgical
treatment factor. In Example 7.2 there were 4 groups/levels/treatments in the
harvest system factor.
7.16 Module 7. Analysis of Variance (ANOVA)
For the rest of this module we will refer to the groups in a single factor as ‘levels’ and
the combinations of levels as ‘treatments’. Based on the example in the table below,
Factor 1 has 3 levels (A, B, C), Factor 2 has 4 levels (D, E, F, G). Each case belongs
to one level of Factor 1 and one level of Factor 2 so there are a total of 12 possible
factorial treatments (see table below). When using an ANOVA we are interested in
knowing if the variation in our response variable can be explained by our factor(s)
and/or the interaction between them.
Example 7.3
A controlled laboratory experiment has been carried out to look at the effects of
four different chemicals on the control of a noxious weed by limiting its growth.
Since it was unknown whether the chemical effect would depend on the stage of
development of the plant at the time of application, some plants were treated
in an early stage of growth and others later in their growing process. Each
chemical was applied to six different experimental plants, three in early growth
stage and three in late growth, giving a total of 24 experimental plants. Each
chemical was applied at the concentration recommended by the manufacturers.
For confidentiality purposes, the chemicals are known simply as A, B, C and
D. The variable measured was dry matter production at six weeks of age. The
results for the 24 individuals are given below.
7.3. Two-way Factorial ANOVA 7.17
Chemical
Time of application
A B C D
5.9 2.6 4.7 2.0
Early 4.7 3.7 4.3 2.4
5.6 3.4 3.8 1.9
2.5 2.7 3.1 5.9
Late 0.6 2.3 4.1 4.6
1.0 2.9 2.8 5.5
Factor 1 ‘Time of Application’ has 2 levels (Early and Late), Factor 2 ‘Chemical’
has 4 levels (A, B, C, and D) and there are 2x4=8 treatments (e.g. Early-A,
Early-B, Early-C.....Late-C, Late-D). In each of the 8 treatments 3 different
plants where used and the dry matter production at 6 weeks recorded. The
values for each plant are shown in the table above. In total there are 3x8=24
plants. Earlier in this Module we mentioned that all cases that are measured
must belong to only one group/level in each factor. In this example the plants
are the cases and each plant was treated with one Chemical applied at one Time.
No plant was measured more than once and so each plant can be described as
an ‘independent experimental unit’.
Our total sample size N is the total number of independent experimental units
(N=24). Within each of the 8 treatments we have 3 ‘replicates’.
We can calculate the mean of each of the 8 sets of 3 measurements and plot
these mean values:
7.18 Module 7. Analysis of Variance (ANOVA)
The differences between the effect of the chemicals on dry matter production
clearly depends on whether the application occurred during the early or late
stage of growth. If application is in the early growth stage, chemical D appears
to be more effective as it produces much less dry matter of weed; but if ap-
plication is in the later growth stage it is chemical A which is more effective,
and chemical D is the worst of the chemicals (weeds treated with D Late have
a high average production of dry matter). Therefore, it looks like the effect of
chemicals interacts with the time of application — the effect of the chemicals
(which works best) depends on which time of application is being considered.
In the factorial model, instead of the single ‘Treatment’ term we used in the
one-way ANOVA, we now need two terms in the ANOVA table — ‘main effect’
terms and ‘interaction term(s)’.
SSA M SA
Factor A a−1 SSA M SA =
a−1 M SE
SSB M SB
Factor B b−1 SSB M SB =
b−1 M SE
SSAB M SAB
AB Interaction (a − 1)(b − 1) SSAB M SAB =
(a − 1)(b − 1) M SE
SSE
Error N − ab SSE M SE =
N − ab
Total N −1 TSS
When interpreting this output we must consider the significance of the inter-
action term first. In this example the Time x Chemical interaction has a
P -value < 0.001 which is less than the standard significance level of 0.05 and
therefore we reject H0 . We conclude that there is a significant interaction be-
tween the levels of Chemical and Time – the variation in dry matter produced
depends on the interaction between the two factors. Recalling the earlier plot
showing the relationship between Time and Chemical this seems reasonable.
When the interaction term is significant, the main effects would need to be in-
terpreted very cautiously because the significant interaction term has already
alerted us that the main effects depend on each other – interpreting them inde-
pendently must be done only with strong justification for why their individual
effects on the dependent variable can be considered when we already know they
7.20 Module 7. Analysis of Variance (ANOVA)
influence each other in how they affect the dependent variable. If the inter-
action is significant then pairwise tests between all treatment levels would be
needed to determine which treatments are different to each other.
If we imagine for a moment that the interaction term in our example had not
been significant, then we could confidently interpret each of the main effects.
In this example there is a significant difference in mean dry matter between
levels of Time (p < 0.05), however there are no significant differences among
the levels of Chemical (p > 0.05).
We can look at the plots of each main effect to see if these results make sense
(continuing to imagine for a moment that the interaction term was not signifi-
cant). The plot below shows all 12 measures of dry matter for Early and Late
Time of application. The line represents the change in mean values between
the groups. Although there is a lot of overlap in dry matter values between the
two levels, because we had a relatively large number of replicates in each level
(n=12) the ANOVA was able to determine that the observed difference in the
means is unlikely to be due to sampling variation alone, and more likely due
to actual differences in average dry matter production between the two levels.
The plot below shows the 6 replicates within each of the 4 levels of Chemical.
Although the change in mean values across the 4 levels looks similar in magni-
tude to that seen for Time of application, in this analysis there are more levels
to compare and fewer replicates in each level. Therefore, the ANOVA could
not distinguish this variation in means from the variation that might occur just
by chance due to our sampling within each level. While the means are different
in each level, we do not have enough evidence to reject H0 .
7.3. Two-way Factorial ANOVA 7.21
Our interpretation is complicated by the fact that the interaction term was
significant in this example and so we need to look at pairwise t-tests between the
factor level combinations. The spss output is too large to include here, but an
example of the first part is shown below. The first and second columns show the
combinations of Time and Chemical being compared. Row 1 is Time1 (Early)
and Chemical1 (A) compared to Time1 Chemical2 (B) – the mean difference in
dry matter production is 2.1667 and this is a significant difference (sig=0.000,
which is much less than 0.05), so reject the null hypothesis. Time1Chemical1 is
significantly different to all other combinations of Time and Chemical, except
for Time2Chemical4 which has a mean difference of 0.0667 and significance of
p=0.870 (which is larger than 0.05), so we cannot reject the null hypothesis.
In the next section of the first column we can see Time1Chemical2 and it
is compared to all other treatment combinations listed in column 2. For all
comparisons where the Sig. value is < 0.05 we reject H0 and conclude that the
treatment means are significantly different. Where the Sig. values are > 0.05
we do not have enough evidence to reject H0 and so conclude that there is no
statistical difference.
7.22 Module 7. Analysis of Variance (ANOVA)
Exercise 7.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 26.5, 26.7, 26.13.
The One-way ANOVA is a natural way to extend the t-test for testing the means of
two groups to compare several groups when those groups represent different levels of
a single factor.
7.4. Closing Comments 7.23
Two-way factorial ANOVA compares means across groups when those categories form
levels of two main factors. Factors can also interact with each other and the factorial
ANOVA can be used to try to understand the influence of both main effects and the
interaction effect on the variation seen in the dependent variable.
7.24 Module 7. Analysis of Variance (ANOVA)
Suppose the National Transportation Safety Board (NTSB) wants to examine the
safety of compact cars, midsize cars, and full-size cars. It collects data on the pressure
applied to the driver’s head during a crash test and runs three separate crash tests
for each of the car sizes. The SST and the SSE , were found to be 86049.56 and
10254 respectively. Use this information to answer the following:
(a) State the response variable, the factor and the treatments.
(d) What are the degrees of freedom for the treatment, error and total ANOVA table
terms?
(e) What are the M ST and M SE values? (try drawing up the ANOVA table to help
you find these values)
To study the performance of a newly-designed light plane it was timed (in mins) over
a marked course under three wind conditions, calm, moderate and windy. Use the
following output to state the hypotheses and interpret the results:
In investigating electrical outages at three different locations, the total outages (in
mins) per month for five months were examined. The following output was obtained:
(a) What is the P -value to answer the question ‘Do the data provide evidence of
different mean levels of outage at the different locations?’
(b) What is the answer to the question ‘Do the data provide evidence of different
mean levels of outage at the different locations?’
(d) A person trying to answer the questions above carried out two one-way ANOVA’s,
the output of one of which is given below. Fill in the indicated gaps in the output
below.
The time difference between actual and scheduled arrival times of trains at a large
city station are recorded for three days in each of two weeks, and an analysis included
the following output.
7.26 Module 7. Analysis of Variance (ANOVA)
The dry shear strength of birch plywood bonded with different resin glues was studied
with a completely randomised designed experiment. The data from the experiment
is as follows:
(a) For the shear strength plywood data, what are the treatment and error degrees
of freedom?
(b) For the shear strength plywood data, what are the SSE , M ST and M SE ?
(a) Enter the data into spss. Make sure you set up your Data View and Variable
View correctly.
(b) Perform a One-way ANOVA in spss. In your output include Descriptives and
Post-hoc Tukey tests.
(c) State the null and alternative hypotheses, and interpret your analysis.
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
7.28 Module 7. Analysis of Variance (ANOVA)
complete and read an ANOVA table, including identifying df, MSS and the
F -statistic.
interpret the spss output for one-way and a two-way ANOVA and subsequent
pairwise t-tests for significant effects.
Module 8
Association
between
Categorical
Variables
8.2 Module 8. Association between Categorical Variables
Contents
8.1 Contingency Table . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.4
8.2 Associated or Not? . . . . . . . . . . . . . . . . . . . . . . . . . . 8.7
8.3 The Chi-Square Test of Independence . . . . . . . . . . . . . . . 8.10
8.4 The Chi-Square Goodness-of-Fit Test . . . . . . . . . . . . . . . 8.23
8.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.26
8.6 Tutorial Module 8 . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.28
8.7 Module 8 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 8.34
8.3
Module Objectives
On successful completion of this module students should be able to:
interpret and construct, using spss, a stacked bar chart for two categorical
variables;
interpret and construct by hand and using spss the joint, marginal and condi-
tional distributions from a contingency table;
carry out, by hand and using spss, a chi-square test of independence and inter-
pret the results;
carry out, by hand and using spss, a chi-square test of goodness-of-fit and
interpret the results;
calculate the standardised cell residuals, both by hand and by using spss, after
a chi-square test and interpret them;
state and assess the assumptions and conditions necessary for the validity of a
chi-square test; and
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
A variable is categorical if it describes some specific attribute or characteristic of
individuals that can only be categorised, e.g., hair colour (with two ‘values’ dark
and fair), or smoking habit (with three ‘values’ never smoked, past smoker, current
smoker) as responses in a survey.
The ‘by rows’ or ‘by columns’ percentages produce two sets of conditional distribu-
tions. If a set of conditional distributions are different, we argue that an association
exists between the two categorical variables. The problem with looking for a differ-
ence of course is to decide how different is different. In this module we answer that
question.
De Veaux, Velleman & Bock, (5th edition) cover more than we need for
this topic. Also it presents the material in a different order to the way we
do it. Consequently it’s important to follow closely what’s written below
in this module. Reference is made to the textbook, but only to reinforce
what’s written here, not to replace it.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 3, Sections
1-3.
The idea of describing the association between two categorical variables by comparing
conditional distributions is important. Here’s an example from a study of statistics
students.
8.1. Contingency Table 8.5
Example 8.1
Data from a survey of statistics students several years ago gave the following
results:
(a) Describe and interpret the joint distribution, the marginal distributions,
and the conditional distributions of smoking and job possession.
(b) Does whether or not a student has a job explain their smoking habit?
Solution
Firstly we need to calculate the marginal and grand totals as follows:
The four cell percentages adding to 100% comprise the joint distribution.
The 52.4%, for example, may be interpreted as 52.4% of the statistics
students do not have jobs and do not smoke.
There are two marginal distributions, one for whether or not a student
smokes and one for whether or not a student has a job. The 76.2% and
23.8% comprise the first and the 69.1% and 30.9% the second. It is usual to
8.6 Module 8. Association between Categorical Variables
Conditional distributions
There are four conditional distributions in this example, one for each row
and one for each column of the table. The first is obtained by dividing
each cell frequency by the total in its row as shown in the first table below.
The second table below is obtained by dividing each cell frequency by the
total in its column.
with 76.8% and 24.1% with 23.2%) we would conclude the differences are
too small to substantiate the existence of an association.
Exercise 8.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 3.3, 3.5, 3.9,
3.15, 3.29.
The main aim of this module is to do a little better than this by using the ideas of
inference to come up with an objective way of deciding whether or not there’s an
association between two categorical variables.
We all recognise that hair colour and eye colour are associated because dark-haired
people tend to have dark coloured eyes and fair-haired people tend to have blue or
light coloured eyes. Sure, there are exceptions to this ‘rule’, but ‘on average’ that’s
the state of affairs. In other words, the colour of your hair gives some idea of the
colour of your eyes (and the colour of your eyes gives some idea about the colour of
your hair).
What we mean then when we say two categorical variables are associated is that
there is some dependence between them in the sense that knowing the value of one
helps decide the value of the other, or some values of one variable tend to go with
particular values of the other. So, knowing a person is blue eyed increases the chance
the person is fair haired but doesn’t help much in deciding if the person is male or
female, because hair and eye colour are associated whereas eye colour and gender are
not.
Note the language. When two variables are associated we say they are dependent on
each other. And when two variables are not associated we say they are independent
of each other. The words ‘associated’ and ‘dependent’ are used interchangeably.
It’s not enough for us to decide from logical considerations that two variables are
associated or not. After all, what we’re interested in finding out is if two variables
are associated, not from our preconceived notions about what we think the state
of the world is, but directly from data concerning the variables. Also, we need to
8.8 Module 8. Association between Categorical Variables
remember that just because the values of one variable might tend to associate with
particular values of the other variable in our data, that doesn’t necessarily mean the
two variables are definitely associated. It could be that sampling variation has raised
its ugly head and is trying to mislead us. We need to be sensitive to this in coming
up with a method of establishing association.
Example 8.2
Suppose one hundred statistics students are randomly chosen. Of these, sup-
pose 40 are male and 60 are female, and 70 are dark haired and 30 are fair
haired. The contingency table representing these data therefore has row and
column totals as shown here.
We haven’t seen any information yet about how many female students are dark-
haired, or how many male students are fair haired, so we can’t fill the cells of
the table in.
Let’s assume though that gender and hair colour are independent of each other,
as we believe they are. If this is the case then we should be able to more or
less ‘guess’ what the cell frequencies are.
Before reading on, see if you can do this on the basis of what you understand
by independence.
Independence means that the value of one variable doesn’t help determine the
value of the other. In other words, the distribution of one variable is the same
regardless of the value of the other variable. For this example, that means the
distribution of hair colour is the same for both males and females. Also, the
distribution of gender is the same for both dark haired and fair haired students.
How does this help determine the expected cell frequencies? Give it some more
thought if you haven’t determined them yet.
To answer the question, the distribution of gender for all 100 students is 60%
male, 40% female and the distribution of hair colour is 70%, 30% for dark and
fair hair. These distributions are expected to stay the same regardless of hair
colour and regardless of gender. That is, 70% of females are expected to have
dark hair and 30% fair. Similarly, 70% of males are expected to have dark
hair and 30% fair. But there are 60 females and 40 males. Therefore 70%
8.2. Associated or Not? 8.9
Notice, in case you’re wondering, the same expected counts are obtained by
thinking in terms of the gender distribution rather than the hair colour distri-
bution. In other words, 60% of the 70 of the dark-haired students are female
and 40% of the 30 fair-haired students are male giving the same frequencies as
above. This is no accident. It will always be the case.
These cell counts or frequencies that we worked out assuming independence
are called expected counts or expected frequencies. Even if the two variables
are not in fact independent we can still work out expected frequencies using
this logic. There’s a formula that allows us to do this directly from the totals
in the table without going through this logic each time. It is
row total × column total
Exp = (8.1)
table total
Applying this formula to the top left-hand cell (dark-haired female) gives
60 × 70
Exp = = 42
100
as found above. You should confirm this formula works for the other three
expected counts in the table.
Of course, the actual or observed counts may be quite different to the expected
counts. Because of sampling variation, we’d be surprised (or suspicious) if the
observed counts exactly match the expected counts even when the two variables
are independent. We’d be asking, whatever happened to sampling variation?
What is being demonstrated here is the meaning of independence in statistical
terms. If two categorical variables are independent then the respective marginal
and conditional distributions are similar. In this example we saw that this is
the same as noting that two variables are independent if the observed and
expected frequencies are similar.
If the observed and expected frequencies are a lot different however, we’d sus-
pect that the two variables are dependent or associated. How we decide what
we mean by a ‘lot’ different is coming up.
Can you feel a hypothesis test coming on?
8.10 Module 8. Association between Categorical Variables
We now turn our attention to deciding whether or not differences between the ob-
served and expected counts could be due to sampling variation.
Incidentally, this test is more accurately called the chi-square test of association be-
cause we are trying to find out if two variables are associated, rather than not associ-
ated. Regardless it’s universally referred to as the chi-square test of independence or
Pearson’s chi-square test of independence in honour of the statistician whom we can
blame for its existence.
As in all hypothesis testing, we have a null and an alternative hypothesis. The general
form of these for a chi-square test of independence is always the same; namely
H0 : A and B are not associated
Ha : A and B are associated,
where A and B are categorical variables. Don’t get these around the wrong way. Just
remember that a null hypothesis indicates an absence of an effect—in this case, an
absence of association, i.e., no association between A and B.
It’s also fine to use the word independence in place of ‘not associated’ and ‘dependent’
in place of ‘associated’. So ‘A and B are not associated’ can be expressed as ‘A and
B are independent’ and ‘A and B are associated’ can be expressed as ‘A and B are
dependent’.
Of course we always put these statements into the context of the data being tested.
So, for the smoking versus job data from Example 8.1, we would write something like
H0 : Whether a student smokes is independent of whether or not the student has a job
Ha : Smoking status and whether or not a student has a job are associated
Observed counts
We now use the totals to compute the expected frequencies from formula (8.1). Check
these out for yourself.
Expected counts
Notice that two decimal places have been included. (Expected frequencies do not
have to be whole numbers since they can be thought of as long run averages, and it
is recommended they be expressed to two decimal places of precision.)
Also note that marginal totals have been included. They will be the same (allowing for
roundoff) as the totals in the previous table if the calculation of expected frequencies
is correct. You should check the row and column totals once the expected frequencies
have been worked out as a check on your calculations.
where Obs stands for observed count and Exp for expected count. Notice how this
formula is based around the difference between the observed and expected counts in
each cell of the contingency table.
8.12 Module 8. Association between Categorical Variables
What does this mean for our data? There are four cells in the table. Therefore there
are four contributions of the sort (Obs − Exp)2 /Exp to calculate. The formula asks
us to add these to give
(161 − 161.59)2 (73 − 72.41)2 (51 − 50.41)2 (22 − 22.59)2
χ2 = + + +
161.59 72.41 50.41 22.59
= 0.029
The value of χ2 depends on the size of the differences between the observed and
expected counts. It doesn’t matter whether the difference is positive or negative.
Squaring the difference gets rid of the negatives. The bigger the differences, the
bigger the value of χ2 . Or putting it around the other way, if the observed and
expected frequencies are similar in each cell, the value of χ2 will be small.
That’s correct. Zero. If the corresponding observed and expected counts are exactly
the same, χ2 would be zero. Mind you, if that happened, we would wonder what
became of sampling variation, and we might be a bit suspicious of the data. But in
theory it could happen. Chi-square then must be positive, or zero at least.
For our data you will have noticed the observed counts and the expected counts are
quite similar. This suggests, even without working out the chi-square value, that
we are not going to be able to reject the null hypothesis of no association. The
small chi-square value of 0.029 confirms this observation. We’ll continue through the
procedure though to show how to convert the test statistic into a P -value as is usual
in hypothesis testing.
How do we find the P -value? Well, as with the t statistic, someone has kindly
worked out what sort of values are to be expected for χ2 when the null hypothesis is
true. That is, they’ve come up with a distribution of the values of χ2 when the two
categorical variables defining a contingency table are independent. This distribution
can be found as Table X at the back of De Veaux, Velleman & Bock. We can look up
the value we found for our test statistic and find a range in which the P -value lies.
Have a look at Table X. Notice the picture of the distribution. What shape has it
got? What values does it take on? Also notice we need to know what degrees of
freedom (df) are involved before we can use the table, just like we did for Table T.
Degrees of freedom
Chi-square measures the difference between the actual counts we observe and the
counts we expect to get if there is no association. If chi-square is ‘small’ (close to
8.3. The Chi-Square Test of Independence 8.13
zero) then the observed and expected counts are similar and we won’t be able to reject
H0 . If chi-square is ‘large’, then the observed and expected counts differ considerably
and we would be able to reject H0 .
Well that depends not only on the size of the differences between the observed and
expected counts but also on the number of cells in the table.
After all, the larger the table (i.e., the more cells), the larger potentially chi-square can
be, because every cell contributes to chi-square. We need therefore to take account
of the number of rows and columns in the table when using Table X. The degrees of
freedom does this because it measures the size of the table.
The formula for degrees of freedom is easy. If there are r rows and c columns in the
contingency table we have
df = (r − 1)(c − 1) (8.3)
So, for example, if the table has two rows and four columns (excluding labels and
totals of course), df = (2−1)(4−1) = 3. The smallest possible contingency table with
just two rows and two columns (as in the gender/job example) has just one degree of
freedom.
To use Table X then we need to go to the row with the appropriate number of degrees
of freedom and look across that row and locate the whereabouts of the chi-square
value. From this the figures at the top of the table indicate the range for the P -value.
If df = 3 and χ2 = 10.5, for example, then the chi-square value is between 9.348 and
11.345 according to Table X. The figures at the top of these columns, 0.025 and 0.01,
are P -values. The P -value then for χ2 = 10.5 with 3 degrees of freedom is between
0.01 and 0.025 or 1% and 2.5%.
As another example, if χ2 = 15.8 and df = 3, then because 15.8 is greater than 12.838,
the P -value is less than 0.005.
Notice that as the χ2 value gets larger, the P -value gets smaller. This is typical of
test statistics and P -values. The larger the magnitude (i.e., ignoring the sign) of a
test statistic such as t or z or χ2 , the smaller the P -value and vice versa. This is
to be expected. A large magnitude for the test statistic indicates the data does not
agree with the null hypothesis—in other words there is only a small probability (the
P -value) that sampling variation alone is the cause.
Back to our example with smoking and jobs. The degrees of freedom for our 2×2 table
is one and χ2 = 0.029. From the df = 1 row in Table X we see that 0.029 is smaller
than the smallest entry, 2.706. We conclude that the P -value is greater than 10%.
Consequently, as we thought earlier, there is insufficient evidence to conclude that
8.14 Module 8. Association between Categorical Variables
smoking and whether or not a student has a job are associated, i.e., the differences
between the observed counts in the sample and the expected counts based on marginal
distributions are most likely due to sampling variation.
Example 8.3
See De Veaux, Velleman & Bock, (5th edition) R6.42 (Review Section VI after
Chapter 23). The data in this question has to first be converted into a table of
observed counts as follows.
Follow-up analysis
It’s not altogether satisfactory to conclude the existence of an association without
investigating the nature of the association or the major source(s) of the association.
An appropriate way of doing this involves looking carefully at the differences between
the observed and expected counts. These differences are called residuals. Therefore
the residual for each cell is given by
residual = observed count − expected count
Make sure you understand where these are coming from. We’ve simply subtracted the
table of expected counts from the table of observed counts. Notice that the sum of
the residuals in each column and in each row is zero. This is typical of what happens
with residuals. You might recall that the sum of the residuals in a regression analysis
is zero.
Now it turns out that although these residuals tell us directly which observed counts
are bigger than the corresponding expected count and which are less, the importance
of each residual in the overall significance of the association depends also on the
expected number in each cell itself. Putting this another way, a residual of 10 (or
−10) is more important if the expected cell count is small, like 40 say, than if the
expected count is large, like say 100. To take account of this we convert the residuals
into standardised residuals by dividing each by the square root of its expected count:
observed count − expected count
standardised residual = √ (8.4)
expected count
The standardised residuals for the data in Example 8.3 are then
< 3 years HS 3+ years HS Some college
Planned −3.16 0.87 3.58
Unplanned 2.70 −0.74 −3.06
Compare the formula for chi-square (8.2) with the one for standardised residuals (8.4)
and you’ll see that a standardised residual for a cell is just the square root of the chi-
square contribution for the cell. Or putting it another way, the value of the chi-square
test statistic is just the sum of all the standardised residuals squared.
8.16 Module 8. Association between Categorical Variables
For example, in Example 8.3, the contribution to chi-square of the (Planned, Some
college) cell is (137 − 101.05)2
√ /101.05 = 12.79 and the value of the standardised
residual is (137 − 101.05)/ 101.05 = 3.58, which is the square root of the chi-square
contribution (allowing for round-off).
The relative size of the standardised residuals (ignoring any negative sign) indicates
the relative importance of the cell to the chi-square value. If the chi-square value
is significant (i.e., it is ‘large’ with a small P -value) then the cell with the largest
standardised residual is the main contributor. By considering the direction of the
residual for this cell, the main source of the significant association can be described
in the context of the data.
Example 8.4
Following-up on the analysis in Example 8.3 we see that the largest standard-
ised residual ignoring the negative sign is 3.58 for the (Planned, Some college)
cell. (If there was a negative sign, it would indicate that the observed count
is less than what is expected if there were no association between unplanned
pregnancies and education level.) We conclude that the main reason a statis-
tically significant association is observed is that more educated women tend to
have fewer unplanned pregnancies.
Note that it only makes sense to do a follow-up analysis of the residuals if there is a
statistically significant association. Only the largest one or two standardised residuals
need to be looked at and interpreted.
This module is self-contained so the following reading only reinforces material already
covered.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 3.
2. The sample is not more than 10% of the population (the 10% condition);
3. The expected counts are all at least five (the expected count condition).
The last condition insists in effect that the sample is reasonably large so that each
cell is expected to contain at least 5 individuals. Note, it’s not the observed counts
we worry about—some of these could even be zero—but the expected counts.
What do we do if some of the expected counts are less than 5? One approach is
to reduce the number of rows or columns of the contingency table by combining
categories together. A row (or column) containing an expected count less than 5 is
combined with another row (or column) so that the problem cell is eliminated. This
may need to be repeated to eliminate all problem cells. It would also need to make
sense in the context of the situation to combine categories in this way. Obviously a
2 × 2 contingency table can’t be reduced any further and there are other methods of
dealing with this problem which are not covered in this course.
We only need to know about the technique of combining rows or columns. Even then,
none of the set problems will require you to do this. Also, as usual, it’s not necessary
to state and check these conditions unless specifically requested in a problem.
This module is self-contained so the following reading only reinforces material already
covered.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 4.
For example, if smoking and having a job are independent, the conditional distri-
butions by rows will be the same. But that’s the same as saying the proportion of
smokers with a job is the same as the proportion of non-smokers with a job; and the
proportion of smokers without a job is the same as the proportion of non-smokers
without a job. And since the proportions without a job can only be equal if the pro-
portions with a job are equal, we could state the null hypothesis as H0 : p1 = p2 where
p1 is the proportion of smokers with a job and p2 is the proportion of non-smokers
with a job.
Of course, independence also means the conditional distributions by columns are the
same. Therefore, the null hypothesis could also be stated as H0 : p1 = p2 where p1
is the proportion of those with a job who smoke and p2 is the proportion of those
without a job who smoke.
This idea extends to larger tables. For Example 8.3, the null hypothesis could be
written either as
H0 : Unplanned pregnancies and education level are independent, or
H0 : The distribution of unplanned and planned pregnancies is the same at each
education level,
and this second way of stating the null hypothesis could be re-expressed in terms of
proportions. That’s a bit messy because there’s quite a few proportions involved and
it doesn’t add to the idea of this section.
The point is, a test of independence is the same as a test of equality (i.e., homogeneity)
of (conditional) distributions which is the same as testing equality of one or more sets
of proportions.
De Veaux, Velleman & Bock, (5th edition) treat the test of independence and test
of homogeneity of proportions separately (see Chapter 22, Section 2). There are
occasions where it makes sense to do that, but we don’t make the distinction in this
module and you won’t be penalised for not making the distinction. Consequently, if,
for example, interest is in assessing whether there is a difference between males and
females in the proportions who smoke, the appropriate null and alternative hypotheses
are
H0 : Gender and smoking are not associated
Ha : Gender and smoking are associated
Or, if we’re interested in whether the grade distribution of on-campus and online
8.3. The Chi-Square Test of Independence 8.19
Using SPSS
For any but a small contingency table, a chi-square test requires a lot of calculation.
You will not be expected to calculate chi-square ‘by hand’ for Tables with more than
six cells.
Exercise 8.2
Watch the following spss videos. After watching the videos try to
replicate the analysis in Example 8.5 below.
Weight Cases
When the data is given in a summarised form, such as a table of counts as in Exam-
ple 8.3, these spss videos and the description below indicate the process for analysing
this data using a chi-square test. If the data is given in raw form, i.e., data values are
given for all cases, you will not need to apply Weight Cases. The following example
provides a brief outline of the method for summarised data such as might be found
in textbook exercises.
Example 8.5
Examples 8.3 and 8.4 are repeated here using spss.
To enter the data, each of the observed counts is identified by a row and column
number (see Figure 8.1). So, for example, 391 is in Row 2, Column 1. To make
the spss output look good of course, include variable value and data value
labels for the ‘Row’ and ‘Col’ variables and their coded values in the Variable
View window.
8.20 Module 8. Association between Categorical Variables
Weighting is now applied as shown in Figure 8.2 using the Data/Weight Cases
procedure and ‘count’ as the weighting variable (see the spss video Weight
Cases). This tells spss that there are 200 individuals with < 3 years HS
who had planned pregnancies, 271 with 3 or more years HS who had planned
pregnancies, etc. (If this is not done, spss thinks each row of the Data View
spreadsheet represents just one individual.)
Figure 8.3: spss dialog box for contingency table and chi-square test procedure.
Notice the Observed, Expected Counts and Standardised Residuals are dis-
played followed by the value of the test statistic (called ‘Pearson’s Chi-Square’),
40.715, degrees of freedom and P -value (‘Asymp. Sig. (2-sided)’) = .000 (i.e.
it is 0.000 when rounded off to three decimal places). Note that the chi-square
test is a two-sided test.
Notice how spss also does a check on the expected cell frequency condition.
Of course, it won’t stop reporting the results of the test even if the condition
fails. We have to be aware of what is valid and what isn’t.
8.22 Module 8. Association between Categorical Variables
In general terms, what does the null hypothesis state in a test of independence?
What about the alternative hypothesis?
What is the formula for the degrees of freedom for the test of independence?
Exercise 8.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 22.9 (by hand),
22.39 (use spss), 22.41 (use spss).
De Veaux, Velleman & Bock, (5th edition) describe three chi-square tests: a goodness
of fit test, a test of independence and a test of homogeneity of proportions. The last
two have been discussed in this module. Now it is time to turn our attention to the
Goodness-of-Fit test.
Example 8.6
The manufacturer of Beanies (a fruit-flavoured gummy-textured sweet in the
shape of a broad bean seed) is concerned that their mixing machine is not
calibrated correctly. If the mixing machine is properly calibrated to reflect
advertised ratios of colours, the distribution of colours of Beanies should be
30% red, 30% green, 20% orange and 20% purple. A random sample of 180g
bags of Beanies were removed from the packaging machine and yielded the
following data:
Colour
Red Green Orange Purple Total
Count 902 1096 654 806 3458
Percent 26.1 31.7 18.9 23.3 100%
Expected % 30 30 20 20 100%
8.24 Module 8. Association between Categorical Variables
How ‘good’ does the data fit the probability model we would expect if the
mixing machine was correctly calibrated?
Step 1: Write hypotheses
H0 : The distribution of colours of Beanies follows the manufacturer’s require-
ment of 30% red, 30% green, 20% orange and 20% purple.
Ha : The distribution of colours of Beanies is different to the manufacturer’s
requirement of 30% red, 30% green, 20% orange and 20% purple.
Step 2: Assuming that the distribution is as required, calculate the χ2 test
statistic
X (Obs − Exp)2
χ2 =
Exp
all cells
However, before we can do this we need to calculate expected counts based on
the manufacturer’s required distribution of colours.
Colour
Red Green Orange Purple
Observed Count 902 1096 654 806
Expected Count 1037.4 1037.4 691.6 691.6
SPSS Output
The initial output for performing the Chi-square Goodness-of-Fit test in spss
gives only the Sig. or P -value of .000 and the decision to reject the null hypoth-
esis. Notice that it says that this is a ‘One-sample Chi-Square Test’ which indi-
cates that the steps to doing this test in spss require you to click on ‘Analyze >
Nonparametric > Legacy Dialogs > One Sample’ to access the Goodness-of-Fit
test.
If you click on this table, it will show more detailed output that includes the
sample size, test statistic, degrees of freedom and P -value.
8.26 Module 8. Association between Categorical Variables
This module is self-contained so the following reading only reinforces material already
covered.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 1.
Exercise 8.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 22.19 (by hand
and using spss), 22.21.
This module covers the material on contingency tables and chi-square tests. First
we dealt with converting the table to percentages in three different ways and talked
about the idea of association. We’ve now got a method of assessing the statistical
significance of an association – Chi-square Test of Independence. In addition we can
also test whether a frequency distribution fits a specific pattern – Goodness-of-Fit.
One final point. Remember that the chi-square test is applied to frequency data. Even
though we often convert frequencies to proportions for descriptive purposes and to
improve our appreciation of the relationships amongst the groups under consideration,
the chi-square calculations must be performed on the original frequencies or counts,
not the percentages.
8.28 Module 8. Association between Categorical Variables
The following data was collected and organized into the Contingency Table below to
see if there is a relationship between amount of coffee consumed (stated as ‘3 or less
coffees per day (Light)’ and ‘more than 3 coffees per day (Heavy)’) and the number
of hours slept per night (stated as ‘7 or less hours per night’ and ‘more than 7 hours
per night’) for 151 statistics students.
(a) Display the joint distribution of level of coffee consumed and level of sleep in
terms of percentages.
(b) Using the joint distribution in (a), write 4 sentences describing what each of these
4 percentages mean (i.e. one of the sentences could be something like ”90% of
statistics students drink 3 or less cups of coffee per day and have 7 or less hours
of sleep per night”).
(e) Sketch a suitable graph to display the relationship between the two variables,
using the information from (d) (there are two ways of doing this, either using
the information from your answer to d(i) or the information from your answer to
d(ii)).
8.30 Module 8. Association between Categorical Variables
(f) From the data in the study, what percentage of statistics students are heavy coffee
drinkers AND sleep 7 or less hours per night?
(g) What percentage of students who are light coffee drinkers also sleep 7 or less
hours per night?
(h) What percentage of students who sleep 7 or less hours per night are light coffee
drinkers?
NOTE on wording:
Saying that two categorical variables are not associated (or not related) is
the same as saying they are independent.
Saying that two categorical variables are associated (or related) is the same
as saying they are NOT independent.
(m) In (k) we examined the association between two categorical variables (level of
sleep and level of coffee consumed). To further test your understanding of cate-
gorical variables and association, can you suggest:
(i) Two other categorical variables that you would expect to be independent?
(ii) Two other categorical variables that you would expect to be associated?
8.6. Tutorial Module 8 8.31
Students in a statistics lecture were asked to clasp their hands together, fingers inter-
twined, and then note which thumb was on top (right thumb equates to Right Hand
Dominance). They were then asked to cross their legs, one ankle over the other, and
note which ankle was on top (right ankle equates to Right Leg Dominance). Using the
answers to these questions, the following two-way table was produced (Figure 8.8):
(a) Conduct an appropriate hypothesis test by hand to answer the research question:
Step 0: Compute the conditional distribution of leg dominance for right hand
dominance and then left hand dominance. Compare these two distributions and
give a tentative answer to the question.
Step 1: Write the hypotheses for answering this question.
Step 1a: What assumptions/conditions needed to be satisfied to perform this
test?
Step 1b: Compute a table of the expected frequencies/counts.
Step 2: Calculate the test statistic and give the degrees of freedom for the test.
Step 3: Calculate the P -value.
Step 4: Write a conclusion in context.
Step 5: Compute the standardised residuals and interpret the relationship be-
tween Hand Dominance and Leg Dominance.
(b) Now repeat the hypothesis test using spss to complete the Steps 1b, 2 and 3 as
in (a). Compare your answers with those in (a).
8.32 Module 8. Association between Categorical Variables
Use this output and the data above to answer the following questions:
(a) Write the hypotheses for testing if Education Level Attained is associated with
Age Group.
(b) What assumptions/conditions are needed to be satisfied to perform this test?
8.6. Tutorial Module 8 8.33
(c) State the test statistic, df and P -value (from the spss output) and use these to
give a conclusion in the context of the question.
(d) Now reproduce these results in spss for yourself (the data file for this question
can be found in the Module 8 materials on the StudyDesk ). Note that you will
be required to perform some analyses in spss for your assignment.
Bill, Clara and Zoe are playing a board game, but Zoe suspects that the die they are
using may have been damaged in the manufacturing process such that outcomes of
a roll, the numbers 1 to 6 are not equally likely (as one would expect they should be
with a fair die). She decides to test her suspicion by tossing the die 300 times and
obtains the following results after analysing her data in spss:
(b) State any assumptions/conditions that need to be satisfied to conduct this test.
(d) Give a conclusion for this test in the context of the question.
Do the calculations for this hypothesis test by hand (show all working):
(e) Assuming the null hypothesis is true, calculate the test statistic.
(g) Determine the P -value for this test using the χ2 Distribution Table.
Question 5: Differences
How is the Chi-square Goodness-of-Fit Test different to the Chi-square Test of Inde-
pendence?
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
carry out by hand and using spss a chi-square test of independence and interpret
the results;
carry out by hand and using spss a chi-square test of goodness-of-fit, and in-
terpret the results;
calculate the standardised cell residuals both by hand and by using spss after
a chi-square test and interpret them;
state and assess the assumptions and conditions necessary for the validity of a
chi-square test;
Contents
9.1 Inference for Simple Linear Regression . . . . . . . . . . . . . . 9.4
9.1.1 The Simple Linear Regression Model . . . . . . . . . . . . . . . . . 9.5
9.1.2 Assumptions & Conditions . . . . . . . . . . . . . . . . . . . . . . 9.5
9.1.3 Inference about Regression Parameters . . . . . . . . . . . . . . . . 9.7
9.2 Multiple Regression Model . . . . . . . . . . . . . . . . . . . . . 9.15
9.2.1 Parameter Estimation by OLS Method . . . . . . . . . . . . . . . . 9.16
9.2.2 Assumptions for a Multiple Regression Model . . . . . . . . . . . . 9.18
9.2.3 Inferences on the Regression Parameters βi . . . . . . . . . . . . . 9.20
9.2.4 Comparing Multiple Regression Models . . . . . . . . . . . . . . . 9.23
9.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.24
9.4 Tutorial Module 9 . . . . . . . . . . . . . . . . . . . . . . . . . . 9.25
9.5 Module 9 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 9.27
9.3
Module objectives
Upon completion of this module students should be able to:
define the difference between the sample regression line and the true regression
line;
state the assumptions necessary for inferential t procedures to be valid for simple
linear regression;
use appropriate graphs to check the assumptions where possible for the validity
of inferential regression procedures;
interpret the meaning of the intercept, slope and error standard deviation in
the simple linear regression model;
perform and interpret the results of a test about the slope using spss;
determine and interpret a confidence interval for the slope using spss;
determine, using spss, a confidence interval for the mean response at a given
value of x and a prediction interval for an individual response at a given value
of x;
use a simple regression model to predict a particular value of a future response;
determine the residuals and interpret the residual plot;
define a linear multiple regression model;
define the parameters of a multiple linear regression model;
state the assumptions of a multiple regression model;
determine estimates of the parameters of a multiple regression model;
find a best fitted regression line of a multiple regression model;
interpret the coefficients of a multiple regression model;
predict a particular value of a future response of a multiple regression model;
find the residuals and interpret the residual plot;
find a confidence interval for the parameters of a multiple regression model;
perform hypothesis tests on the parameters of a multiple regression model;
find the adjusted coefficient of determination and interpret it; and
check the assumptions of a multiple regression model.
9.4 Module 9. Linear Regression Analysis
Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
Introduction
The idea of simple linear regression and correlation were introduced in Module 3.
It may be worth reviewing that material before studying this module. We assume
familiarity with the idea of explanatory and response variables, the equation of a
regression line and of being able to determine, using spss, the correlation coefficient
and least-squares regression line and be aware of the potential impact of outliers
and influential observations on a regression line. We build on material presented in
Module 3 to discuss inferences involving simple regression and then extend this to
multiple regression.
De Veaux, Velleman & Bock, (5th edition) covers the material in this module in
Chapters 23 & 24.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Section 1
The main idea here is that the points we see in a scatterplot of two quantitative
variables are just a sample from a population. The line we put through the scatterplot
may not be the same as the ‘true’ line – the one we would put through the population
of points if we knew what they were. The statistics b0 , b1 , yb and e are all to do with
the line through the sample, and we can calculate all of them from the data. The
parameters β0 , β1 , µy and (notice they’re all represented by Greek letters) are all to
do with the population, and we don’t know the values of any of them! Inference about
regression is all about trying to figure out the values of these parameters. To do this,
we need to make some assumptions, including, as usual one concerning normality,
because we’re going to use the t distribution model again.
9.1. Inference for Simple Linear Regression 9.5
y = β0 + β1 x + (9.1)
In this case, this random or unexplained variability would include effects on weight
not related to height, for example, bone density, body shape, etc. Thus we would
not expect that all people with the same height would have the same weight. So
although we have postulated a model it will not, in general, fit any particular set of
data perfectly.
To complete the statistical model we need to know the properties of and relationship
between the response and explanatory variables.
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Sections 2
and 3.
What assumptions about the response variable y are required to proceed with
inference calculations?
How is R2 interpreted?
9.6 Module 9. Linear Regression Analysis
(i) Linear relationship: the response and explanatory variables must be linearly
related. This is often checked by visual observation of the scatter plot.
(ii) Independence of errors: the error terms are independent of each other. This
also implies that the responses are independent. This is checked by a residual
plot.
(iii) Equal error variance: the error variance is constant (or equal). This means
that the spread of the response (y) is the same for all values of the explanatory
variable (x). That is, V ar[] = σ 2 . This is checked by a residual plot.
(iv) Normal distribution of errors: the errors around the idealised regression line
follow a normal distribution for all values of x. This is checked by the histogram
or normal probability plot of the residuals.
The main point here is that whenever we do some inference there are assumptions or
conditions involved. We can anticipate the normality condition because we’ll be using
the t distribution again, as we did in Modules 5 and 6. And just as mentioned in
those modules, the more data we have the less important is the normality assumption.
Independence always appears somewhere in any assumptions. We are not saying that
the explanatory and response variables are independent – if they were we wouldn’t be
bothering with a regression at all! We’re saying that the n pairs of measurements of
the x and y values are independent. This often can’t be checked except by knowing
how the data was measured or collected. Often the best we can report is that we
assume the n pairs of measurements of x and y were made independently. The equal
variance assumption or equal spread assumption is a new type of condition. All it’s
saying is that the scatter of points in the vertical direction is approximately uniform
along the regression line. You may see this referred to as homoscedasticity.
The usual trick in checking assumptions for regression is to make good use of the
residuals. The routine suggested in Chapter 23, Section 2 of De Veaux, Velleman &
Bock, (5th edition) is fine. Of course, judgement is involved in deciding whether to
go ahead with inference. That won’t be a major issue for the problems asked in this
course – we’ll be pretty specific about what has to be done. In the bigger scheme of
things though, checking assumptions sets apart the user of statistics from the abuser
of statistics.
9.1. Inference for Simple Linear Regression 9.7
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Section 3.
How many degree of freedom has the Student’s t-model when used in simple
regression?
What are the two types of prediction intervals that can be produced? What is
the difference between them?
Why do we call one interval a confidence interval, and the other a prediction
interval?
Which of the formulas mentioned should we know about or are we going to be asked
to use? Firstly, from the previous section (and from Module 3), we should be aware
that se represents the standard deviation of the residuals and is given by
rP
(y − yb)2
se = (9.2)
n−2
9.8 Module 9. Linear Regression Analysis
This statistic describes the amount of scatter around the line (in the vertical direction)
and is a fundamental quantity. It’s given a few different names – the error standard
deviation, the residual standard deviation or the standard error of the estimate in
spss. If its value were zero then all the points would be on the regression line.
Otherwise, like any standard deviation or standard error it’s positive. The larger it
is, the more scattered the points around the line and the larger will be the margins
of error in our intervals. The n − 2 in the denominator of (9.2) is called the degrees
of freedom for se , and as such is the degrees of freedom for all simple regression
calculations. We need to know the formula for a confidence interval for the slope β1 :
b1 ± t∗n−2 × SE(b1 ) (9.3)
Also, we need to know how to calculate the test statistic for H0 : β1 = 0; that is
b1 − β1
t= (9.4)
SE(b1 )
There’s a formula for SE(b1 ) but we don’t use it. When we need to use (9.3) or (9.4)
we will be given SE(b1 ) somewhere (probably in some computer output). That’s it.
We don’t need to do any inference on β0 , and when asked for prediction intervals for
y or confidence intervals for µy , we should be aware that they have the form
yb ± t∗n−2 × SE (9.5)
but we won’t ever have to use these formulas. Usually we’ll just have to read the
interval from computer output.
Example 9.1
Students in a Statistics tutorial measured forearm length and head circumfer-
ence in centimetres as shown:
Forearm Head Forearm Head
Subject length (cm) c’ference (cm) Subject length (cm) c’ference (cm)
1 43.5 60 13 48.0 58
2 40.0 57 14 44.0 58
3 44.0 57 15 47.0 59
4 52.0 59 16 47.0 59
5 46.0 58 17 38.0 56
6 46.5 62 18 41.0 58
7 43.5 56 19 43.0 56
8 44.5 57 20 47.5 57
9 40.0 57 21 45.0 59
10 43.5 56 22 46.5 56
11 46.0 57 23 49.5 58
12 46.0 63 24 48.0 59
9.1. Inference for Simple Linear Regression 9.9
Tick the ‘Save’ and ‘Plots’ options in the spss Regression procedure as indicated
in Figures 9.2 and 9.3. One of the results is that we get six more variables in
the Data View Window: ‘PRE 1’ (= yb), ‘RES 1’ (= e, the residuals), ‘LMCI 1’
and ‘UMCI 1’ (95% confidence interval bounds for µy ), and ‘LICI 1’ and ‘UICI
1’ (95% prediction interval bounds for y). Part of the new Data View is shown
in Figure 9.4.
Make a scatterplot of the residuals ‘RES 1’ (or ‘unstandardised residuals’ as
spss calls them) against the predicted values ‘PRE 1’ (‘unstandardised predic-
tions’ in spss). We get Figure 9.5. There is a lot of scatter but it’s plausible to
accept that there is no systematic thickening or thinning so the equal spread
assumption is reasonable. Also there are no obvious outliers. No information is
given as to the order in which the data was collected so the residuals cannot be
plotted against time. We can therefore only state there is no reason to assume
the observations are not independent.
9.10 Module 9. Linear Regression Analysis
A check on the normality condition is given in Figures 9.6 and 9.7. Normality
is acceptable on the basis of these plots.
We feel justified now in interpreting the results of the regression procedure,
shown in Figure 9.8. Ignore the table titled ‘ANOVA’. Notice the correlation
coeffcient is r = 0.380 and R2 = 14.4% telling us that approximately 14% of the
variability in head circumference is explained by variability in forearm length
for these data. This confirms the description of the relationship as being weak
– 86% of the variability in head circumference is explained by other variables
apart from forearm length.
9.12 Module 9. Linear Regression Analysis
The error standard deviation or ‘standard error of the estimate’, as spss calls
it, is se = 1.719 cm. This is really the estimated standard deviation of head
circumferences for individuals of a particular forearm length. A good way of
knowing whether se is ‘large’ or ‘small’ is to compare it with the standard devia-
tion of all 24 head circumferences that were measured. That value is sy = 1.818
9.1. Inference for Simple Linear Regression 9.13
\
head circumference = 48.316 + 0.215 × forearm length
which is fine except it should be made clear also that head circumference and
forearm length are measured in centimetres. The standard error of b0 and b1 are
given in the ‘Coefficients’ table. The important one is SE(b1 ) = 0.112 because
this tells us how precisely the slope of the line has been measured which leads
9.14 Module 9. Linear Regression Analysis
to a confidence interval and hypothesis test. From 9.3 we can obtain a 95%
confidence interval for the true slope β1 as
The most important thing about this interval is that it doesn’t exclude zero.
Therefore we are unable to reject H0 : β1 = 0 in favour of Ha : β1 6= 0 at the
5% level of significance. In other words, the P -value for this test will be bigger
than 5%.
It’s easy to confirm this from the ‘Coefficients’ table where we see t = 1.927 is
the value of the test statistic with (2-sided) P -value = 0.067. Putting this result
into words we can conclude from these data there is only weak or slight evidence
in favour of arguing that head circumference can be predicted from forearm
length. We should also be aware how the formula (9.4) applies. Substituting
b1 and SE(b1 ) into (9.4) gives
b1 − β1 0.215
t= = = 1.92
SE(b1 ) 0.112
as reported by spss (within round off error).
Predictions are produced by spss in the Data View as we see in Figure 9.4.
So, for example, a 95% prediction interval for head circumference of a student
with forearm length of 46 cm is (54.57, 61.86) cm. In other words, with 95%
confidence a student with forearm length of 46 cm can be expected to have a
head circumference somewhere between 54.6 cm and 61.9 cm.
In the same vein, with 95% confidence the mean head circumference of all
Statistics students with forearm length of 46 cm is expected to be somewhere
between 57.5 cm and 59.0 cm.
What say we want to make a prediction for an x value that is different to those
in the data? For example, a forearm length of 45.4 cm. All we need do is
include that value (or those values) of x at the end of the ‘x’ column in the
Data View and run the regression procedure as before. The analysis won’t
change but prediction intervals will be generated for the ‘new’ data. Check
yourself that a 95% confidence interval when x = 45.4 is (57.35, 58.82) and a
95% prediction interval is (54.45, 61.73).
2. Interpretation
b1 is the slope of the fitted regression line. It is the rate of change in ŷ for
one unit change in x. In Example 9.1, the increase in head circumference per
additional cm increase in forearm length is b1 = 0.215 cm.
b0 is the intercept of the fitted line. This is the estimated value of yb when x = 0.
Usually there is no or very little interest in it from a statistical point of view.
Here for example, it is nonsensical to say that forearm length of zero, relates to
a head circumference of 48.3 cm.
Exercise 9.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 23.1, 23.3, 23.5,
23.7, 23.9, 23.11, 23.23, 23.57 (use spss).
Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Sections 5
& 6 and Chapter 24.
The multiple regression model and inferences on it are covered in this section.
In real life the response variable y depends on several explanatory variables. If there
are two or more explanatory variables in a linear regression model, it is called a multi-
ple regression model. Let us consider the linear regression of y on k(≥ 2) explanatory
variables, say X1 , X2 , . . . , Xk , then the multiple regression model is defined as
y = β0 + β1 x1 + β2 x2 + · · · + βk xk + , (9.7)
where
Notes:
1. β0 , β1 , . . . , βk are estimated using n sets of observations
(yi , xi1 , xi2 , . . . , xik ), i = 1, 2, . . . , n where n ≥ k + 1. This will allow us to ‘fit’
the model to the data.
However, in all of the above, the x-values are assumed to be measured without error.
If the x-variables are measurements or counts as in (a) and (b) they are said to
be quantitative variables, whilst those in (c) are called qualitative (or categorical)
variables. In this module, we will only consider quantitative variables.
yb = b0 + b1 x1 + b2 x2 + · · · + bk xk , (9.8)
Example 9.2
To predict the house price (response variable, y) one could use the floor area
(x1 ), number of bedrooms (x2 ) and number of bathrooms (x3 ) as explanatory
variables. The following data set is typical of multiple regression analysis.
Here all three explanatory variables are numerical/quantitative.
Figure 9.9: House Price Data (n = 10) for Multiple Regression Analysis
The fitted regression line is obtained by using the regression coefficients pro-
duced by spss as follows:
yb = 543.382 + 0.076x1 − 76.905x2 + 15.832x3 , (9.9)
where the intercept is b0 = 543.382, coefficient of x1 is b1 = 0.076, coefficient
of x2 is b2 = −76.905, and coefficient of x3 is b3 = 15.832.
(i) Linear relationship: The response and explanatory variables must be linearly
related. This is checked by visual observation of the scatter plot. From Fig-
ure 9.12, the house price has strong linear positive relationship with both living
9.2. Multiple Regression Model 9.19
area and number of bathroom, but negative very weak linear relationship with
number of bedrooms. Multicollinearity – linear relationship between pairs of
explanatory variables can be an issue. From Figure 9.12, we can see that mul-
ticollinearity is not an issue in this case as none of the explanatory variables is
linearly related to any other explanatory variable.
(ii) Independence of errors: The error terms are independent of each other. This is
checked by a residual plot. From Figure 9.13, there is no trend or pattern in
the residual plot, and hence there is no concern of violation of this assumption.
(iii) Homoscedasticity: The error variance is constant (or equal), i.e., the spread of
the response (y) is the same for all values of the explanatory variable (x). This
is checked by the residual plot. From Figure 9.13, there is no funnel type shape
in the residual plot, and hence there is no concern of violation of the equal
variance assumption.
(iv) Normal error distribution: The errors around the idealised regression line follows
a normal distribution for all values of x. This is checked by the histogram
or normal probability plot of residuals. From Figure 9.14, the histogram of
residuals does not show any serious violation of the normality assumption. A
Q-Q plot also supports the same.
9.20 Module 9. Linear Regression Analysis
Since the assumptions are satisfied, we could proceed with the inference on the re-
gression parameters.
When the assumptions of the model hold good, one could find confidence intervals
and test hypotheses about the true values of the regression parameters βi .
To test H0 : βi = βi0 , any fixed value, against Ha : βi 6= βi0 , use the test statistic
(bi − βi0 )
t=
SE(bi )
bi ± tn−k−1,α/2 × SE(bi ).
Example 9.3
Refer back to the house price data in Example 9.2.
1. Test that the living area is not significant, (i.e., β1 = 0) using a two-sided
alternative at α = 0.05; and
2. Find a 95% confidence interval for the coefficient of living area, β1 .
Solution
Again from the spss output, the two-sided P -value is 0.018 which is less
than α = 0.05, so H0 is rejected. That is, there is sufficient evidence to
claim that the β1 is significantly different from zero at the 5% level of
significance.
2. Note here df = n − k − 1 = 10 − 3 − 1 = 6, hence tα/2 = t6, 0.05/2 = 2.447.
A 95% confidence interval for β1 is
Example 9.4
Refer back to the house price data in Example 9.2. spss Output from the
Regression procedure for this overall test is shown in Figure 9.15.
With a P -value of 0.006, the null hypothesis is rejected and so at least one co-
efficient is not zero (we have shown in Figure 9.11 that the non-zero coefficients
are β1 and β2 ).
Example 9.5
Using the fitted regression model 9.9, predict the house price for a property with
living area x1 = 4000, number of bedrooms x2 = 3 and number of bathrooms
x3 = 2.5.
Solution
The predicted value of house price (b
y ) is given by
Warning
Do not predict the response for values of the explanatory variables outside the range
of observed values as the linear relationship of the data may not be valid beyond the
observed data.
9.2. Multiple Regression Model 9.23
The value of R2 increases as the number of explanatory variables in the model in-
creases but models with too many explanatory variables are difficult to understand
and interpret. The adjusted
R2 allows us to compare multiple regression models with
different numbers of explanatory variables.
If you have different models for the same response variable with different numbers
of explanatory variables, then the one with the largest value of Radj 2 is preferred.
Normally the explanatory variables with insignificant coefficients should be removed
from the model.
Example 9.6
Find the value of R2 and Radj
2 and comment on those values.
Solution
For the house price data, the Regression procedure in spss produces the model
9.24 Module 9. Linear Regression Analysis
summary as in Figure 9.16. Here the value of R2 = 0.861, that is, 86.1% of the
variation in house prices is accounted for by the three regressors or explanatory
variables in the model. As R2 does not adjust for the number of regressors in
the model, we would always use the adjusted R2 instead.
From the spss output, Radj2 = 0.792, so, we conclude that 79.2% of the variation
in house prices is accounted for by the explanatory variables living area, number
of bedrooms and number of bathrooms in the house.
Exercise 9.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 23.13, 23.15,
23.17, 23.19, 24.3, 24.5, 24.7, 24.9.
You should be aware that linear statistical models play a major role in the study of
Statistics and that, if you proceed to the course in STA3301 Statistical Models, the
ideas introduced in this module will be expanded upon and models involving more
parameters and different assumptions about the error term will be considered.
9.4. Tutorial Module 9 9.25
The maintenance cost (in dollar) and distance driven (in hundred km) of 10 vehicles
are given in the following table.
1 2 3 4 5 6 7 8 9 10
Maintenance Cost 456 828 500 489 387 553 531 560 475 474
Distance Travel 112 173 160 127 137 124 153 143 112 123
d) Enter the data into spss and run the Regression procedure. Use the spss output
to answer the following questions.
i) Predict the value of the maintenance cost if a vehicle drove 14300km distance.
l) Find the 95% prediction interval for a the maintenance cost of a single vehicle if
it drove 14300km.
n) Check the assumptions of the simple regression model with reference to the residual
plot.
o) Check the near normal distribution of the response variable.
p) State the value of the coefficient of determination.
q) Interpret the value of the coefficient of determination.
r) Comment on how ‘good’ the distance travelled is in predicting the maintenance
cost.
The data set [Link] contains information on the grade point average (GPA)
along with SATMath, SATVerbal, HSMath and HSEnglish scores of a random sample
of 30 students. In the context of a multiple regression analysis, run the Regression
procedure in spss and use the output to answer the following questions. Include
relevant spss output in your answers.
a) Write the equation of an appropriate multiple regression model for the data.
b) State all the parameters of the model.
c) State the estimated regression parameters.
d) Write the equation of the fitted model.
e) Test if the regression is significant.
f) Interpret the estimated coefficient of SATMath.
g) Find the estimate of the error variance σ 2 .
h) Test the significance of the regression parameter of SATMath at α = 5%;
i) Calculate and save the unstandardised residuals and report the first three values.
j) Check the assumptions of the multiple regression model using the residual plot.
k) Check the near normal distribution of the response variable.
l) Interpret the value of the adjusted coefficient of determination.
m) Comment on how ‘good’ the predictors are in predicting the GPA of students.
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
9.5. Module 9 Checklist 9.27
formulate a simple linear regression model and write the linear model equation;
find the best fitted simple linear regression model for a given data set;
find and interpret the standard error of the estimators of simple regression
parameters;
find a confidence interval of the mean response for given values of explanatory
variable(s);
find prediction interval for a single response for given values of explanatory
variable(s);
Contents
10.1 Sample size, P -values and Effect size . . . . . . . . . . . . . . . . 10.3
10.2 P-hacking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10.6
10.3 Type I and II errors and their relationship to Power . . . . . . 10.9
10.4 Parametric and Non-parametric tests . . . . . . . . . . . . . . . 10.10
10.5 Choosing the best test to use . . . . . . . . . . . . . . . . . . . . 10.11
10.6 Tutorial Module 10 . . . . . . . . . . . . . . . . . . . . . . . . . . 10.15
10.7 Module 10 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 10.17
10.1. Sample size, P -values and Effect size 10.3
Introduction
In each of the weekly modules in this course we have worked through material related
to specific topics or statistical methods. Each Module built on the content covered
in previous modules, however the focus was always on the specific module material.
In this module we will expand on some important statistical concepts and will also
consider how to evaluate scenarios that may require deeper synthesis of materials
from across different modules. The majority of the content for this module is covered
in this study book.
In the formulae of each test statistic for each of the various statistical tests you should
recognise that the sample size (n) is included in the calculations. For example, n might
be used in the calculation of the mean and standard deviation or in the standard error.
Sample size is also a key component of determining the P -value through the use of
degrees of freedom. So, how would we expect a P -value to change as sample size
changes?
Generally, the larger a sample size, the more likely a study will find a relationship
or difference statistically significant if the relationship or difference actually
exists. We will come back to the bold part of the previous sentence in a moment,
but first, we will give an example to show the relationship between samples size and
P -values.
Example 10.1
If we have two samples from the same population with the same sample size
and standard deviation but different means such that:
Does this relationship between sample size and small P -values mean that as long as
we have large samples that statistical tests will always give a significant result, even
when there is no real difference or relationship?
Earlier in this section we mentioned that generally larger sample sizes lead to signif-
icant results (small P -values) if the relationship or difference actually exists.
This bold section is important. If the null hypothesis (H0 ) is true and no difference
between the means actually exists then no increase to the sample size will produce a
significant result. The relationship only holds if H0 is false (and if the assumptions
of the test have been met). This is important because we do not want to incorrectly
reject a true H0 (Type I error).
However, this relationship between sample size and P -values does mean that as sample
size increases a statistical test can detect smaller and smaller differences between
means or subtler relationships between variables. Again, this is not a flaw in the
methods; we want the tests and P -values to work this way. Another example:
10.1. Sample size, P -values and Effect size 10.5
Example 10.2
In the first example, when the difference between the means was 2cm and the
sample size was 50 per group, the test was significant with P -value < 0.05. If
the difference between means was only 0.5cm such that:
the sample size would need to be approximately 800 per group for the test to
detect a difference that small (ony 0.5) and be significant (P -value < 0.05).
This is a good and cautious approach; small differences should require more
supporting evidence (that is larger sample sizes) before we reject the H0 .
But what if a difference of 0.5 cm is not practically or contextually meaningful?
The difference between statistical and practical significance is why P -values
should always be treated as just one step in the final interpretation of any
analysis – they are not the only thing that should be considered and analyses
should always be interpreted in context of the initial question and hypotheses,
and the data.
Remember that statistical significance just means that the P -value was less
than some predefined alpha (α) level. The P -value is not an indication of the
importance of the result. Small unimportant differences, in a practical sense,
can be statistically significant. Practical significance refers to the magnitude
of the effect that has been measured, or the effect size. An understanding of
the smallest effect size that would have some practical meaning is important
when interpreting any statistical analysis and often we need to refer to previous
research or literature to help determine this.
It is important to remember that sample size is not the only thing that can influence
P -values in hypothesis tests. In either of the examples above, if we had also altered
the SD, the measure of variability around the means of each of our groups, this
also would have influenced the calculation of the test statistic and the subsequent
P -value. Think about how SD contributes to the calculation of the test statistic and
what would happen if it were larger or smaller.
Objectivity does not mean the removal of uncertainty. There is always some un-
certainty because we rely on samples to estimate statistics about our populations of
10.6 Module 10. Synthesis and Consolidation
10.2 P-hacking
P-hacking is a term used to describe attempts to find a significant result when pre-
ceding analyses were not significant. P-hacking can be in the form of increasing the
sample size, trying different methods of analysis, trying different subsets of the data
or adjusting the measurement scale of the data etc. As shown in Figure 10.1, each of
these approaches (as well as additional ones) are legitimate approaches to data prepa-
ration and analysis, but not if they are done only because prior approaches have not
provided a statistically significant result.
If we try varying the data and the analysis repeatedly only because we have not
achieved a statistically significant result, then we are inserting bias into the process.
If these are legitimate variations for a particular data set, then the variations should
be part of the initial analysis plan. If analysis is stopped once a significant result is
found then we never explore the effect of these variations, which results in selective
reporting of only significant results. This violates one of the ethical tenets described
in Module 2.
Perform careful data cleaning and descriptive analysis before starting any in-
ferential analysis to ensure that you have objectively considered the potential
influence of outliers or the scale of the variables etc.
Perform all checks to ensure test assumptions are met.
Where analysis of data begins without clear research questions or hypotheses,
or new questions arise during the analysis process, recognise that this form of
analysis is Hypothesis Generating rather than Hypothesis Testing. We should
ensure that this is clear in any reporting and label results as preliminary (need-
ing follow-up studies to add weight to the conclusions). The same sample data
set should not be used to both generate and test a hypothesis as the data is
biased towards the hypothesis.
Apply a correction where multiple testing or multiple comparisons of sub-groups
of the original data set is needed. Correction for testing multiple comparisons
may even be needed when comparison of a series of subgroups within the data
was the purpose of the original analysis plan. The xkcd comic in Figure 10.2 is a
good example of using multiple tests in search of significant results. We will not
cover the different approaches to correcting for multiple testing in this course.
However if you are curious google ‘Bonferroni correction’ for one example.
10.2. P-hacking 10.7
Figure 10.2: Multiple testing, Hypothesizing after the result is known, statistical
and practical significance, claiming causality. (Reprinted from [Link]
under the CC BY-NC 2.5 license.)
10.3. Type I and II errors and their relationship to Power 10.9
Figure 10.2 above is a good example of several analysis ‘problems’. Multiple testing
has been mentioned in the fourth point above. It is also an example of point 3: hy-
pothesizing after the result is known – claiming or giving the impression in reporting
that the significant result was the original hypothesis when in fact the process of find-
ing the significant outcome was part of a hypothesis generating process. In addition,
the final claim of the research seems to be drawing a causal link between green jelly
beans and acne. Think back to Module 2. What did you learn about our ability to
claim causal relationships? What type of study design would have been needed to
make this claim? We will come back to these questions in the Tutorial 10 questions
below.
De Veaux, Velleman & Bock, (5th edition) gives a good review of this topic in Chapter
19 – read from the section titled ‘Errors’ to the end of the chapter.
The relationship between Power and Type I and II errors means that minimising one
can lead to a trade off with the others. In reality the only parameter we can overtly
control is the significance level, α which is the probability of incorrectly rejecting a
true H0 (Type I error). The standard accepted α is 0.05 (5%); that is, a 5% chance
of incorrectly rejecting H0 . We discussed earlier in this Module that increasing the
sample size will reduce the P -value which effectively also reduces α. However, as we
reduce the probability of a Type I error we increase the probability of a Type II error.
In words, as we reduce the chance of incorrectly rejecting H0 we make it harder to
reject H0 in general. So, if H0 is in fact false we have increased the chance that we
will fail to reject it.
Referring to Figure 19.2 in De Veaux, Velleman & Bock, (5th edition) we can see that
if we move the vertical line to the right to reduce α (probability of Type I error), β
(probability of Type II error) increases. As Power is defined as 1 − β, as α decreases
so does Power.
The relationship between Power and Type I and II errors can be a mind-bender and
when first learning these concepts it is easy to go around in circles a bit. Don’t worry,
we have all been there.
10.10 Module 10. Synthesis and Consolidation
Where the assumptions are met the parametric method is a more powerful test and
the preferred approach. However, when using inferential methods we prefer to always
err on the side of caution. In practice this means that we prefer that a non-significant
conclusion be drawn rather than incorrectly concluding a significant result. From a
statistical perspective this means that it is often ‘worse’ to make a Type I error than
a Type II error and so where assumptions are not met or are uncertain, the use of
a non-parametric test is preferable. As non-parametric tests have fewer assumptions
they place a larger ‘burden of proof’ on the data to demonstrate a clear pattern before
the test will indicate a significant result, so the use of these methods still minimises
the chance of a Type I error.
Figure 10.3 gives some examples of non-parametric tests that could be used when
a parametric test is not the best option. Other than the Chi-square tests already
mentioned, we will not cover any of these other non-parametric tests any further in
this course. Note that not all parametric tests have a non-parametric equivalent and
there is no non-parametric equivalent of confidence intervals.
10.5. Choosing the best test to use 10.11
To determine the appropriate statistical methods that you should use, you will need
to consider the following:
(i) What types of data are the variables you wish to analyse? Are they categorical
(nominal), ordinal or quantitatve (ratio/scale)?
(ii) What are your dependent and independent variables?
10.12 Module 10. Synthesis and Consolidation
Figure 10.4 gives an idea of the broad level decisions that need to be made. In
practice, inferential methods should always be preceded by some descriptive methods
and generally both numeric and graphical descriptive methods are needed. Let’s look
at an example.
Example 10.3
Scenario: What is the difference (if any) in the average annual income of
accountants and real estate agents?
Although we don’t have a lot of information, we know enough about types of
studies, types of variables and different statistical methods to map out a quite
detailed approach to the analysis.
(i) The interest seems to be in comparing one group to another without any
treatment or intervention being considered. Therefore this data is proba-
bly collected from an observational study rather than an experiment.
(ii) Annual income would probably be measured on a continuous scale in
dollars or thousands of dollars. The sample data would need to include
values for one group (of accountants) and values for the other group (of real
estate agents). These would be different groups of people so each person
would belong to only one occupation and would only be ‘measured’ once.
Therefore, the data would be independent and not paired.
10.5. Choosing the best test to use 10.13
(iii) The scenario indicates that the interest is in the difference between av-
erage annual incomes. From the tests we have covered in this course we
could use a t-test or a confidence interval to consider differences between
means. We have already determined that the occupation groups are not
paired so we could consider two- independent-samples t-test with annual
income as the dependent variable and occupation type (accountant or real
estate agent) as the independent variable.
(iv) Before performing the t-test analysis, descriptive numerical summaries of
the mean and standard deviation of income for accountants and real estate
agents separately should be calculated. Histograms of the distribution of
yearly income for each group should also be considered to help identify
any outliers and evaluate the assumptions of the relevant t-test.
(v) If there are concerns about the assumption of normality not being met
then a non-parametric equivalent test could be considered.
An additional helpful resource is this YouTube video - it is less than 10min long.
Choosing which statistical test to use:
[Link]
10.14 Module 10. Synthesis and Consolidation
3. What is P-hacking?
Exercise 10.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 19.9, 19.25,
19.28, 19.30, 19.31.
If you require further explanation related to these questions, please ask in the Forum
on the StudyDesk.
10.6. Tutorial Module 10 10.15
Discuss this statement: In order to show that there is a statistically significant differ-
ence between my control group and my treatment group (to prove that the treatment
is effective), I just need to make sure my sample size is big enough.
Describe your understanding of the ethical concerns related to P-hacking and the
reporting of statistical results. You may need to refer back to the Supplementary
Material in Module 2.
Explain how Power and the significance level (α) are related.
For each of the following scenarios describe what descriptive and inferential analysis
(from those covered in this course) you think would be most appropriate and why?
Make sure you mention what type you think the mentioned variables are, what you
think are the dependent and independent variables and whether you think a non-
parametric method should also be considered.
(a) What is the difference (if any) in the average annual income between two business
types, dairy farmers (n=15) and banana growers (n=10)?
(b) What linear relationship (if any) exists between the angle of inclination of a beach
retaining wall and the amount of sand deposited along the beach, measured at
20 predetermined positions along the beach.
(c) Predict the concentration of nitrous oxides in smoke plumes from the height above
the smoke stack.
(d) What is the strength and direction of the linear relationship between the weight
of children under 6 years of age and their height?
10.16 Module 10. Synthesis and Consolidation
(e) What is the difference (if any) between the average scores of students on two
pieces of assessment within a statistics course?
(f) What is the relationship between voters party affiliation (Labour, LNP, Greens)
and their opinion on whether Australia should be a republic (Yes, No)?
(g) A sample of 20 cards has been drawn from the deck of cards. Was each suit
equally likely to be drawn?
You need to describe to a colleague how a particular statistical test works but you
must do it within 100 words. Give your best description of each of the following with
as much important information included as possible.
Scenario: Data was collected to determine if revenue for small businesses can be
predicted from the advertising spending (both variables measured in dollars).
Describe any problems you can identify in each of the following approaches to address
this scenario. Give detailed answers.
(a) A bar graph was used to compare the total advertising spending for each business
and then a paired samples t-test was used to see if there was a difference between
average revenue and average advertising spending.
(b) A histogram for advertising spending and then another histogram for revenue were
used to explore the data and remove outliers. A correlation was then performed
to determine the strength of the relationship between advertising and revenue.
(c) A histogram for advertising spending and then another histogram for revenue were
used to explore the data and then a two-independent-samples t-test was used to
see if there was a difference between average revenue and average spending on
advertising.
(d) A scatterplot was used to explore the relationship between the spending and
revenue variables. A confidence interval using two-independent-samples method
was then performed.
10.7. Module 10 Checklist 10.17
(e) A scatterplot was used to explore the relationship between the spending and
revenue variables and then regression analysis was used to model changes in
advertising spending (dependent variable) from changes in revenue (independent
variable).
The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
state what a P -value actually represents and what factors influence its value;
describe the difference between statistical and practical significance and why
this is important in research;
state the implications of Type I and Type II errors on the Power of a test;
define the difference between a non-parametric and parametric test and when
you should use each;