0% found this document useful (0 votes)
9 views248 pages

Statistics for Quantitative Research

This document is a study book for quantitative researchers at the University of Southern Queensland, covering various statistical concepts and methodologies. It includes modules on exploring data, using normal models, statistical inference, and regression analysis, among others. The content is structured with objectives, tutorials, and checklists to facilitate learning and application of statistics.

Uploaded by

jo.lamarca.13
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views248 pages

Statistics for Quantitative Research

This document is a study book for quantitative researchers at the University of Southern Queensland, covering various statistical concepts and methodologies. It includes modules on exploring data, using normal models, statistical inference, and regression analysis, among others. The content is structured with objectives, tutorials, and checklists to facilitate learning and application of statistics.

Uploaded by

jo.lamarca.13
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistics for

Quantitative
Researchers
School of Mathematics, Physics and Computing

Faculty of Health, Engineering and Sciences


University of Southern Queensland

Study Book
ii

©University of Southern Queensland, 2023

Published by

University of Southern Queensland


Toowoomba Qld 4350
Australia

[Link]

Copyrighted materials reproduced herein are used under the provisions of the Copyright Act
1968 as amended, or as a result of application to the copyright owner.

No part of this publication may be reproduced, stored in a retrieval system or transmitted


in any form or by any means electronic, mechanical, photocopying, recording or otherwise
without prior permission.

Produced using LATEX in the USQ style.


Table of Contents

1 Exploring and Understanding Data 1.1

1.1 What is/are Statistics? . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.4

1.2 About Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.6

1.3 Categorical Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.8

1.4 Quantitative Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.9

1.5 Working with spss – data entry . . . . . . . . . . . . . . . . . . . . . . 1.10

1.6 Describing Distributions with Numbers . . . . . . . . . . . . . . . . . . 1.13

1.7 Comparing distributions . . . . . . . . . . . . . . . . . . . . . . . . . . 1.16

1.8 Using Technology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.18

1.9 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.20

1.10 Tutorial Module 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.21

1.11 Module 1 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.23

2 Using the Normal Model 2.1

2.1 Ethics and Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3

2.2 The Normal Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.6

2.3 Using the Normal Model . . . . . . . . . . . . . . . . . . . . . . . . . . 2.10

iii
iv Table of Contents

2.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.13

2.5 Tutorial Module 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.14

2.6 Module 2 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.16

3 Exploring Relationships Between Quantitative Variables 3.1

3.1 Scatterplots, Association, and Correlation . . . . . . . . . . . . . . . . 3.4

3.2 Simple Linear Regression . . . . . . . . . . . . . . . . . . . . . . . . . 3.6

3.3 Summary of Correlation and Regression . . . . . . . . . . . . . . . . . 3.8

3.4 Cautions and Precautions . . . . . . . . . . . . . . . . . . . . . . . . . 3.11

3.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.13

3.6 Tutorial Module 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.14

3.7 Module 3 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.18

4 Gathering Data 4.1

4.1 Understanding Randomness . . . . . . . . . . . . . . . . . . . . . . . . 4.4

4.2 Sample Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.4

4.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.6

4.4 Using Random Numbers to Select a Sample . . . . . . . . . . . . . . . 4.8

4.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.12

4.6 Tutorial Module 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.14

4.7 Module 4 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.19

5 Statistical Inference for One Mean 5.1

5.1 Sampling Distribution of a Sample Mean . . . . . . . . . . . . . . . . . 5.6

5.2 Statistical Inference for a Mean . . . . . . . . . . . . . . . . . . . . . . 5.16

5.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.29

5.4 Tutorial Module 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.30

5.5 Module 5 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.32


Table of Contents v

6 Statistical Inference for Comparing Means 6.1

6.1 Inference for Paired Samples . . . . . . . . . . . . . . . . . . . . . . . . 6.4

6.2 Independent Groups . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.11

6.3 Parametric and Nonparametric Tests . . . . . . . . . . . . . . . . . . . 6.18

6.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.20

6.5 Tutorial Module 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.21

6.6 Module 6 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.23

7 Analysis of Variance (ANOVA) 7.1

7.1 One-way Analysis of Variance (ANOVA) . . . . . . . . . . . . . . . . . 7.3

7.2 One-way ANOVA and spss . . . . . . . . . . . . . . . . . . . . . . . . 7.9

7.3 Two-way Factorial ANOVA . . . . . . . . . . . . . . . . . . . . . . . . 7.15

7.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.22

7.5 Tutorial Module 7 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.24

7.6 Module 7 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.28

8 Association between Categorical Variables 8.1

8.1 Contingency Table . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.4

8.2 Associated or Not? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.7

8.3 The Chi-Square Test of Independence . . . . . . . . . . . . . . . . . . 8.10

8.4 The Chi-Square Goodness-of-Fit Test . . . . . . . . . . . . . . . . . . 8.23

8.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.26

8.6 Tutorial Module 8 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.28

8.7 Module 8 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.34


vi Table of Contents

9 Linear Regression Analysis 9.1

9.1 Inference for Simple Linear Regression . . . . . . . . . . . . . . . . . . 9.4

9.2 Multiple Regression Model . . . . . . . . . . . . . . . . . . . . . . . . . 9.15

9.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.24

9.4 Tutorial Module 9 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.25

9.5 Module 9 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.27

10 Synthesis and Consolidation 10.1

10.1 Sample size, P -values and Effect size . . . . . . . . . . . . . . . . . . . 10.3

10.2 P-hacking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10.6

10.3 Type I and II errors and their relationship to Power . . . . . . . . . . 10.9

10.4 Parametric and Non-parametric tests . . . . . . . . . . . . . . . . . . . 10.10

10.5 Choosing the best test to use . . . . . . . . . . . . . . . . . . . . . . . 10.11

10.6 Tutorial Module 10 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10.15

10.7 Module 10 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10.17


Module 1
Exploring and
Understanding
Data
1.2 Module 1. Exploring and Understanding Data

Contents
1.1 What is/are Statistics? . . . . . . . . . . . . . . . . . . . . . . . . 1.4
1.1.1 Population and Sample Datasets . . . . . . . . . . . . . . . . . . . 1.5
1.2 About Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.6
1.3 Categorical Data . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.8
1.4 Quantitative Data . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.9
1.5 Working with spss – data entry . . . . . . . . . . . . . . . . . . . 1.10
1.5.1 Working with spss – raw data . . . . . . . . . . . . . . . . . . . . 1.12
1.6 Describing Distributions with Numbers . . . . . . . . . . . . . . 1.13
1.7 Comparing distributions . . . . . . . . . . . . . . . . . . . . . . . 1.16
1.8 Using Technology . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.18
1.8.1 A Note about Graphs . . . . . . . . . . . . . . . . . . . . . . . . . 1.19
1.9 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.20
1.10 Tutorial Module 1 . . . . . . . . . . . . . . . . . . . . . . . . . . . 1.21
1.11 Module 1 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 1.23
1.3

Module Objectives
On successful completion of this module students should be able to:

ˆ identify the cases and variables in a set of data;

ˆ distinguish between categorical, ordinal and quantitative variables;

ˆ use spss to enter data, open a data set and label variables and values;

ˆ use spss to construct frequency and relative frequency tables;

ˆ interpret frequency and relative frequency tables;

ˆ use spss to construct a bar chart and pie chart display of a categorical variable;

ˆ use spss to produce a histogram of a quantitative variable;

ˆ describe a histogram in words in terms of shape, centre, spread and unusual


features;

ˆ identify outliers by eye in a quantitative distribution;

ˆ use spss to calculate the five number summary of a data set, including the
median, lower and upper quartiles, and interquartile range (iqr) and interpret
these measures;

ˆ use spss to construct boxplots;

ˆ interpret distributions displayed in boxplots;

ˆ use spss to calculate the mean and standard deviation of a set of data; and

ˆ recognise the impact of outliers and skewness on measures of centre and spread.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
1.4 Module 1. Exploring and Understanding Data

Introduction
This first module of the Statistics for Quantitative Researchers Course covers material
in Chapters 1, 2, and 4 of De Veaux, Velleman & Bock, (5th edition)). It also
introduces the statistical software spss. Note that material in Chapter 3 is covered
in Module 8.

Install spss on your computer or have access to it via [Link]


[Link]/ before working through this module.

In general terms the module covers the relevance and importance of Statistics as
a discipline and describes types of data. The basic ideas and definitions in this
module are used throughout the course. The module also covers the description of
distributions of variables by showing how data can be summarised using graphs and
numbers.

1.1 What is/are Statistics?


Statistics has a reputation for being challenging and is not necessarily a subject of
choice for some students. However statistics is used regularly to help us gain an
understanding of all of the information which is available to us on a daily basis. The
root of this information is data - ‘any collection of numbers, characters, images, or
other items that provide information about something’ (De Veaux, Velleman & Bock,
(5th edition), Chapter 1). The need to understand this data is what drives Statistics.

This course will help you develop the skills to become proficient at interpreting and un-
derstanding data; and more importantly communicating this information and knowl-
edge to others. De Veaux, Velleman & Bock (5th edition, Chapter 1) gives some
examples of how Statistics is important in everyday life.

Throughout this Study Book there are references to readings and exercises from the
prescribed textbook (De Veaux, Velleman & Bock, (5th edition)). These provide
additional explanations of content and practice exercises.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 1 (Section 1)

The brief introduction to Statistics contained in this reading can be expanded on


considerably, so be sure to refer to the lecture videos and lecture handout on the
Course StudyDesk and complete the tutorial worksheet at the end of this module.
1.1. What is/are Statistics? 1.5

After each prescribed reading throughout the Study Book, some questions are asked
related to the reading and are designed to highlight some of the important points in
the reading. Even if you don’t formally write down the answers to these questions,
at least check the answers off in your mind and read back through the material if
necessary. We suggest the points raised might be used as the basis for your own notes
about the subject. Answers to these questions are not provided but, where necessary,
can be requested by posting to the appropriate forum on the Course StudyDesk.

ˆ What is Statistics?

ˆ What are statistics?

ˆ Complete the sentence: Statistics is about . . . .

ˆ Why is Statistics important?

1.1.1 Population and Sample Datasets


Data on certain variables is central to any statistical analysis. Any given or observed
dataset is either population data or sample data.

ˆ Population: In Statistics, a population is the collection of all individuals or


items or objects of interest. If a dataset contains information or ‘values’ on all
members of a population then it is population data. This kind of dataset comes
from a census or 100% count of the members of a population. Any numerical
characteristic of a population is called a parameter. For example, the mean of
a population dataset is a parameter and denoted by the Greek letter µ (read as
‘meuw’). Population datasets are very rare in real life, and hence usually the
values of the parameters are unknown. Details on these are covered in Module
4.

ˆ Sample: A sample is a subset of any population. To make the sample repre-


sentative, items are selected from the population by a random sampling method.
You will learn more about random sampling in Module 4 and other later Mod-
ules. Any characteristic of a sample is called a statistic. Essentially any value
calculated from a sample dataset is a statistic. For example, the mean of a
sample dataset is a statistic and is denoted by y (read as y-bar).

ˆ Estimation: Often statistics from the sample are used to estimate the pop-
ulation parameters. For example, the value of a sample mean y (is a known
number) and is an estimate of the unknown population mean µ. In real life,
most of the datasets are from samples.
1.6 Module 1. Exploring and Understanding Data

1.2 About Data


Data does not have to only consist of numbers; it can be represented by characters
such as names or other labels. Data values, however are useless without context.
Journalists establish the ‘Five W’s’: Who, What, When, Where and (if possible)
Why. Often How is added as well. By answering these questions, as Journalists do,
the context for the data values is established. Thus data is defined as systematically
recorded information in context and can be in the form of numbers or labels.

As stated earlier, one aspect to better understanding of statistics is familiarity with


statistical language. Some terminology when collecting and recording data are listed
below:

ˆ Cases – in answering the Who question when collecting data we are defining
the ‘cases’ of the data set – the individuals or objects that the data describes.
If information is being collected on people then these people are the cases; if the
information is recorded for trees then the trees are the cases. Often the cases
that we describe in a dataset are a subset (sample) from some larger group of
interest (population).

ˆ Variables – the information or characteristics recorded for each individual or


object are called the variables. If we record the height, age and gender for a
group of individuals then the people are the cases and height, age and gender
are the variables. Variables can be quantitative, categorical or ordinal:

– Quantitative – the variable in question records the quantity of what is


measured, and generally has associated units; a variable for which it makes
sense to do arithmetic, such as find an average. For example, height (cm),
age (years) or the number of books an individual owns.
– Categorical – the variable in question records to what group or category
an individual belongs and does not have an associated unit. For example,
mode of study (with values oncampus or online), smoking status (smoker
or non-smoker) or program of study (science, business, arts, etc.).
– Ordinal – An ordinal variable is a categorical variable for which the order
of categories has meaning. For example, survey questions which ask for a
response from a likert scale such as: 1 = strongly disagree; 2 = disagree;
3 = neutral; 4 = agree; 5 = strongly agree.
– Identifiers - sometimes each case has a unique identifer or Case ID (for
example, a university ID). While such variables in some sense are categor-
ical, each individual or case has a unique value for the identifier and so
analyzing them is irrelevant.
1.2. About Data 1.7

The word distribution recurs throughout the course. Every time we see the word
we should recall it as a description of all the values a variable can take and how often
those values occur.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 1 (Sections
2 & 3).

After this reading write down answers to the following questions:

ˆ What are data?

ˆ What is a case? Describe at least two types of cases.

ˆ What is a variable? What is a value?

ˆ What are units?

ˆ What is a categorical variable? Give an example.

ˆ What is a quantitative variable? Give an example.

ˆ What is an ordinal variable? Give an example.

ˆ What is a data table?

Remember that a variable takes on values which are variable; i.e., they can vary from
one case to the next. Don’t confuse the variable with its values. ‘Program of Study’
is a variable; ‘business’ and ‘science’ are two of its values. When asked to describe
a variable, make sure your definition is precise. For example, ‘ice-cream’ is not a
variable, ‘flavour of ice-cream’ is.

Exercise 1.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 1.11, 1.13.

Note: 1.11 means De Veaux, Velleman & Bock, (5th edition) Chapter 1,
Exercise 11 at the end of the chapter.

Answers to the odd-numbered exercises are in the back of the textbook. Teaching
staff are able to check answers to even-numbered exercises where necessary. Post
questions of this type to an appropriate forum on the StudyDesk.
1.8 Module 1. Exploring and Understanding Data

1.3 Categorical Data

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2 (Section 1).

A variable is categorical if it describes some specific attribute or characteristic of


individuals that can only be categorised, e.g., health status (with two ‘values’ dis-
eased and healthy), or smoking habit (with three ‘values’ never smoked, past smoker,
current smoker).

Data for a single categorical variable can be represented by a graph such as a bar
chart or pie chart.

Data on two categorical variables can be represented by a contingency table, also


known as a two-way cross-table. Contingency tables display the relationship between
two categorical variables. The entries in the cells of a contingency table are the counts
or frequencies of the combination of categories of the two variables.

Example: To determine the effectiveness of a drug (serum) to treat arthritis, 400


arthritic patients were divided into two equal groups. Group 1 received the treatment
(serum) and the other group received a placebo (a dummy treatment – a treatment
that looks like the serum but without the active ingredient). In the treatment group,
117 patients improved and in the placebo group 74 patients improved. In this exam-
ple, Intervention is a categorical variable with two ‘values’ – treatment and placebo.
Also, Condition of patient is the other categorical variable with two ‘values’ – im-
proved and unimproved.

The above count data can be represented in the following contingency table:

Condition Treatment Placebo Total


Improved 117 74 191
Unimproved 83 126 209
Total 200 200 400

In this module we have described categorical data and briefly mentioned contingency
tables. Module 8 covers contingency tables and the association between two categor-
ical variables extensively, exploring marginal and conditional distributions, and the
test of independence.
1.4. Quantitative Data 1.9

Exercise 1.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.1, 2.3, 2.5,
2.35, 2.41, 2.43.

1.4 Quantitative Data

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2 (Sections 2
& 3).

After this reading write down answers to the following questions:

ˆ Which types of graphs can be used to display the distribution of a quantitative


variable?

ˆ What is the essential difference in appearance between a bar graph and a his-
togram?

ˆ Why are bins of equal width desirable in a histogram?

ˆ Which three features are used to describe a histogram?

ˆ Which three features can be used to describe shape?

ˆ What is the mode of a histogram?

ˆ What is an outlier?

ˆ What is meant by a symmetric distribution? Give an example of an approxi-


mately symmetric distribution.

ˆ What is meant by ‘skewed to the right’ ? Give an example of a skewed to the


right distribution.

ˆ What is meant by ‘skewed to the left’ ? Give an example of a skewed to the left
distribution.

ˆ What does rounding entail? Why is it used?


1.10 Module 1. Exploring and Understanding Data

An important idea from this reading is that of a visual description of the distribution
of a quantitative variable. Shape, centre and spread are the key elements of this
description. Shape can in turn be broken down into number of modes, degree of
symmetry, and outliers or unusual features. A description based on these elements is
intended to provide sufficient information for a reader to be able to roughly sketch
the histogram of the distribution, including a scale on the horizontal axis.

With this in mind, if a visual description is required, it should be concise and stick to
the main features. However, don’t over-interpret and rely only on a visual inspection
of a histogram. In Section 1.6 more mathematically precise numerical measures of
centre and spread are considered. As regards a visual inspection though, the measures
we discuss in Section 1.6 are not relevant here.

Similar comments apply when comparing two or more distributions visually (see
De Veaux, Velleman & Bock, (5th edition), Chapter 4, Section 1). However, we
prefer to use another type of graph for this procedure as we shall see in Section 1.7
of this Study Book.

In closing this section it should be emphasised as regards histograms that:

ˆ there are no gaps between the bars of a histogram (unless there is a class of
zero frequency);

ˆ the histogram bins should be of equal width; and

ˆ the bin width (or number of bins) is chosen to most clearly reveal the overall
shape of the distribution.

Exercise 1.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.33, 2.45, 2.47.

1.5 Working with SPSS – data entry


Is spss installed yet? When spss opens, two tabs appear at the bottom of the spss
window, one called ‘Data View’, the other called ‘Variable View’. Click on ‘Data
View’ if it’s not already showing. The spreadsheet that appears is designed to be
filled with data in data table format; i.e., each column containing values of a variable
and each row describing a case. Data can be typed directly into the cells of this
spreadsheet. For example, type in data so that the ‘Data View’ is as in Figure 1.1.
The empty cells appear to have dots in them. This is spss’s way of denoting missing
values. All we have to do is leave the cell empty. Notice that data input into spss
1.5. Working with spss – data entry 1.11

Figure 1.1: spss Data View after entering data.

Figure 1.2: spss Variable View after entering variable information.

defaults to two decimal places even for data values that are whole numbers. This can
be adjusted in the ‘Variable View’.

Now click on ‘Variable View’ and enter the information as shown in Figure 1.2. Not
all the information you need is visible on this figure. The values of the variable
‘gender’ are 1 for male and 2 for female; for ‘faculty’, the values are 1 = Business,
2 = Sciences, 3 = Other; and for ‘mode’, 1 = On campus, 2 = Online. It may help
to refer to the spss video Inputting data in SPSS on the Course StudyDesk when
entering this information.

Notice the distinction made in spss between Variable Name, Variable Label, and
Variable Values. Also notice spss should be told the type of each variable, whether it
is ‘scale’ (i.e., quantitative), ‘nominal’ (i.e., categorical) or ‘ordinal’. After completing
the ‘Variable View’, click again on the ‘Data View’ and check that it looks the same
as in Figure 1.3.

These data are part of a dataset collected in a past survey of Statistics students. The
complete dataset is made use of in the spss videos which can be found in the spss
Resources link on the Course StudyDesk.

Notice that all the data in this example are numeric, even though three of the five
variables (gender, faculty, and mode) are categorical. The values of the categorical
variables have been coded. In other words, categories such as ‘male’ have been
1.12 Module 1. Exploring and Understanding Data

Figure 1.3: spss Data View after entering variable information.

given numeric labels. This is a common and strongly recommended practice. Coding
increases the speed and accuracy of data entry and can simplify the manipulation of
data once it is entered into the computer. Certainly, coding makes the data harder to
read but spss allows us to toggle between the coded and uncoded versions at the click
of the icon called ‘Value Labels’, 2nd active icon from the right in the top toolbar.
Coding is made use of extensively in the data files used in this course.

The data can be saved as an spss data file using the ‘File/Save As’ command. The
extension on the file name that spss creates is ‘sav’. Try this on the current example—
call the file ‘[Link]’ say. If spss is now closed and the [Link] name double-
clicked in Windows Explorer, spss will automatically open and load [Link], and
all the variable information that was inserted will still be there.

1.5.1 Working with SPSS – raw data


A brief description of how to use spss to construct a bar or pie chart is given in
the Tech Support at the end of Chapter 2 of the textbook. This only applies for
the simple bar and pie charts when the data is in standard format (i.e. each row
represents a case). This type of data is sometimes called raw or ungrouped data
to distinguish it from grouped data in which the data is only given as a frequency
table. Note that the assignment data is given in a raw data format, but some text
book questions give data in a grouped data format. Refer to the spss videos to see
how to investigate the relationship between two categorical variables when in raw
data format using ‘Contingency Tables’ and in grouped data format using ‘Weight
Cases’ and then ‘Contingency Tables’. This will be covered in detail in Module 8.
1.6. Describing Distributions with Numbers 1.13

Exercise 1.4
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Starting with spss

ˆ Inputting data in spss

ˆ Importing data into spss

ˆ Saving output in spss

ˆ Bar Graph

ˆ Pie Graph

ˆ Stacked Bar Graph

ˆ Contingency Tables

ˆ Weight Cases

1.6 Describing Distributions with Numbers

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 2, Sections 4
& 5.

After this reading you should be able to write down the answers to the following
questions:

ˆ Which are the two most commonly used measures of centre?

ˆ What is the problem with midrange as a measure of centre?

ˆ What is the median of a distribution?

ˆ Describe the steps in finding the median of a set of numbers.

ˆ What is a quartile?
1.14 Module 1. Exploring and Understanding Data

ˆ State three ways of measuring spread.

ˆ How is range defined?

ˆ What is another name for the median?

ˆ How are the first and third quartiles determined?

ˆ What does iqr stand for? What is the formula for the iqr?

ˆ What is the range of a set of numbers? Why is range not commonly used as a
measure of spread?

ˆ What is the formula for the mean? Interpret this formula in words.

ˆ How can the mean of a distribution be interpreted?

ˆ What symbol is used to denote a mean?

ˆ What statistics comprise the five-number summary?

ˆ Sketch a typical boxplot, pinpointing the positions of the five-number summary


measures.

ˆ Which symbol is used to denote standard deviation? To denote variance?

ˆ How is the standard deviation found from the variance?

ˆ When do the mean and median of a distribution agree closely?

ˆ The standard deviation should be used in conjunction with which measure of


centre?

ˆ When is the standard deviation zero? When is it positive? When is it negative?

ˆ If the units of measurement of the observations is metres, what are the units of
measurement of each of the mean, median, standard deviation and variance?

ˆ Which descriptions of centre and spread are best for a skewed distribution?
Which measures are best for a reasonably symmetric distribution?

This reading concerns the calculation, interpretation and usage of statistics (plural);
that is, numbers calculated from data. Commonly the types of statistics described in
this chapter are called summary statistics in that they provide a concise numerical
summary of the distribution of a quantitative variable.

To enhance understanding and interpretation of these summary statistics it is helpful


to be familiar with the formulae for the mean, variance, standard deviation, median,
lower quartile, upper quartile, iqr and range of a set of data. These formulae are in
the textbook but for convenience we repeat them here.
1.6. Describing Distributions with Numbers 1.15

P
y
mean = y = (1.1)
n

(y − y)2
P
2
variance = s = (1.2)
n−1


rP
(y − y)2
standard deviation = s = s2 = (1.3)
n−1

Note: it is useful to know how to calculate y and s on your calculator in statistics


mode, although you will primarily be using spss to calculate all of the summary
statistics. You won’t be asked to do these calculations by hand using the formulae.
n+1
median = m = value in the position of ordered data (1.4)
2

lower quartile = Q1 = median of lower half of ordered data (1.5)

upper quartile = Q3 = median of upper half of ordered data (1.6)

inter quartile range, IQR = Q3 − Q1 (1.7)

range = max − min (1.8)

Notice we have used the notation Q1 and Q3 for lower and upper quartiles. Sometimes
these are called instead the first and third quartiles—then the median is the same as
the second quartile.

The distribution of data for a single variable is often conveniently summarised in


terms of two statistics, a measure of centre and a measure of spread. The commonly
used pairs of measures are the mean and standard deviation or the median and
iqr. The mean should not be used in conjunction with the iqr or the median with
the standard deviation. The mean, standard deviation pairing is preferred to the
median, iqr pairing primarily because they turn out to be the most amenable to
further analysis as we shall see in later modules. However, unless the distribution
of interest is reasonably symmetric the mean, standard deviation pairing may be
misleading measures of centre and spread. In the presence of moderate or severe
skewness or outliers, the median and iqr are often the better indicators of centre and
spread respectively.
1.16 Module 1. Exploring and Understanding Data

The text mostly uses y to describe the mean of a sample. There’s no reason why x
couldn’t be used, or t or u for that matter. Using y or x are just conventions and it’s
best to stick to them unless there’s a good reason not to. In Module 6, for example,
we deal with differences between pairs of numbers and it makes sense to call the mean
of the differences d. Just make sure you make it clear to the marker in any of your
assignment answers what your notation represents.

Statistics (singular) is about variation. In many ways then measures of spread such
as iqr (which measures the range of the central 50% of a set of data) and standard
deviation (which has no easy interpretation except as the square root of variance
which in turn is approximately the mean squared deviation) are more important
than measures of centre such as median and mean. Certainly it should be clear by
now that it would be quite misleading to report the centre of a distribution only,
without mention of its spread.

The word ‘average’ is in common usage of course. From a statistical point of view
however it’s preferable to be more specific and use ‘mean’ or ‘median’. Otherwise
we could be accused of being misleading knowing that the mean and median can be
considerably different because of skewness or outliers.

It may help to remember median as a measure of the middle of a distribution by


noting that both the words meDian and miDDle contain a middle D. In the same
vein, mOde can be remembered as concerning mOst cOmmOn. For mean, you’re on
your own!

Exercise 1.5
Do De Veaux, Velleman & Bock, (5th edition), exercises 2.15, 2.17, 2.21,
2.49, 2.51, 2.55, 2.77, 2.83 (use spss for calculations)

1.7 Comparing distributions

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 4, Sections 1
& 2. Note Sections 3 & 4 are optional.

A boxplot is a quick and dirty way of displaying the distribution of a quantitative vari-
able compared to the more refined method described earlier; namely the histogram.
1.7. Comparing distributions 1.17

Figure 1.4: Side by side boxplots of rainfall data.

As such we tend not to use it to display the distribution of a quantitative variable.


However, if interest is in comparing two or more distributions (typically when we
have a number of groups involved, all having been measured on the one variable),
side-by-side boxplots, like those displayed on the next page (Figure 1.4) are appro-
priate. The reason for preferring this method of comparison over that given at the
start of Chapter 4 of the text (i.e., comparing two histograms), is that boxplots focus
more directly on basic shape, centre and spread than a set of histograms, which can
be more difficult to interpret in these terms.

Example 1.1
An experiment was carried out in Tasmania, Australia between 1964 and 1971
to ascertain the effectiveness of cloud-seeding to promote rain (Miller, A.J. et
al (1979) ‘Analyzing the results of a cloud-seeding experiment in Tasmania’,
Communications in Statistics—Theory & Methods A8(10),1017–1047). Some
of the data from this experiment are displayed in Figure 1.4. Compare in fewer
than 75 words words, the distributions of rainfall from the seeded and unseeded
clouds.
Solution
Both distributions are right skewed. The seeded rainfall distribution has a lower
centre but a larger spread than the unseeded rainfall distribution. The median
of the seeded rainfalls is about 3.5 cm and of the unseeded rainfalls about
4.5 cm. The interquartile range is about 6 cm for the seeded and about 4 cm
for the unseeded rainfalls. Both distributions have one or two high outliers.
1.18 Module 1. Exploring and Understanding Data

Notice that what we’re doing here in general terms is describing a relationship between
a categorical variable and a quantitative variable. In this example the categorical
variable is ‘whether or not seeding occurred’ and the quantitative variable is ‘amount
of rainfall’.

Check out also the spss video on boxplots (on the StudyDesk – ‘Graphing in spss’
then boxplot).

Exercise 1.6
Do De Veaux, Velleman & Bock, (5th edition), exercises 4.15, 4.25, 4.41
(use spss), 4.49.

1.8 Using Technology


In the tutorial exercises and assignments you are expected to use spss to calculate
summary statistics for the given datasets, namely the mean and standard deviation,
median, quartiles and iqr, and graphing. However, scientific calculators may be used
for calculations with small datasets, if you prefer. Also any units associated with an
answer should be included or marks may be docked in assignments.

Calculators and computers may give answers with 8 or 10 significant figures. That’s
fine part way through a calculation, but make note of the rule that says you should
not round numbers until the final answer is reached. Common sense dictates the
degree of rounding of the final answer. Statistics is not an exact science. We are
often estimating values based on limited or approximate information or using an
approximate model to describe the real world. As such, statistics like means and
standard deviations can only be expected to be slightly more precise than the original
data and answers based on the normal distribution are approximations to the truth.
The larger the quantity of data, the relatively more precise statistics can sensibly be.
Exact rules about this exist but are not needed in this course because a fair amount
of latitude is given. Just demonstrate commonsense and round numbers correctly.

Calculators distinguish the sample standard deviation (probably indicated by σn−1 or


xσn−1 or s on your calculator) from the population standard deviation (σn or xσn or
σ on your calculator). Incidentally, we won’t be using this last button in this course
so you can ‘blank’ it out on your calculator. A similar distinction between sample
mean (y) and population mean (µ) is not needed on a calculator because the same
formula is used for both. Use the correct terminology though when writing down the
value of a mean, whether it be a µ or a y.
1.8. Using Technology 1.19

spss is available at all times during the course. It should be used to draw graphs,
produce tables, and calculate statistics. Before getting carried away with including
computer output in assignment answers however, read the instructions given in the
assignments carefully. spss output often contains more than we need to answer a
particular question. In assignments you need to demonstrate that you can select
relevent information from the spss output to answer the question that was posed.
Determination of relevant summary statistics is quite straightforward in spss and is
discussed in the spss video Quantitative Descriptive Statistics located in the spss
Resources link on the StudyDesk. Note that a major aim of this course is to gain
familiarity with statistical software.

1.8.1 A Note about Graphs


An important point not really emphasised so far is that, as far as possible, a graph
or table should be self-contained. In other words a viewer should be able to tell, with
as little effort as is reasonable, what is the nature of the information that the writer
intends to convey. Hence we should include in all our tables and graphs:

ˆ A contextual title above a table or below a figure;

ˆ Meaningful labels (rather than codes or uncommon abbreviations);

ˆ Units as appropriate; and

ˆ Number of individuals or cases (e.g., n = 150), if not otherwise apparent.

As regards the last item, the point is that sample size gives the reader some idea of how
reliable the results shown in the graph are. For example, if relative frequencies rather
than frequencies are displayed and the reader sees a figure such as 50%, knowing that
50% is based on 200 individuals is far more compelling than knowing it’s based on
just 10.
1.20 Module 1. Exploring and Understanding Data

Exercise 1.7
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Quantitative Descriptive Statistics

ˆ Selecting Cases

ˆ Boxplot

ˆ Histogram

1.9 Closing Comments


The terms and concepts covered in this module recur throughout the course. Getting
familiar with them will therefore put you in good stead for later work.

Make sure that you keep moving forward and try to keep ahead. There’s a lot of
material in this course. It pays to keep up even at the expense of things which may
still be a mystery. They will make more sense after doing later work.

The secret of success in mastering the course is problem solving. It’s not enough
just to read the material through and look at examples. Ultimately you need to be
problem solvers.

A minimum set of problems from the text have been prescribed. Do as many as you
are able in the time available, concentrating on areas which you feel less comfortable
with. Another source of practice questions is the tutorial worksheets. These have an
added benefit in that the worked solutions to these worksheets (which will be made
available on the StudyDesk) will give you guidance on how to express your answers
in assignments. There is no substitute for tackling problems!

The terms, concepts and graphs covered in this module recur throughout the course.
Again, getting familiar with them will therefore put us in good stead for later work.
1.10. Tutorial Module 1 1.21

1.10 Tutorial Module 1


The following section contains Tutorial 1 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1:
The following represent the results (out of 100) for a class of 35 statistics students on
their final exams and the symbol (B or S) next to each student result indicates the
student’s program of study (B – Business Studies, S – Applied Sciences):

85(B), 38(B), 92(S), 93(S), 100(B), 88(S), 78(B), 79(S), 96(B), 81(B), 62(S), 69(S),
77(B), 78(S), 83(S), 91(S), 90(S), 85(S), 99(S), 97(S), 43(B), 100(S), 90(S), 73(B),
84(B), 85(B), 93(S), 63(B), 94(S), 94(S), 85(B), 74(B), 92(S), 86(B), 90(S)

For this data complete the following questions:

(a) Enter the data into spss and create a histogram of results for the whole class.

(b) Referring to the graph only, describe the distribution in terms of shape, centre
and spread (and outliers).

(c) What measures of centre and spread would you use to describe this distribution?
Why?

(d) Using spss,

(i) Calculate the mean of the data set.


(ii) Calculate the standard deviation of the data set.
(iii) Calculate the median of the data set.
(iv) Calculate the IQR of the data set.

(e) What is the five number summary of this data set?

(f) Why are the mean and median different/similar?

Question 2:
Refer to the relevant data in Question 1 to complete the following questions:

(a) Using spss, create an appropriate graph to compare the distribution of class
results for Business Studies students with that of Applied Sciences students.

(b) Referring to your graph only, comment on the differences in these two distribu-
tions in relation to shape, centre, spread (and outliers).
1.22 Module 1. Exploring and Understanding Data

(c) What measures of centre and spread would you use to describe each of these
distributions? Why?

Question 3:
The statistics class was also asked about their usual mode of transport to campus.
From this class of 35 students, 12 indicated that they walked to campus, 5 rode a
bicycle, 14 came by car and 4 came on the bus.

For this data complete the following questions:

(a) Create a bar graph for this data using spss (note that you will need to use Weight
Cases because the data is already grouped).

(b) Describe what this graph tells you.

(c) Use another type of graph to display this data using spss.

(d) Describe what this graph tells you that is different from the previous graph.

Question 4:
For the histogram of body mass index (BMI) in Figure 1.5 answer the following
questions:

Figure 1.5: Histogram of Body Mass Index data.

(a) What is the shape of the distribution of BMI?


1.11. Module 1 Checklist 1.23

(b) From visual inspection, state the centre of the distribution.

(c) From visual inspection, state the spread of the distribution.

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

1.11 Module 1 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ identify the cases/individuals and variables in any data set;

ˆ classify a variable as categorical or quantitative;

ˆ identify the units in which a quantitative variable is measured;

ˆ identify a variable as being categorical or ordinal or quantitative;

ˆ produce bar graphs and pie graphs for categorical data;

ˆ choose an appropriate display for quantitative data;

ˆ summarise and describe the distribution of a categorical variable with a fre-


quency table or relative frequency table;

ˆ visually describe the distribution of a quantitative variable in terms of shape,


centre and spread from a histogram;

ˆ recognize the impact of outliers and skewness on measures of centre and spread;

ˆ produce and interpret a five-number summary;

ˆ compare groups using boxplots;

ˆ calculate summary statistics using spss.


Module 2
Using the
Normal Model
2.2 Module 2. Using the Normal Model

Contents
2.1 Ethics and Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3
2.2 The Normal Model . . . . . . . . . . . . . . . . . . . . . . . . . . 2.6
2.3 Using the Normal Model . . . . . . . . . . . . . . . . . . . . . . . 2.10
2.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.13
2.5 Tutorial Module 2 . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.14
2.6 Module 2 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 2.16
2.1. Ethics and Statistics 2.3

Module Objectives
On successful completion of this module students should be able to:

ˆ standardise the value of a variable;

ˆ use standardised values to make comparisons;

ˆ understand what is meant by a normal model;

ˆ sketch a normal curve; and

ˆ make appropriate use of standard normal tables in problem solving.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
This module introduces guidelines around ethical practice in statistics. These ethical
considerations are designed to help guard against the propogation of misinformation.

In addition, this module covers material in Chapter 5 of De Veaux, Velleman &


Bock, (5th edition). The notion of modelling data using a normal distribution is
introduced. The basic concepts of modelling data discussed in this module will be
revisited a number of times in the second half of this course, so now is the time to
put in a bit of extra effort as it will pay dividends later in the semester.

2.1 Ethics and Statistics


Whether we are consumers of statistical information or the person responsible for pro-
ducing statistical analyses and results, standards of ethical practice help us navigate
potential pitfalls which could lead to the spread of misinformation.
2.4 Module 2. Using the Normal Model

The Committee on Professional Ethics of the American Statistical Association (ASA)


published their Ethical Guidelines for Statistical Practice in 2018. [Link]
[Link]/ASA/Your-Career/Ethical-Guidelines-for-Statistical-Practice.
aspx. These guidelines are intended to help users of statistics make ethical decisions
and inform those who rely on statistical analysis of the standards they should expect.

The ASA states:


‘Good statistical practice is fundamentally based on transparent assumptions, repro-
ducible results, and valid interpretations. In some situations, guideline principles may
conflict, requiring individuals to prioritize principles according to context.’

Discussion around ethical use of data and/or research often focuses on important
aspects such as:

ˆ Reproducibility and open data;

ˆ Security of data storage and data transmission;

ˆ Ethical treatment of human and animal participants (this includes confidential-


ity and informed consent in the area of human research);

ˆ Intellectual property.

There are also key ethical considerations relating to:

ˆ Quality of data, where the data may have been compromised by a conflict
of interest held by those collecting, analysing, reporting or commissioning the
reporting of the data and/or resulting research outcomes. For example:

– Where known problems with the collection method that could distort in-
terpretation are intentionally not reported.
– Where details of data collection methodology are not adequately described
allowing for later misuse or misinterpretation of the true nature of the data.
This might not be done to intentionally misinform; however, simple lack
of clarity and incomplete disclosure can lead to ethical issues.

ˆ Use of statistical methods appropriate to the data and questions that hope to
be answered:

– Verification of assumptions;
– Appropriate use of statistical methods if more than one is available;
– Appropriate use of visualisation (more than one type of graph) if more
than one approach is available.
2.1. Ethics and Statistics 2.5

ˆ Fair reporting of results:

– Sample size and its effect on statistical significance (we will cover this in
more detail in Module 10);
– Reporting significance without contextualising the meaning and impor-
tance of the effect size (e.g. the size of the difference between means or the
strength of the relationships between variables. More details in Module
10);
– Applying multiple hypothesis tests to subsets of the data and selectively
reporting only the significant results (we will cover this in more detail in
Module 10);
– Over-interpreting results e.g. inferring cause-and-effect based on an obser-
vational study, discussing results in the context of a population for which
the sample is non-representative

Even if your future work does not require you to produce statistical analyses and
results, you may need to rely on statistics produced by others. As consumers of sta-
tistical information we rarely have access to the original data to check if the analyses
and claims are true. Ethical guidelines help us identify the information that should
be available to us in order to trust the claims made.

When considering statistical claims in research publications, the ethical statistician:

ˆ Acknowledges statistical and substantive assumptions made in the execution


and interpretation of any analysis. When reporting on the validity of data used,
acknowledges data editing procedures, including any imputation and missing
data mechanisms.

ˆ Reports the limitations of statistical inference and possible sources of error.

ˆ In publications, reports, or testimony, identifies who is responsible for the sta-


tistical work if it would not otherwise be apparent.

ˆ Reports the sources and assessed adequacy of the data, accounts for all data
considered in a study, and explains the sample(s) actually used.

ˆ Clearly and fully reports the steps taken to preserve data integrity and valid
results.

ˆ Where appropriate, addresses potential confounding variables not included in


the study.

ˆ In publications and reports, conveys the findings in ways that are both honest
and meaningful to the user/reader. This includes tables, models, and graphics.
2.6 Module 2. Using the Normal Model

ˆ In publications or testimony, identifies the ultimate financial sponsor of the


study, the stated purpose, and the intended use of the study results.

ˆ When reporting analyses of volunteer data or other data that may not be repre-
sentative of a defined population, includes appropriate disclaimers and, if used,
appropriate weighting.

ˆ To aid peer review and replication, shares the data used in the analyses when-
ever possible/allowable and exercises due caution to protect proprietary and
confidential data, including all data that might inappropriately reveal respon-
dent identities.

ˆ Strives to promptly correct any errors discovered while producing the final re-
port or after publication. As appropriate, disseminates the correction publicly
or to others relying on the results.

2.2 The Normal Model

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 5.

After this reading write down answers to the following questions:

ˆ What is meant by standardising (or should that be standardizing)?

ˆ How is data standardised?

ˆ What does a standardised score of −2.0 indicate?

ˆ What notation is commonly used for standardised data?

ˆ If an observation is measured in metres, what are the units of the standardised


observation?

ˆ What is the purpose of standardising?

ˆ What impact does adding a constant to every data value have on measures of
centre? What about measures of spread?

ˆ What impact does multiplying or dividing every data value by a constant have
on measures of centre? What about measures of spread?
2.2. The Normal Model 2.7

ˆ What impact does standardising data in a distribution have on shape, centre


and spread?

ˆ Describe the overall shape of a normal curve.

ˆ Which two values are needed to specify a particular normal distribution? What
notation is used?

ˆ Why do we need new notation for the mean and new notation for the standard
deviation?

ˆ What is a parameter?

ˆ What is the distribution of a variable with a normal distribution after it has


been standardised?

ˆ What is the mean of the standard normal distribution? What is the standard
deviation of the standard normal distribution?

ˆ How is the standard normal distribution denoted?

ˆ In the normal model, what percentage of values fall within one standard devia-
tion of the mean? Within two standard deviations of the mean? Within three
standard deviations of the mean?

ˆ What happens to a normal distribution as µ changes?

ˆ What happens to a normal distribution as σ changes?

ˆ What rule describes a common property of all normal distributions?

ˆ What does this rule mean?

ˆ Describe the notation used to specify a normal distribution.

ˆ What do areas under normal curves represent?

ˆ What does Table Z provide?

Standardising provides a way of comparing values from two different distributions.


In the previous module we compared two (or more) distributions using boxplots.
Notice though, these distributions all involved the one variable. We were really just
comparing one variable across a number of groups. With z-scores we can compare
‘apples’ and ‘oranges’. This is a very useful device.

There are two formulas that we will need to use quite often. The first one is called
the standardising formula because it turns the value y of a quantitative variable into
a z-score or standard score:
2.8 Module 2. Using the Normal Model

Figure 2.1: How to sketch a normal curve.

y−µ
z= (2.1)
σ
The second is called the unstandardising formula because it turns a z-score back into
a y value:
y =µ+z×σ (2.2)

One equation is simply a rearrangement of the other.

Standardising and the normal distribution are two different concepts. Keep them sep-
arate in your mind. Apart from being used as a means of comparison, standardising
is also useful as a way of enabling us to make use of the normal model for describing
many variables of interest. The z-score in this situation is really just a means to an
end, not an end in itself.

By convention, everyday Roman letters are used to describe the summary measures
of actual data and Greek letters are used to describe summary measures of model
distributions such as the normal distribution. For example, the model distribution
counterparts of y, s and s2 are µ, σ and σ 2 .

spss is not designed to easily find percentages associated with the normal model. It
can be done but is hardly worth the effort when Normal Tables such as Table Z in
the back of the text are readily available. Thoroughly check through the step-by-step
examples of using the normal model to solve problems in De Veaux, Velleman & Bock
(5th edition, Chapter 5, Section 4 – Working with Normal Percentiles).

It’s always a good idea to sketch the normal distribution if a normal model
is being used to help answer a problem. By shading in an area under the curve,
remembering the total area under the curve is one (or 100%), we can estimate what
the answer should be before doing the calculation. A reasonably accurate sketch is
needed though. Figure 2.1 shows how to produce one. Assuming a mean µ = 60 and
standard deviation σ = 15, we’ve marked off points on a horizontal axis at 45, 60 and
75 representing µ − σ, µ and µ + σ. Three crosses are located as shown directly above
the three points and a convex upwards curve put through these crosses as shown with
convex downwards curves extending out each tail.
2.2. The Normal Model 2.9

Many aspects of using Statistics involve judgement. We have already been asked,
for example, to assess symmetry, recognise outliers, choose appropriate measures of
centre and spread, and round-off answers appropriately. In this section we are asked
to assess whether data can be modelled using the normal distribution. In other words,
given a set of values of a variable, does it appear reasonable to assume that if we had
access to all possible values of that variable, the distribution of all those values would
have a normal distribution. Strictly speaking we could safely say no because we might
argue that nothing in real-life has a perfectly normal distribution. But insisting on
exactness isn’t going to help us solve real-life problems. We need to be prepared to
accept that a model is an idealised or theoretical description of reality. If it comes
reasonably close to the reality it will be useful in providing approximate answers to
relevant questions. The Normal Probability Plot described in De Veaux, Velleman &
Bock (5th edition, Chapter 5, Section 5), available in spss as a P-P Plot, provides
a suitable visual means of checking the reasonableness of using a normal model in a
real situation if some data is available. spss also has a Q-Q plot which essentially
does the same job as the P-P plot. Sometimes, of course, no data is available, and
we need to appeal to other information or our’s or other people’s experience to make
a judgement.

The decision as to whether data are normally distributed or not can have important
implications as to how we proceed with further analysis.

To assess normality we would initially construct a histogram and look for features
which would suggest data are not normally distributed, before proceeding to a normal
probability plot. Those features in the histogram would include things like skewness
or lack of symmetry, outliers and absence of a central peak.

Remember, what we are assessing is normality of the population of data from which
our sample of data is selected. A histogram of the sample itself will never look
perfectly normal even when the population from which it is drawn is normal. Natural
variation will cause anomalies in the shape of the graph. When a sample is small (say
less than 40), the sample distribution may be quite misshaped but still be consistent
with a normal population. For larger samples we would of course expect a shape
closer to normality.

There is considerable subjectivity involved in assessing normality using just the ap-
pearance of a histogram. While more sophisticated methods do exist, there is no
way to know if the population is normally distributed for sure. For small samples
especially, the most we can often say is that the sample is consistent with a normal
population (but what is unsaid is that the sample is consistent with other non-normal
population distributions also!).

On a more positive note, as we shall see in Module 5 and beyond, the issue of normality
in choosing a method of analysis generally decreases in importance as sample size
increases. Often, even with samples as small as 15, only approximate normality is
2.10 Module 2. Using the Normal Model

required and beyond 30, only gross departures from normality are worth worrying
about.

This information is not of much use for very small samples. Often in practice we know
something more about the distribution of the data under consideration than just the
data itself either from experience in similar contexts or from theoretical consideration
of the mechanism that causes the data to be variable.

Two important categories of data which are known to be nearly normal are firstly,
measurements of physical stature (height, weight and so on) of many organisms (an-
imals, plants, etc), and secondly, errors in making measurements. For example, the
heights of human adult males are well approximated by a normal distribution. Also
the errors in these measurements of height will tend to be nearly normal. Both these
statements can be justified from both empirical evidence (experience based on data)
and from theoretical considerations.

2.3 Using the Normal Model


There is no simple formula to give the areas under normal curves. We must be able
to find such areas using Table Z from the Statistical Tables link on the StudyDesk or
in De Veaux, Velleman & Bock.

Example 2.1
Assuming the height of adult humans is normally distributed with mean of
172 cm and standard deviation of 5 cm, what proportion of adults are between
166 cm and 180 cm tall?
Sketch a normal curve and locate the points 166 and 180 on the horizontal axis
(see Figure 2.2). The area shaded under the curve between these two points
equals the proportion required.
2.3. Using the Normal Model 2.11

Figure 2.2: Area under the normal curve for Example 2.1.

Construct a z axis as shown below the x (height) axis. Note the mean of
the distribution 172 corresponds to z = 0 and the points 172 − 5 = 167 and
172 + 5 = 177 correspond to z = −1 and z = 1 respectively.
Using formula (2.1), the z-scores for the endpoints of interest, 166 and 180, are
respectively (166 − 172)/5 = −1.2 and (180 − 172)/5 = 1.6.
From Table Z, the area to the left of z = −1.2 is 0.1151 and the area to the left
of z = 1.6 is 0.9452. The shaded area is therefore 0.9452 − 0.1151 = 0.8301.
Hence the proportion of adults between 166 cm and 180 cm tall is approximately
83%.
(Note: these probabilities can also be checked using spss; see the spss video
Normal Probabilities under the spss Resources link on the StudyDesk).

Example 2.2
Using the information as given in Example 2.1, what would be the height of a
doorway that requires 5% of adults to stoop in order to pass through it?
Since only the tallest 5% of adults are forced to stoop, the value (height)
required, x, is at the upper end of the distribution as shown in Figure 2.3. The
area to the left of x is 1 − 0.05 = 0.95.
2.12 Module 2. Using the Normal Model

Figure 2.3: Area under the normal curve for Example 2.2.

From Table Z, the entry closest to 0.95 in the body of the table is 0.9505. This
is the entry corresponding to z = 1.65. (Yes, z = 1.64 would be equally valid.)
This value is positive and greater than one as expected for a height at the
upper end of the distribution.
Using formula (2.2) for unstandardising, we have x = 172 + (1.65)(5) = 180.2.
An adult has to be at least 180 cm tall before being forced to stoop.

Exercise 2.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 5.1, 5.3, 5.11,
5.17, 5.25, 5.43 (use spss to construct a histogram and a P-P plot), 5.53,
5.55.

Exercise 2.2
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Normal Probabilities

ˆ P-P plot
2.4. Closing Comments 2.13

2.4 Closing Comments


The concepts covered in this module recur in a variety of contexts in the material
covered in the second half of the course. The effort you put in now to understand these
concepts well will put you in good stead for later work. A tip that has been found to
be invaluable: a diagram is vital when doing problems involving the normal model.
It helps you to organise the information in a problem and determine more clearly
what is required.

Note that a Quick Review and additional exercises for the first two modules are given
after Chapter 5 in the text book (after the exercises).
2.14 Module 2. Using the Normal Model

2.5 Tutorial Module 2


The following section contains Tutorial 2 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1: Using Standard Normal Tables


Draw a picture of the standard Normal curve. Mark in the position and values of the
mean and standard deviation.

Now use your diagram of the Normal distribution and Table Z to answer the following
questions:

(a) What percentage of a standard Normal model is found in each region?

(i) z > 1.5


(ii) z < 2.3
(iii) z > −1.72
(iv) z < −0.56
(v) −2 < z < 1.34
(vi) 2.3 < z < 3.1

(b) In a standard Normal model, what value (s) of z cut(s) off the region described?

(i) the highest 10%


(ii) the highest 60%
(iii) the lowest 2%
(iv) the lowest 40%
(v) the middle 95%
(vi) the middle 50%

Question 2: How do the males in Western Europe stack up?


Based on a Normal model with mean 179 cm and standard deviation of 7 cm describ-
ing the heights of males in a particular Western European country:

(a) What proportion of males will be taller than 184 cm?

(b) Below what height will the shortest 10% of the male population of this country
be?

(c) What percentage of males from this country are between 175 and 185 cm tall?
2.5. Tutorial Module 2 2.15

(d) Calculate the interquartile range of heights of males from this country.

(e) Above what height are the tallest 5% of the males in this country?

Question 3: Do they get what they pay for?


Management at a cereal company is concerned that the filling machine at their factory
may need some further adjustment to obtain a better balance between the number
of satisfied customers (they are getting the amount of cereal that they have paid
for based on the advertised weight) and minimizing overfilling (loss of profit). The
filling machine currently fills boxes based on a Normal model with mean 720 grams
and standard deviation of 30 grams describing the weights of boxes of cereal of a
particular variety.

(a) If boxes of this cereal are labelled as containing 700 grams, what proportion of
customers are going to be dissatisfied with their purchase?

(b) Below what weight will the lightest 5% of boxes of this cereal be?

(c) What percentage of boxes of this cereal are between 700 grams and 750 grams?

(d) Above what weight are the heaviest 10% of boxes of this cereal?

Question 4: Comparing two groups based on summary statistics


Who should get the prize? Is there much chance that anyone would have scored a
better result than the prizewinner?

The local school gives one prize for Social Studies at the end of the school year. Two
subjects are studied in the Social Studies discipline – history and geography. George
studied history and obtained a result of 86 marks out of 100 overall for the year.
Ringo studied geography and obtained a result of 89 marks out of 100 overall for the
year.

(a) Before Speech Night some students in the class thought that Ringo was a certainty
to get the Social Studies prize.
Explain the students’ reasoning. Do you agree?

(b) The Social Studies HOD (Head of Department) studied statistics at university
and thought that she had a fair solution. She knew that the distribution of marks
in the history class had a mean of 63 with a standard deviation of 12 and the
distribution of marks in the geography class had a mean of 60 with a standard
deviation of 16.
Explain how she could use this information to arrive at a fair solution.

(c) Applying the teacher’s procedure, determine who got the Social Studies Prize.
2.16 Module 2. Using the Normal Model

(d) What is the probability that anyone would have scored better than the Social Sci-
ences prize winner (if the results on the examinations followed a Normal model)?

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

2.6 Module 2 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ standardise the value of a variable;

ˆ use standardized values to make comparisons;

ˆ understand what is meant by a normal model;

ˆ state the difference between µ and y; and σ and s;

ˆ use Table Z to find proportions in problems associated with a normal distribu-


tion;

ˆ use Table Z in reverse to find values in problems associated with a normal


distribution.
Module 3
Exploring
Relationships
Between
Quantitative
Variables
3.2 Module 3. Exploring Relationships Between Quantitative Variables

Contents
3.1 Scatterplots, Association, and Correlation . . . . . . . . . . . . 3.4
3.2 Simple Linear Regression . . . . . . . . . . . . . . . . . . . . . . . 3.6
3.3 Summary of Correlation and Regression . . . . . . . . . . . . . 3.8
3.4 Cautions and Precautions . . . . . . . . . . . . . . . . . . . . . . 3.11
3.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.13
3.6 Tutorial Module 3 . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.14
3.7 Module 3 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 3.18
3.3

Module Objectives
On successful completion of this module students should be able to:

ˆ use spss to produce a scatterplot where appropriate;

ˆ distinguish between the explanatory and the response variable where appropri-
ate;

ˆ describe a scatterplot in terms of form, direction, scatter and unusual features;

ˆ calculate using spss and interpret a correlation coefficient;

ˆ understand when the use of correlation is appropriate;

ˆ understand what is meant by a linear model;

ˆ use spss to determine a regression line;

ˆ plot a regression line on a scatterplot;

ˆ interpret the slope and intercept of a regression line;

ˆ interpret R2 , the square of the correlation coefficient;

ˆ calculate and interpret residuals;

ˆ list and check the conditions required for using a regression line;

ˆ appreciate the limitations of regression analysis, including the impact of lurking


variables, extrapolation, subsets, and outliers; and

ˆ appreciate that association does not necessarily imply causation.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
3.4 Module 3. Exploring Relationships Between Quantitative Variables

Introduction
Modules 1 and 2 mainly concentrated on describing data for just one variable. We’re
seldom however interested in just one variable in isolation. Life is not one-dimensional.
Life is multidimensional. Many variables interact with each other in complex ways. To
help understand these interactions requires exploration of the relationships amongst
variables.

The relationship between two categorical variables (as in Chapter 3 of De Veaux,


Velleman & Bock, (5th edition)) is covered in Module 8. The linear relationship
between two quantitative variables is covered here in Module 3 as simple linear re-
gression. In Module 9 this is extended to include inference for linear regression and
also multiple regression.

Content in Module 3 is covered by Chapters 6, 7 & 8 of De Veaux, Velleman & Bock,


(5th edition).

Much data is multivariate; that is the values of two or more variables are collected
for each of a number of people, countries, industries, or whatever. The discipline of
Statistics largely concerns methods of exploring, describing and measuring the type
and strength of relationships amongst variables. Module 1 dealt with methods of
graphing or calculating summary statistics of individual variables (as appropriate)
and introduced some useful ways to summarise the location, spread and shape of the
distributions of such variables. However, the reason for measuring or observing more
than one variable in a given situation is usually because a relationship or association is
suspected to exist amongst some or all of them. Our interest in this module is in using
appropriate means to detect and measure that relationship. We restrict attention to
pairs of variables not only for simplicity but also because graphs and tables don’t
easily lend themselves to dealing with more than two variables simultaneously. Later
modules in this course extend some of these concepts to more than two variables.

3.1 Scatterplots, Association, and Correlation

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 6; the sec-
tions on Measuring Trend: Kendall’s Tau, Nonparametric Association:
Spearman’s Rho and Straightening Scatterplots is optional.

Sorry about the American examples. We need an Australian version of this text!
After this reading write down answers to the following questions:
3.1. Scatterplots, Association, and Correlation 3.5

ˆ What is the purpose of a scatterplot?

ˆ What type of variables are displayed in a scatterplot?

ˆ In describing scatterplots, what three features should be looked at? What else
might also be looked for?

ˆ Define the respective meanings of positive and negative direction.

ˆ What is meant by a linear form?

ˆ What are the generic names commonly used for the horizontal and vertical
axes?

ˆ What is the difference between an explanatory variable and a response variable?


Give an example.

ˆ Which variable should be graphed on the horizontal axis? Which variable should
be graphed on the vertical axis?

ˆ What does correlation measure?

ˆ What is the formula for the correlation r?

ˆ What variable types are involved in correlation?

ˆ What are the units of correlation?

ˆ What range of values can the correlation coefficient take on?

ˆ How does r indicate a positive association? A negative association?

ˆ What does a scatterplot look like if r = 1? If r = −1?

ˆ What happens to r if the units in which the variables are measured are changed?
Why is this?

ˆ What happens to r if the variables on the two axes are swapped around?

ˆ In what way can a graph be misleading in displaying the level of correlation?

ˆ What effect do outliers have on correlation?

ˆ What is a lurking variable?

The full title of the correlation discussed here is Pearson’s correlation coefficient, also
called Pearson’s product moment.

Notice that the eXplanatory variable goes on the X-axis. Also that X comes before
Y in the alphabet and similarly we assume the explanatory x comes before or drives
the response y.
3.6 Module 3. Exploring Relationships Between Quantitative Variables

It is not always clear in a problem which is the explanatory and which is the response
variable. For example, in dealing with weights and heights of human beings, the
role of the two variables is unclear without more information being provided. We
know that r is the same regardless of which variable is put on the x-axis so if all
we’re interested in is determining correlation it doesn’t matter. However, in the next
section on regression we take things further and look at predicting one variable from
the other. We find then that the variable being predicted is always the y or response
variable. Consequently, if a problem involves correlation, look further at the questions
and ascertain if regression is involved. If so, information will exist to determine the
explanatory and response variables.

Exercise 3.1
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Scatterplot

ˆ Regression and Correlation

Exercise 3.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 6.3, 6.7, 6.11,
6.13, 6.19, 6.43 (a) (Reminder: use spss for graphs and calculations).

3.2 Simple Linear Regression

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 7. (Read
through the Math Box for understanding ONLY. You are not required
to reproduce this).

After this reading write down answers to the following questions:

ˆ What is a linear model?


3.2. Simple Linear Regression 3.7

ˆ What is the equation of a line that passes through the origin?

ˆ What purpose does a regression line have?

ˆ What is the equation of a regression line in terms of z-scores?

ˆ What is the equation of a regression line in terms of x and y.

ˆ What does a ‘hat’ symbol indicate?

ˆ What symbol is used for the intercept? What about the slope?

ˆ What is the formula for the slope of the line?

ˆ What is the formula for the intercept of the line?

ˆ Why is the regression line described as ‘least-squares’ ?

ˆ Write down the general form of the regression line.

ˆ How do y and yb differ?

ˆ Give an interpretation of the slope of a regression line.

ˆ Give an interpretation of the intercept of a regression line.

ˆ What is a residual? How is it calculated?

ˆ Which symbol denotes a residual?

ˆ What is the purpose of finding the residuals?

ˆ What is the mean of the residuals?

ˆ What is the formula for the standard deviation of the residuals? What symbol
is used for this standard deviation?

ˆ What is a residual plot?

ˆ What is the ideal pattern for a residual plot?

ˆ Give an interpretation of the square of the correlation, R-squared.

ˆ What range of values can R2 have?

ˆ What can be concluded if R2 is 100%? If R2 is 0%?

ˆ In what sense is the regression line best?


3.8 Module 3. Exploring Relationships Between Quantitative Variables

Remember the first thing to do before doing any correlation or regression calculations
is to produce a scatterplot. Correlation and simple linear regression are only relevant
if the form of the scatterplot is linear.

The word ‘simple’ is not used in the text but is used in the heading of this section. It
indicates that we are interested in one explanatory variable and how it relates to one
response variable. In a later module we will deal with multiple regression, meaning
more than one explanatory or x variable is involved.

Computers are ideal for coping with the amount of calculation required in the methods
described in this module. Our task is not to get bogged down with formulae or
calculations. Those things are best left to a computer. We need to be able to judge
the suitability of the techniques here, know when to apply them, know how to get
spss to perform them and be able to interpret the results.

3.3 Summary of Correlation and Regression


In this Module, the basic concepts of regression and correlation are introduced. In-
ference for regression and multiple regression will be covered in Module 9.

Correlation Coefficient
The direction and strength of the linear relationship between two quantitative vari-
ables are measured by Pearson’s correlation coefficient. r, the sample correlation
coefficient, is a number between −1 and 1. The value of r is 1 if the associa-
tion/relationship between the two quantitative variables is a perfect positive rela-
tionship. The value of r is −1 if the association/relationship between the two quan-
titative variables is a perfect negative relationship. r = 0 implies that there is no
linear relationship between the two variables. The slope, b1 , of the regression line
and the correlation coefficient, r, share the same sign/direction (both positive or
both negative).

Coefficient of Determination, R2
The square of r, often denoted by R2 , is called the coefficient of determination which
indicates the proportion of variation in the response variable that is explained by
its relationship with the explanatory variable. Usually, it is expressed as a percent-
age, and R2 close to 100% indicates that X is a good predictor of Y , provided the
scatterplot indicates a linear relationship.

Simple Linear Regression Analysis


The linear relationship between two quantitative variables, namely the response or
dependent variable denoted by Y and the explanatory or independent (predictor)
variable denoted by X, is represented by a simple regression model
Y = β0 + β1 X + ,
3.3. Summary of Correlation and Regression 3.9

where β0 = intercept parameter, β1 = slope parameter and  is the error term.

The fitted or estimated regression line is found by using n pairs of sample/observed


data points (x, y), as
yb = b0 + b1 x,
where yb represents the predicted or estimated value of the response, b0 is the estimate
of β0 (the y-intercept of the regression line) and b1 is the estimate of β1 (the slope of
the regression line).

The above regression equation is found by using the ordinary least squares method
that yields the estimate of slope parameter β1 to be
sy
b1 = r ×
sx
and the estimate of the intercept parameter β0 as

b0 = y − b1 x,

where sy is the standard deviation of y, sx is the standard deviation of x, y is the


mean of y and x is the mean of x.

The above regression equation is used to predict a value of Y , the response variable,
for a give value of X = x, the explanatory variable.

Interpretation of the slope and intercept


The estimated slope, b1 is the rate of change in the value of Y for one unit change in
the value of X. The estimated intercept, b0 is the value of Y when X = 0.

Residuals
The difference between the observed response y and predicted response yb (for a given
value of x) is called a residual. So, the residual is defined as

e = y − yb.
P
The
P 2sum of residuals is always zero, that is, e = 0. But the sum of squared residuals,
e , is not zero, and is minimised to find the least squares estimators of the slope
and intercept.
3.10 Module 3. Exploring Relationships Between Quantitative Variables

Method of Ordinary Least Squares

Figure 3.1: Method of Ordinary Least Squares.

The line to be fitted is yb = b0 + b1 x and it is unlikely that for a given x value


the observed value of y will exactly equal the predicted value yb. However, the least
squares method ensures that the overall (squared) deviations of all the observed y 0 s
n
from the fitted y 0 s, that is, e2 = i=1 (yi − ŷi )2 is minimised.
P P

Assumptions for linear regression model


1. The two variables are quantitative

2. The relationship between the two variables is linear

3. The error term  (equivalently, response Y ) is approximately normally dis-


tributed

4. The variance of the error term is constant (for all values of X)

A residual plot (residual versus predicted response) is used to check the validity of
the assumptions of linear regression model.

Exercise 3.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 7.5, 7.27, 7.29,
7.33, 7.63, 7.69.
3.4. Cautions and Precautions 3.11

3.4 Cautions and Precautions

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 8.

After this reading write down answers to the following questions:

ˆ How might subgroups be revealed in a regression analysis?

ˆ How might nonlinearity be revealed in a regression analysis?

ˆ What is extrapolation?

ˆ Why is caution needed when extrapolating?

ˆ Describe three types of outliers.

ˆ What is leverage? What do high leverage points do?

ˆ How might the impact of outliers be judged?

ˆ How is a correlation based on averaged data likely to relate to the correlation


based on individual data?

ˆ What is a lurking variable?

ˆ Is the strength of the association a factor in deciding whether or not an associ-


ation implies causation?

ˆ Give an example of a study involving a lurking variable. What is the lurking


variable?

There is not much in the way of formulas or calculations in this reading but the
message is strong. This chapter of the textbook brings us back to reality after reading
about the elegance and power of regression in the previous chapter. It reveals traps for
the unwary (or should that be tools for the unscrupulous) and gives us a perspective
about Statistics as a science involving a large dose of commonsense, good judgement,
and a lot of work looking beyond just the application of some formula or blackbox of
some statistical software. Notice the importance of exploring the data and residuals
of the data using plots—just the sort of thing computers are good at producing.

The caution that the existence of an association between two variables (whether they
be quantitative variables as in this module or categorical variables to be covered in
Module 8) does not necessarily imply that changing one variables causes the other
3.12 Module 3. Exploring Relationships Between Quantitative Variables

one to change is so important it deserves mentioning again (and again). In the next
module we delve into what we might do about it. Here however we need to be clear
that if we are asked to explain an association on the basis of a lurking variable,
we should not just describe some potential lurking variable but also give a plausible
reason why it might be associated with both the explanatory variable and the response
variable. It’s not enough to postulate an association with just one of these variables
and not the other.

Example 3.1
Does driving with the car’s headlights on during the day reduce traffic ac-
cidents? Explain clearly (in fewer than 100 words) why data showing that
drivers who leave their lights on do have fewer accidents is not necessarily good
evidence that turning on lights causes fewer accidents.
Solution
Those drivers who turn on their lights during the day (in the belief that it
makes them more visible to other traffic on the road) may well have fewer
accidents than those who don’t because drivers who turn on their lights may
tend to be more cautious and it is the cautiousness rather than any increase
in visibility that results in fewer accidents. In other words, fewer accidents
and lights-on driving may be a common response to the overall caution of the
driver.

Exercise 3.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 8.15, 8.25, 8.29,
8.35, 8.45.

Exercise 3.5
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Scatterplot

ˆ Regression and Correlation

ˆ Residual Plot
3.5. Closing Comments 3.13

3.5 Closing Comments


We are definitiely not expected to do correlation and regression calculations by hand
but we need to be able to instruct the computer correctly and interpret the output.
Graphs are an essential part of correlation and regression. In fact we should never do
a regression or correlation without first graphing the data. Unfortunately that ‘rule’
gets broken all too often in statistical practice in ‘real-life’.

We’ve looked at how to display the relationship between a categorical and quantitative
variable in Module 1 (see boxplots in Section 1.7). The relationship between two
categorical variables is covered in Module 8 (see contingency tables in Section 8.1).
In this module we’ve seen that scatterplots are appropriate to display the relationship
between two quantitative variables. A summary of these results is shown in Table 3.1.

Categorical Quantitative

Categorical Contingency Table Side-by-side boxplots

Quantitative Side-by-side boxplots Scatterplot

Table 3.1: Displaying the relationship between two variables

Just a comment about using the correct language in dealing with relationships. The
word ‘correlation’ is used all over the place in real-life. Strictly it should only be
used to describe a linear relationship between two quantitative variables. The word
‘association’ is more general and can be used to describe a relationship where one
exists between any two variables, categorical or quantitative.

Note that a Quick Review and additional exercises for elements of Module 3 are given
after Chapter 9 in the text book (after the exercises).
3.14 Module 3. Exploring Relationships Between Quantitative Variables

3.6 Tutorial Module 3


The following section contains Tutorial 3 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1: Investigate the following research question


Can we predict the energy content of a chocolate bar from the fat content?

The following spss output was obtained for a sample of 16 chocolate bars (Figure 3.2
and Figure 3.3):

Figure 3.2: Scatterplot of Energy (in calories) versus Fat Content (in grams) for
Chocolate Bars.

(a) Identify the variables, the type of each variable and the units of measure of each
variable.

(b) Describe the scatterplot from the SPSS output in Figure 3.2 in terms of form,
direction and scatter.

(c) Do you initially think a linear model is appropriate? Why?

(d) From the spss output (Figure 3.3), what is the correlation coefficient for this
data, and explain what it means in this context.
3.6. Tutorial Module 3 3.15

Figure 3.3: SPSS output of Energy (in calories) versus Fat Content (in grams) for
Chocolate Bars.

(e) From the spss output, what is the value of the slope (b1 ) of the regression line
for this data?

(f) Give an interpretation for what the slope means in this context.

(g) From the spss output, what is the value of the intercept (b0 ) of the regression
line for this data?

(h) Give an interpretation for what the intercept means in this context.

(i) What is R2 , and explain what this means in context.

(j) Write down the regression line for predicting Energy from Fat Content (be sure
to define the explanatory variable and the response variable in your equation).

(k) I have a chocolate bar with 40 grams of fat, what would be the predicted energy
rating of this chocolate bar? (i.e., make a prediction of ‘y’ when x = 40).

(l) Do you think this is a valid prediction? Justify your answer.

(m) Using the residual plot (Figure 3.4), comment on the suitability of using a linear
regression model for this chocolate data.
3.16 Module 3. Exploring Relationships Between Quantitative Variables

Figure 3.4: Residual Plot (unstandardised residual versus Fat Content).

Question 2: Investigate the following research question


Does the age of a Corolla have an impact on how much you can ask for it when you
sell it?

To answer this research question data was collected on 15 cars (see Figure 3.5).

Figure 3.5: SPSS data on Age of Corolla (years) and Price Advertised ($)
3.6. Tutorial Module 3 3.17

(a) Identify the variables, the type of each variable and the units of measure of each
variable.

(b) Enter the data in Figure 3.5 into spss and create a scatterplot to represent the
relationship between the two variables.

(c) Describe the scatterplot from your spss output in terms of form, direction, scatter
and outliers.

(d) Do you initially think a linear model is appropriate? Why?

(e) Using spss, calculate the correlation coefficient for this data, and explain what
it means in this context.

(f) What is R2 , and explain what this means in context.

(g) Using spss conduct a regression analysis.

(h) From your spss output, state the value of the slope (b1 ) of the regression line for
this data?

(i) Give an interpretation for what the slope means in this context.

(j) From your spss output, state the value of the intercept (b0 )?

(k) Give an interpretation of what the intercept means in this context.

(l) Write down the equation of the regression line for predicting advertised price
from age (be sure to define any variables you use). Use spss to draw this line
onto your scatterplot.

(m) I have a Corolla that is 10 years old, how much do you think I would advertise
this for if I were to sell it? (i.e., make a prediction of ‘y’ when x = 10).

(n) Using spss, create a residual plot of these data and use it to comment on the
suitability of using a linear regression model for this Corolla data.

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
3.18 Module 3. Exploring Relationships Between Quantitative Variables

3.7 Module 3 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ construct and interpret a scatterplot for two quantitative variables;

ˆ describe a relationship in terms of direction, form, and scatter;

ˆ recognise when a correlation coefficient is appropriate;

ˆ calculate the correlation coefficient from bivariate data using spss;

ˆ distinguish between the explanatory and response variables where appropriate;

ˆ determine the least-squares regression line from data using spss;

ˆ plot a regression line on a scatterplot;

ˆ interpret the slope and intercept of a regression line;

ˆ give a contextual interpretation of R2 ;

ˆ calculate residuals, and plot them against x and against time order; and recog-
nise unusual patterns.
Module 4
Gathering Data
4.2 Module 4. Gathering Data

Contents
4.1 Understanding Randomness . . . . . . . . . . . . . . . . . . . . . 4.4
4.2 Sample Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.4
4.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.6
4.4 Using Random Numbers to Select a Sample . . . . . . . . . . . 4.8
4.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.12
4.6 Tutorial Module 4 . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.14
4.7 Module 4 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 4.19
4.3

Module Objectives
On successful completion of this module students should be able to:

ˆ understand the concept of randomisation and a random number;


ˆ identify the population and sampling frame in a study;
ˆ understand the concepts of random sampling and census;
ˆ distinguish amongst simple random samples, convenience samples, voluntary
response samples, stratified samples, systematic samples, cluster samples and
multistage samples in survey designs;
ˆ recognise the potential of undercoverage, nonresponse and response biases in
surveys;
ˆ recognise whether a study is observational or experimental;
ˆ recognise the existence and potential effect of confounding variables on both
observational and experimental studies;
ˆ identify the explanatory variables (factors), treatments, response variables, and
experimental units or subjects in an experiment;
ˆ outline the design of a completely randomised experiment specifying group sizes,
specific treatments, and the response variable;
ˆ explain why a randomised comparative experiment can give good evidence for
cause-and-effect relationships;
ˆ recognise the importance of a control group, placebos and the double-blind
technique;
ˆ identify blocks and understand why blocking can improve an experiment;
ˆ recognise and describe the layout of a matched pairs design; and
ˆ use random numbers to select a sample from a population, including a sim-
ple random sample (srs) and a stratified random sample when the strata are
identified; and to randomly allocate experimental units to treatment groups.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.
4.4 Module 4. Gathering Data

Introduction
We now take a step backwards. It’s all very well to describe data and relationships
between variables like we’ve been doing but where did the data come from in the first
place. Somebody, somewhere has collected it. How should it be gathered? What do
we need to be aware of before running a survey or an experiment designed to produce
data? Statistics has a lot to say about these things as we see in this module.

4.1 Understanding Randomness


This is fairly gentle stuff after some of the previous chapters and is designed to
get us thinking about the concept of randomness. Collecting data ‘at random’ is a
cornerstone of Statistics without which data intended to establish information about
the world at large would be of doubtful reliability.

Some calculators have random number generators on them, probably designated by


RAN#. If you have one, check it out. It probably generates random numbers with
three decimal places between 0.000 and 0.999 such that each of these thousand possible
numbers occurs with equal frequency in the long run. The uniform random number
generator in spss does a similar thing.

4.2 Sample Surveys

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 10.

After this reading write down answers to the following questions:

ˆ What is meant by a population?


ˆ What is meant by a sample?
ˆ What is meant by bias? Give an example.
ˆ Why is randomising important in choosing a sample from a population?
ˆ What matters in obtaining a representative sample from a population, no matter
what the size of the population.
4.2. Sample Surveys 4.5

ˆ Give two disadvantages associated with a census.

ˆ What is a statistic? What is a parameter?

ˆ What is a statistic often used for?

ˆ What symbols are used for mean, standard deviation, proportion, correlation,
intercept, and slope, when calculated from a sample?

ˆ What are the corresponding symbols for the population model?

ˆ What does srs stand for?

ˆ How is an srs defined?

ˆ What is a sampling frame?

ˆ What is sampling variability?

ˆ What is a stratified random sample? Give an example.

ˆ What is a cluster sample? Give an example.

ˆ What is a multistage sample?

ˆ What is a systematic sample?

ˆ What is a voluntary response sample? Give an example.

ˆ What is a problem with a voluntary response sample?

ˆ What is a convenience sample? Give an example.

ˆ What is undercoverage? Give an example.

ˆ What is nonresponse? Give an example.

ˆ What does bias mean?

ˆ What is nonresponse bias? What is response bias?

ˆ How might undercoverage or nonresponse effect sample survey results?

ˆ Describe two other potential sources of response bias in sample results.

ˆ What is meant by confounding? Give an example.

We prefer to think of a census, not as a sample that consists of the entire population
(as per the text definition), but as an attempt to sample the entire population. If we
stick to the text definition, the census run by the Australian Bureau of Statistics every
4.6 Module 4. Gathering Data

five years would not really be a census because some people are invariably missed.
An attempt is made however to ‘catch’ everyone.

Strictly speaking the word ‘population’ in ‘population parameter’ is redundant just as


the word ‘sample’ is in ‘sample statistic’. The point is that a parameter is a number
describing some aspect of a population or a model (describing a population), and a
statistic is a number describing some aspect of a sample.

Exercise 4.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 10.17, 10.23,
10.35, 10.39, 10.41, 10.45.

4.3 Experiments

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 11.

After this reading write down answers to the following questions:

ˆ What is an observational study?

ˆ What is an experiment?

ˆ What is a retrospective study?

ˆ What is a prospective study?

ˆ What is an experimental unit?

ˆ What is a subject?

ˆ What is a treatment?

ˆ What is a factor?

ˆ What is a level?

ˆ State the four principles of experimental design.

ˆ What is meant by control?


4.3. Experiments 4.7

ˆ What is meant by randomisation?

ˆ What is meant by replication?

ˆ What is meant by blocking?

ˆ What is anecdotal data?

ˆ Give three advantages of experiments over observational studies.

ˆ What do we call the ideal simple experimental design? Draw a diagram to


represent this design for two treatments.

ˆ When is a difference statistically significant?

ˆ What is a control treatment?

ˆ When is an experiment single-blind? When is it double-blind?

ˆ What is a placebo?

ˆ What is a block? Why is blocking used?

ˆ What is matching?

ˆ Draw a diagrammatic representation of a randomised block design with two


blocks and three treatments.

ˆ What does confounding mean? Give an example.

There’s a lot of jargon here. For students of psychology and those in the physical sci-
ences, especially biology, the ideas and language of experimental design are especially
important. For business students, the language and ideas of sample surveys are more
important. For all of us however, a key point is to recognise that an experimental
study, provided the subjects or experimental units have been randomly assigned, of-
ten allows us to draw a conclusion about cause and effect whereas an observational
study does not.

For example, if tomato plants grown under fertiliser A produce significantly more
tomatoes than plants grown under fertiliser B, and the plants were randomly assigned
to the two fertilisers, we can legitimately conclude (everything else being equal) that
the difference in yield is due to the difference between the fertilisers. If however
one farmer is observed to use fertiliser A and another fertliser B on their respective
tomato plants, a difference in yield cannot necessarily be ascribed to a difference in
the fertilisers because lurking variables might be the cause. For example, other factors
such as different levels of care given by the farmers, or a difference in soil fertility or
climatic conditions might confound the results making it invalid to conclude that the
fertilisers caused the difference in yield (even if they did!).
4.8 Module 4. Gathering Data

Notice the use of the terms lurking variables, factors and confounding in this example.
These terms can be used in the context of any study, observational or experimental.
If we call a lurking variable a confounding variable or a confounding factor, nobody
will complain.
Incidentally, we also don’t need to be as rigid as the textbook concerning the definition
of an experiment. The text suggests an experiment involves manipulation of factor
levels, random assignment of experimental units to treatments (or, equivalently, of
treatments to experimental units) and comparison of responses across treatments. A
more generally accepted definition is simply a study in which an experimenter has in-
tervened in assigning experimental units to treatments, whether or not randomisation
has occurred.

Exercise 4.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 11.21, 11.23,
11.31, 11.47.

4.4 Using Random Numbers to Select a Sam-


ple
How do we go about choosing a simple random sample? The textbook talks about
random numbers and their availability in tables. It doesn’t give an expository example
of how to use them however.
Suppose, for example, we wish to run a survey of university students regarding
their attitude towards proposed new branding of the university (logo and associ-
ated colours). There are about 28000 students making up the university student
population but we are not about to survey all of them because we can’t afford to
and because we don’t need to. An accurate picture of the proportion of the student
population that, for example, favours the proposed new branding, can be obtained
from a simple random sample of the student population. How large should the sample
be?
Later in the course we find that the larger the sample the more precisely we can
estimate a population parameter such as the one we’re interested in here. There are
even formulas that help decide what this size should be. We don’t need to go into
those formulas. Let’s assume that a sample of 500 is appropriate. How do we select
those 500 from the 28000?
The first step is to create a sampling frame. In other words we create or obtain a
list of all students who are eligible (e.g. currently enrolled). Probably this list is
4.4. Using Random Numbers to Select a Sample 4.9

alphabetically ordered, but it doesn’t need to be. It just needs to be complete. The
members of this list are then numbered. The usual numbering system is such that
all members of the sampling frame have a unique number with an equal number of
digits. For example, if there are 28825 members in the sampling frame, we might
number the members consecutively from 00001 to 28825. Or we might number them
from say 10000 to 38824.

To obtain the chosen 500 we just generate 500 random numbers from the numbers
assigned to the sampling frame and these random numbers identify the 500 individuals
in the simple random sample. In principle this can be done using a table of random
numbers. In practice, with all but small random samples, a computer package such as
Excel or spss would be used making use of their built-in random number generators.

To generate a small random sample let’s see how to use the random number table
in the back of De Veaux, Velleman & Bock, (5th edition). The following example
follows the principles described above.

Example 4.1
Use the table of random digits to select a simple random sample of size five
from the following sampling frame:

Bock, D
Carmichael, C
De Veaux, R
Dunn, P
Fahey, P
McDonald, C
Moore, D
Norusis, M
Plank, A
Shi, M
Velleman, P
Khan, S

Number the members of the sampling frame from 01 (for Bock) to 12 (for
Khan). Notice that because there are 12 members in the sampling frame we
need to use two digits to identify each member. Go to the table of random digits
on page 1009 of De Veaux, Velleman & Bock, (5th edition). Pick a starting
point at random on this page (e.g. by letting something fall on the page). Pick
a forwards or reverse direction at random (e.g. by tossing a coin). Suppose we
start at the beginning of line 27 going forwards. Because our sampling frame
numbering uses two digit numbers, we mark off pairs of digits starting with 31,
08, 74, etc. The first pair 31 is ignored because no members of the sampling
4.10 Module 4. Gathering Data

frame have this number. The next pair 08 extracts Norusis, M. as the first
member of the sample. Continuing on, 74 is ignored as are a long sequence
of numbers until 05 occurs, identifying Fahey, P. as the second member of the
sample. Eventually we obtain the sequence of relevant numbers 08, 05, 05, 04,
08, 04, 04, 02, 11 giving the sample of numbers 08, 05, 04 02, 11 after ignoring
the repeats.
The required srs is therefore Norusis, Fahey, Dunn, De Veaux and Velleman.

Random sampling applies also in experiments. In a completely randomised experi-


ment, subjects are randomly allocated to treatment groups. This randomisation can
be achieved by dividing up the subjects into a number of simple random samples.
The following example shows how, using the table of random digits.

Example 4.2
An experiment compares three treatments A, B and C. Twelve subjects are
available and we allocate four to each group at random. Suppose the subjects
are the same ones as in the sampling frame in the previous example. As before
they are numbered from 01 to 12. The idea is to take a srs of size 4 and
put them in treatment group A, then take a random sample of size 4 from
the 8 remaining subjects and put them in treatment group B, and then the
remaining 4 subjects are put in treatment group C. (This is not the only way
of randomising subjects to treatment groups—you might think of another way
of doing it. The important thing though is that each subject must have an
equal chance of being in each group.)
Suppose we start reading the random table part way along row 16 and decide
to read backwards; i.e., the digits start as 7009754052. . . . Check for yourself
that after discarding repeats, the first eight relevant random number pairs are
09, 06, 11, 10, 08, 12, 02, 04. Therefore 09, 06, 11, and 10 go into group A, 08,
12, 02, and 04 into group B and the rest into group C. Hence the subjects are
allocated to the treatment groups as follows:

Group A Group B Group C


Plank, A Norusis, M Brock, D
McDonald, C Khan, S De Veaux, R
Velleman, P Carmichael, C Fahey, P
Shi, M Dunn, P Moore, D
4.4. Using Random Numbers to Select a Sample 4.11

Solutions to the following problems are given at the end of the module.

Exercise 4.3

The six people listed below are enrolled in a statistics course taught over
the Web. Use the list of random digits:

27102 56027 55892 33063 41842 81868 71035 09001 43367 49497 54580 81506

Starting at the beginning of this list, choose a simple random sample of


three to be interviewed in detail about the quality of the course. Use the
labels attached to the six names as follows:
1. De Veaux
2. Velleman
3. Santner
4. Goel
5. Jones
6. Klein

What is the sample you obtain?

Exercise 4.4

Newspoll wishes to carry out a poll to gauge voter sentiments for an


upcoming by-election. The electoral roll contains 86314 registered voters.
Describe how Newspoll might go about selecting the 1500 voters to be
polled.

Exercise 4.5

Consider a comparative experiment involving three diets and a control.


Randomly assigned 20 subjects in equal numbers to the four treatments.
Suppose the subjects are labelled A, B, C, . . . , T. Use the random digits
table reading forward starting at row 10.
4.12 Module 4. Gathering Data

4.5 Closing Comments


Most of us, at some stage in our studies, will be required to make use of data collected
by others. Many of us will be required to collect our own data. All of us read or
hear in the media reports of studies, usually in the political or health fields, claiming
to demonstrate the efficacy of some new diet or to display the opinions of the public
about some current issue. The material we’ve touched on in this module is obviously
very relevant then to all of us.

The basic principles of good data collection design have been talked about, but in-
sufficient time is available in this course for putting them into practice. This is
unfortunate because, more than any other topic in this course, learning by doing is
invaluable here. Some follow up statistics courses such as Experimental Design in-
clude project work involving data collection. Also, research students will more than
likely be exposed to contexts which will involve data collection and implementation of
the principles discussed here. Anyway, we can at least call ourselves armchair experts
on data collection after this module, if nothing else.

As far as this course is concerned, we are now in a position to turn our attention to
how data, collected according to the principles of good design, can be turned into
useful information.

Note that a Quick Review and additional exercises for elements of Module 4 are given
after Chapter 11 in the text book (after the exercises).

Answers to Exercises

Exercise 4.3
De Veaux, Velleman, Jones. (The first three distinct digits between 1 and 6, reading
from the left are 2, 1, 5 representing these three people.)

Exercise 4.4
The sampling frame (i.e. list of voters on the electoral roll) could be numbered from
00001 to 86314. A sequence of random digits might then be obtained or generated.
For example, many calculators will generate random numbers (e.g. 0.348, 0.019,
0.272, 0.951, 0.877, . . . ) which can then be grouped into lots of five (e.g. 34801,
92729, 51877, . . . ). These are then used to pick out 1500 members of the sampling
frame to include in the srs ignoring any five digit numbers greater than 86314 or
4.5. Closing Comments 4.13

any that have previously occurred (e.g. 34801, 51877, . . . ). This is very laborious in
practice—fortunately there are computer packages that automate this process.

Exercise 4.5
Assuming the subjects are numbered consecutively from 01 to 20, we obtain the
following sequence of relevant random numbers 03, 07, 12, 15, 04, 06, 11, 09, 16, 17,
20, 05, 10, 01, 02. Assuming the subjects are firstly assigned to group 1, then to
group 2, group 3 and finally group 4, we obtain the following allocation:

Group 1 Group 2 Group 3 Group 4


C F T H
G K E M
L I J N
O P A R
D Q B S
4.14 Module 4. Gathering Data

4.6 Tutorial Module 4


The following section contains Tutorial 4 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1: Generating an appropriate sample


Stratified sampling involves taking simple random samples from each of a number of
groups (strata) into which a population has been divided. Suppose a random sample
of eight people is required, stratified according to discipline studied (b = business and
s = science) so as to contain equal numbers of business and science students. The
sampling frame is as follows.

Barry H (b) Bill C (b) Bill D (b)


Bill G (b) Bob D (b) Charles Y (b)
Colin P (s) David H (b) Dawn F (s)
Evonne G (b) Gerry G (s) Gough W (b)
Harriet T (s) Henry K (s) Janet R (s)
Jennifer C (s) Joan S (b) John G (s)
Laura S (b) Leonardo D (b) Manon R (b)
Mario C (s) Michael J (b) Michele P (s)
Nelson M (b) Olivia W (s) Paul D (b)
Paul N (b) Peter M (b) Russell C (b)
Samantha R (s) Wayne G (b) Wynifred M (b)

(a) Using a random number generator, select a simple random sample (SRS) of eight
people from the sampling frame defined above.

(b) How many business students (b) and science students (s) are there in your sample?

(c) Using a random number generator, select a second simple random sample.

(d) How many business students and science students are there in this sample?

(e) Compare your two samples: Did you expect the result you got? Explain your
answer with reference to the sampling process.

(f) Using a random number generator, select a stratified random sample, stratified
according to discipline studied so as to contain equal numbers of business and
science students.

(g) Describe, in detail, the steps you took to produce your stratified sample.

(h) Why might you decide that a stratified sample is more representative?
4.6. Tutorial Module 4 4.15

(Note: [Link] is
one possible random number generator or the function RANDBETWEEN() in Excel).

Question: How do we decide if a study is a good study?

Firstly, we need to determine the type of study being reported to be able to appreciate
its strengths and weaknesses. Does the study include the features of a good design?

Answer the following questions about the given reports, giving reasons for your com-
ments where appropriate. Due to copyright restrictions actual newspaper articles
reporting on real research studies were not able to be used for this tutorial; the
following activities refer to fictitious studies reported in a newspaper style.

Article 1

PROTEIN POST RACE ALLOWS RUNNERS TO


RECOVER QUICKER, SAYS STUDY
Runners who took protein bars and/or drinks post-race improved their
rates of muscle recovery, according to latest research. A study of more
than 1000 United Kingdom runners found they were able to bounce back
quicker for training and racing than competitors who refuelled simply
with water or a banana, or nothing at all. The British Institute of Elite
Joggers surveyed club athletes who competed at race distances from 10km
to 42.2km (marathon). The runners were asked what food and drink they
consumed within one hour of a specific event and whether they thought
it made any difference in subsequent days and weeks. They were also
asked to provide supporting evidence of training logs and race results
immediately after the ‘test’ event. The poll found that:

ˆ 62 per cent of runners had a protein drink and/or protein bar after
a race
ˆ 18 per cent of runners just drank water or soft drink
ˆ 10 per cent of runners consumed fruit, such as a banana, orange or
watermelon
ˆ 8 per cent of runners enjoyed a beer, wine or vodka cruisers post-race
ˆ 2 per cent of runners had nothing

Institute President Steve Coe said they were also asked about how soon af-
terwards they returned to training or racing. Mr Coe said a staggering 95
per cent of runners who consumed protein drinks or bars after races were
back on the track or road within 48 hours. ‘The vast majority reported
rapid recovery and were ready to run again very quickly,’ he said. ‘While
some of it might be mind over matter, there is without doubt proof in the
pudding that an early drink or meal of protein aids muscle recovery and
4.16 Module 4. Gathering Data

prepares a runner to return sooner. There is nothing wrong with water


and a banana after sporting activity. Some would say the more natural
the product the better. But this was not backed up in the figures, which
concluded that protein was certainly more beneficial. They usually took
24-48 hours longer to feel somewhere close to their pre-race state.’

The poll, perhaps not surprisingly, revealed that those who drank alcohol
or consumed nothing post-race were likely to not run again in the following
week. ‘Many reported feeling lethargic, tired, even hung over. They said
they did not usually train or race again for at least 3-4 days,’ Mr Coe said.
The Institute had long advocated consuming protein-rich food or drink in
the 30 to 45-minute window immediately after hard exercise. ‘We don’t
specify any particular product, but anything is better than a beer post-
race ... well maybe not for some,’ he said.
Runners’ Worldly, August, 2015

Article 2

BOOM OR BUST BEFORE YEAR ONE


Start-up businesses have a great shot at making money if they can survive
their first year, says a new study. A survey of 800 small to medium-sized
businesses in Queensland found that one out of every two folded within
the first 12-18 months. The Chamber of Industry and Commerce Group
contacted current and former members who started their business after
January 2014.

Research indicated that 398 SMEs were still trading today, while 402 were
no longer registered. CICG Director Jedidiah Jones said of those 402, al-
most 300 had closed the door before January 2015, while the remaining
100 shut it down in the following 6-9 months. ‘It is tough being in busi-
ness, particularly starting out,’ Mr Jones said. ‘Many start-ups open with
insufficient capital or backing. They hit the doldrums after three months
and simply fail to hang on.’ Mr Jones said, that, conversely, the 50 per
cent or so who make it beyond the first year and are still trading are seeing
a brighter light on the horizon. ‘They have become largely self-sufficient.
The bottom line is in the black, not the red. And they have got through
the tough times and are ready to grow.’

CICG provided support to start-ups to help them get going and hopefully
make it to that all-important second year. But there was little support
from state or federal governments. ‘Politicians need to go out on a limb
and help our struggling start-ups and entrepreneurs. These are the people
4.6. Tutorial Module 4 4.17

with the next great ideas, the Google, Facebook and Twitter creators of
this world,’ Mr Jones said. ‘However, they need financial help to get
through the first year. A 50 per cent fail rate is not good for a modern,
progressive country like Australia.’
The Monthly, CICG, January, 2016

Article 3

CANCABIS – POWERFUL ANTIDOTE TO THE MAN FLU


A new little green pill could prove to be a game-changer for over 70s look-
ing to fight off the common cold. A study in the latest Journal of Medical
Experimentation reveals the drug Cancabis is keeping those nasty sniffles
away in elderly male patients.
Background :
Reports that a cannabis-containing drug, Cancabis, normally used for
cancer treatment, had an unexplained, side-effect: combatting colds.
Method :
Scientists conducted a fully randomised, placebo-controlled, double-blind
experiment. They tested 300 men aged 70 or older: 100 were given Can-
cabis; 100 were given the standard annual flu shot; and 100 a placebo
‘sugar pill’. The participants were asked to note the number of colds suf-
fered over the winter months (May-October) after taking the substances.
Results:
The magic Cancabis pill was a winner in seeing off the sniffles.
Cancabis:

ˆ 75 per cent said they had two colds or less


ˆ 21 per cent said they had 3-4
ˆ 3 per cent said 5-6
ˆ 1 per cent said more than 6 colds.

Flu shot:

ˆ 65 per cent said they had two colds or less


ˆ 15 per cent said they had 3-4
ˆ 15 per cent said 5-6
ˆ 5 per cent said 7+

Control:

ˆ 35 per cent said they had two colds or less


ˆ 25 per cent said they had 3-4
ˆ 25 per cent said 5-6
4.18 Module 4. Gathering Data

ˆ 15 per cent said 7+

Associate Professor Geoff Baggins, from Monashy University, said the


new wonder drug could help pensioners see off the winter blues and help
save them money. ‘It is amazing. We all thought the little blue pill
was the wonder drug. But this green pill could leave it for dead,’ he
said. ‘What started out as pain relief for cancer suffers would appear to
have a surprising impact for mature men struggling to fight off flu.’ The
study said there was anecdotal evidence from participants who said they
had used both Cancabis and the flu shot and this seemed to cancel the
benefits of both out. ‘Green is good in today’s world,’ Assoc Prof Baggins
said. ‘This is very good news for some of our aging citizens. Good luck
to them and hopefully they will not have to put up with the taunts of
man-flu for much longer.’
Journal of Medical Experimentation, December, 2015

Question 2: What type of study is reported in each article?

(a) Article 1: Observational/Experimental? Explain

(b) Article 2: Observational/Experimental? Explain

(c) Article 3: Observational/Experimental? Explain

Question 3: For each of these three articles, answer the questions related
to that type of article.

If the article is about an experiment, answer the following questions:

(a) Identify:

ˆ factors, levels and/or treatments:


ˆ response variable(s):
ˆ group sizes:

(b) What, if any, cause-and-effect conclusion is reasonable?

(c) Are there any lurking variables that might explain the observed association(s)?

(d) Do the principles of control, randomisation, replication and blocking appear to


have been used? Explain.

(e) Draw a diagram that represents the experimental design used in the study.
4.7. Module 4 Checklist 4.19

If the article is about an observational study, answer the following ques-


tions:

(a) Identify:

ˆ the variable(s) of interest:


ˆ the subjects studied:
ˆ how the subjects were selected:

(b) What, if any, cause-and-effect conclusion is reasonable?

(c) Are there any lurking variables that might explain the observed association(s)?

Question 4: What are the flaws?

A study to test the effects of marijuana recruited Australian adults (aged 18-35 years)
who have used marijuana. 20 adults were randomly assigned to smoke marijuana
cigarettes, while 25 adults were given placebo cigarettes. What are the flaws in this
study?

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

4.7 Module 4 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ recognise whether a study is observational or experimental;

ˆ recognise the existence and potential effect of lurking variables on both obser-
vational and experimental studies;

ˆ distinguish between samples and populations;

ˆ distinguish the sampling designs: simple random sampling, stratified sampling,


systematic sampling and multistage sampling;

ˆ recognise the possibility of bias in non-random sampling methods;

ˆ recognise the existence of problems due to undercoverage, nonresponse and


wording of questions in sample surveys;
4.20 Module 4. Gathering Data

ˆ identify the factors (explanatory variables), treatments, response variable(s)


and individuals (or subjects, or cases) in an experiment;

ˆ explain why a randomised comparative experiment can give good evidence of


causation;

ˆ display diagrammatically a completely randomised design specifying group sizes,


treatments, and the response variable;

ˆ recognise the importance of a control group, placebos, the double-blind tech-


nique and of replication;

ˆ recognise and describe the layout of a matched pairs design;

ˆ identify blocks and understand why blocking can improve an experiment;

ˆ use random numbers from a table or computer software to select a simple ran-
dom sample (SRS) and a stratified random sample;

ˆ use random numbers to assign experimental units to treatment groups.


Module 5
Statistical
Inference for
One Mean
5.2 Module 5. Statistical Inference for One Mean

Contents
5.1 Sampling Distribution of a Sample Mean . . . . . . . . . . . . . 5.6
5.1.1 Modelling the Distribution of a Sample Mean . . . . . . . . . . . . 5.6
5.1.2 The Central Limit Theorem (CLT) . . . . . . . . . . . . . . . . . . 5.7
5.1.3 How Large is Large? . . . . . . . . . . . . . . . . . . . . . . . . . . 5.8
5.1.4 z-Score Formulas . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.10
5.1.5 Reporting Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . 5.11
5.1.6 Jargon and Other Things . . . . . . . . . . . . . . . . . . . . . . . 5.13
5.1.7 In summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.15
5.2 Statistical Inference for a Mean . . . . . . . . . . . . . . . . . . . 5.16
5.2.1 Estimating with Confidence for One Mean . . . . . . . . . . . . . . 5.16
5.2.2 Hypothesis Testing for One Mean . . . . . . . . . . . . . . . . . . . 5.19
5.2.3 Sample size determination . . . . . . . . . . . . . . . . . . . . . . . 5.26
5.2.4 More about Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.27
5.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.29
5.4 Tutorial Module 5 . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.30
5.5 Module 5 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 5.32
5.3

Module Objectives
On successful completion of this module students should be able to:

ˆ distinguish between parameters and statistics;

ˆ recognise that a statistic will take different values when sampling is repeated;

ˆ recognise that statistics based on large samples are less variable than statistics
based on small samples;

ˆ state the meaning of a sampling distribution;

ˆ interpret the meaning of a standard error;

ˆ describe the sampling distribution model of a mean;

ˆ state and interpret the Central Limit Theorem;

ˆ use the normal model to approximate probabilities associated with a sample


mean;

ˆ explain what is meant by statistical significance;

ˆ explain when a t-model rather than the standard normal model is appropriate;

ˆ determine by hand and using spss a one-sample t-interval;

ˆ carry out by hand and using spss a one-sample t-test;

ˆ state the assumptions for and check the appropriateness of a one sample t
procedure;

ˆ explain what is meant by a P -value;

ˆ make a decision based on a prescribed level of significance;

ˆ state the meaning of power, Type I and Type II errors;

ˆ explain the difference between statistical significance and practical importance;


and

ˆ use a z procedure to estimate the size of a sample required to estimate a mean


to within a specified margin of error at a specified level of confidence.
5.4 Module 5. Statistical Inference for One Mean

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
So far the course has focused on describing and producing data with a bit about
probability thrown in. The rest of the course concentrates on drawing conclusions
from data about the world at large in a range of different scenarios.

Collectively this is called Statistical Inference. Estimation (in the form of a confi-
dence interval)—estimating a population parameter using a sample statistic is one
aspect of statistical inference. Hypothesis testing—making decisions about popula-
tions (including the values of population parameters) based on sample information is
the other.

Introduction to Hypothesis Testing


Hypothesis testing, also known as significance testing, is introduced here where we
discuss the idea of drawing conclusions about the world at large using the data in a
sample collected from a population. Perhaps the best way to introduce hypothesis
testing is using an example.

Example 5.1
One role of a university student counsellor is to advise students on techniques
to improve their study skills. Good sleep is believed to be part of any good
study routine and caffeine is thought to be a possible issue in maintaining good
sleep patterns.
Our interest is in ascertaining, on the basis of information obtained from 151
randomly-chosen students, whether there is an association between the amount
of coffee consumed and the number of hours of sleep students manage to have
per night. What we are doing then is asking a question about a population (all
students of the university) and trying to answer it on the basis of a represen-
tative sample (our 151 randomly-chosen students) from that population.

This is the idea of hypothesis testing. That description suggests we are testing the
plausibility of some conjecture and that’s essentially what hypothesis testing amounts
5.5

to. We take some statement (a hypothesis) about a population and test the consis-
tency of data with that statement. We don’t go as far as suggesting we’re trying
to establish the absolute truth or otherwise of some statement about a population.
We are seldom in a position on the basis of the limited information available in our
sample to be able to do that. What we hope for is that we can give some measure of
how consistent our data is with the hypothesis and draw a conclusion from that.

Let’s look at how that works with this student data. The hypothesis we are testing
might be stated as ‘There is an association between the amount of coffee consumed
and the number of hours of sleep students manage to have per night’. This is called
the ‘research’ or ‘alternative hypothesis’ (Ha ). The starting point of the test
always is to be sceptical; that is, we assume initially that the alternative hypothesis
does not hold. In other words we assume that ‘There is no association between the
amount of coffee consumed and the number of hours of sleep students manage to have
per night’. This is called the ‘null hypothesis’ (H0 ). The null hypothesis then is a
statement of the absence of whatever it is that the alternative hypothesis is stating.
The word ‘null’ is used to indicate this absence.

The test proceeds by assuming the null hypothesis to be true and testing how
consistent the data is with this assumption. If we decide that the data is inconsis-
tent with the null hypothesis then we would be inclined to believe the alternative
hypothesis rather than the null hypothesis described the population correctly.

How do we decide on how consistent the data is with the null hypothesis?

We need to produce an objective measure of how different is different. This measure


is called the test statistic and is calculated on the assumption that the null hypothesis
is true. The calculated test statistic is based on the context being investigated. If the
calculated test statistic is extreme in the distribution of possible test statistic values
for the particular context, we may be able to reject the null hypothesis in favour of
the alternative hypothesis. A measure of ‘how extreme a test statistic is’ is called the
P -value. The size of the P -value provides guidance on whether our sample gives us
evidence in favour of the alternative hypothesis or whether the difference we observe
is just due to sampling variation. In order to measure the P -value, knowledge of
sampling variation and the distribution of possible test statistic values for a given
context is required.

In this module we introduce these concepts in the context of a population mean. As a


first step to statistical inference for one mean, we consider the sampling distribution
of a sample mean.
5.6 Module 5. Statistical Inference for One Mean

5.1 Sampling Distribution of a Sample Mean


The concept of a sampling distribution is crucial in statistics because it allows us
to make the jump from simply describing data sets (descriptive statistics) to using
data to help us make decisions and answer questions about populations (inferential
statistics). Sampling distributions and probability are important in understanding
the logic behind inferential statistics. Having said that however, don’t think you
need to study this section to the exclusion of all else until it becomes crystal clear—it
is still possible to make use of the methods of inferential statistics without completely
mastering the methods in this particular section. (In fact many users of statistics do
exactly that all the time!) It is however far more satisfying to know why a particular
method works rather than just how to use it.

5.1.1 Modelling the Distribution of a Sample Mean

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 17 (Sec-
tion 1).

After this reading write down answers to the following questions:

ˆ What shape has the sampling distribution of the sample mean if the population
is normal?

ˆ What shape has the sampling distribution of the sample mean in general?

ˆ What is the mean of the sampling distribution of the sample mean?

ˆ What is the standard deviation of the sampling distribution of the sample mean?

ˆ What is the fundamental theorem of Statistics?

ˆ What does the Central Limit Theorem (clt) state?

ˆ Is the variability of a sample mean greater than, equal to or less than that of
an individual observation?

ˆ What is the basic message of the Law of Diminishing Returns?

ˆ What is a standard error?


5.1. Sampling Distribution of a Sample Mean 5.7

Try not to lose sight of the basic concept in this section. The idea of averaging,
whether it’s to work out a mean as a way of estimating something about the world
at large, is best done with larger rather than smaller samples because the larger the
sample the closer we expect the mean to be to the true value. That’s not particularly
brilliant! It’s really just commonsense (or an application of the Law of Large Numbers
if you want to impress!).

This section provides us with a formula that allows us to apply this basic idea to real-
world problems. Notice the presence of the square root of n in the standard error
formula for a mean. This is no accident. It’s what the law of diminishing returns
is getting at—one of the golden rules of statistics is that as n gets bigger things
only improve by the square root of n. For example, how much better off would
we
√ be with a sample that is ten times as large. The answer is that we would be only
10 = 3.2 times as well off. Or putting this another way. If we wanted to be ten
times as well of, we would need 102 = 100 times as much data. Sad, but true—that’s
the nature of the world.

How do we know how well off we are? Well, that’s the idea of the standard error. The
smaller the standard error the better off we are. Recall the standard deviation is like
a ruler measuring how far we are away from the true mean (as described in Module 1
and Chapter 5 of De Veaux, Velleman & Bock, (5th edition)). When we calculate
a mean from some data then we’d like to know how far away the value we’ve got is
from the true mean.

A standard error is just a standard deviation, but in this case it’s the standard
deviation of a mean. Actually, to be more exact, the term standard error is used to
describe our best estimate of the standard deviation of any statistic, whether it’s a
mean, proportion, correlation, intercept, slope, or even another standard deviation!

It follows that we wouldn’t expect a mean to be more than about three standard
errors away from the true mean according to the 68-95-99.7 rule provided the mean
is approximately normally distributed.

5.1.2 The Central Limit Theorem (CLT)


The sample mean (ȳ) varies from sample to sample in repeated sampling. If the sam-
ple size is large then the sampling distribution of the mean is approximately normal
regardless of the shape of the population distribution. This property is known as the
Central Limit Theorem. More formally, if y is random variable with mean µ and
standard deviation σ then for a random sample of size n the sampling distribution
of the sample mean, ȳ, is approximately normal with mean µ(ȳ) = µ and standard
5.8 Module 5. Statistical Inference for One Mean

deviation σ(ȳ) = √σ , provided n is large.


n

The CLT allows us to use the normal model to answer probability questions about
the sample mean even though we don’t know the population distribution.

Example 5.2
The ages of individuals in a srs of size 50 are recorded. Suppose the mean
and standard deviation of age of the population from which the sample was
drawn are 44.5 years and 13.2 years respectively. What is the probability that
the mean age in years of our sample is

(a) greater than 50?


(b) within one year of the population mean?

Solution
By the clt, the sampling distribution of the sample mean Y is approximately
normal with mean 44.5 years and standard deviation 1.8668 years (the standard
error of Y ). This distribution is displayed in Figure 5.1.

(a) The required probability is the area to the right of y = 50. Converting to
a z-score we have z = (50 − 44.5)/1.8668 = 2.95. From Table Z this gives
an area below y = 50 of 0.9984. Hence the probability that the mean age
in years of our sample is greater than 50 is 1 − 0.9984 = 0.0016 ≈ 0.2%.
(b) ‘Within one year of the population mean’ means that Y is between
43.5 years and 45.5 years. The respective z scores for these ages are
−0.54 and 0.54, for which Table Z gives 0.2946 and 0.7054. Hence the
probability that the mean age of our sample is within one year of the
population mean is 0.7054 − 0.2946 = 0.4108 ≈ 41.1%.

Exercise 5.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 17.27, 17.51.

5.1.3 How Large is Large?


How do we know if the normal model applies? Well basically it will if n is sufficiently
large. How large does a sample need to be in order that we can use the normal
distribution to describe the sampling distribution of the mean?
5.1. Sampling Distribution of a Sample Mean 5.9

Figure 5.1: The sampling distribution of mean age.

Statistics texts talk about ‘large samples’ and ‘small samples’. By ‘large samples’
what is usually meant is, large enough so that we can make use of the clt! Conversely,
by small samples is meant, not large enough to safely use the clt! Nobody can tell
us exactly when n changes from being small to large. Many factors are involved not
the least of which is how approximate is approximate. Nonetheless a few guidelines
exist based on experience and theory.

ˆ Heights, weights and numerous dimensional measurements on many populations


of animals, plants and so on in nature are approximately normal in distribution
provided the population is reasonably homogeneous; e.g., animals of the same
species, sex and of similar age. For srs’s from such populations the normal
distribution can be applied to describe the sample mean for any-sized sample.

ˆ Errors in a sequence of measurements read off an instrument such as a ruler,


gauge, scale, or so on are often approximately normally distributed. It follows
then that if ten measurements are made of the distance say from your place of
work to home to sufficient precision so that the measurements vary, they can
be thought of as a sample from an approximately normal distribution.

ˆ For srs’s of size 25 or more, there is seldom any problem in using the clt to
argue that the sampling distribution of the mean is approximately normal.

ˆ We should always plot the data (histogram) from which the mean is being
calculated and make some judgement about it depending on the amount of
skewness and the presence or otherwise of outliers. If the population appears
likely to have just a single peak and unlikely to be severely skewed based on the
sample plot, and the sample contains no outliers, applying the clt to a srs of
5.10 Module 5. Statistical Inference for One Mean

size 15 or more is likely to give reasonable results. Plots of samples of size less
than 15 can suffer a lot from sampling error and are therefore not very reliable.
If no other evidence is available about the ‘true’ distribution apart from such a
plot, it will be necessary to make it clear up front that normality is assumed,
or make use of a method that does not rely on normality (see Section 6.3).
Outliers also can make a big difference to an analysis if the sample is small and
it pays to check out just how big that difference is by doing the analysis with
and without the outlier(s).

5.1.4 z -Score Formulas


Problems dealing with the sampling distribution of a mean lead to calculating a
z-score and looking up Table Z for the answer.

A z-score is just a measure of how many standard deviations the value of interest is
above the mean. In terms of a formula when dealing with sample means we have

y−µ
z= (5.1)
√σ
n

We have found z-scores in a previous situation. In Module 2 we used the formula


y−µ
z= (5.2)
σ
to find proportions or fractions associated with variables described by a normal model.

These formulas involve taking a value, whether it be y or y, subtracting what it’s


expected to be (the mean), and dividing by its standard deviation.

The similarity between (5.1) and (5.2) often causes confusion so is clarified here.

Notice that if we put n = 1 in (5.1), the formula reduces to (5.2) apart from the bar
over the y. Obviously working out the mean of a sample of size one is pretty boring
because there’s only one number involved, but the mean of one number is really just
that same number, right? So we could drop the bar over the y in this case and formula
(5.1) would be exactly the same as (5.2). All that is saying is that (5.1) works for
any size sample n including n = 1.

How does this help us? Well, the formula (5.1) is the general formula and (5.2) is
a special case of it. Thinking in these terms, problems such as 17.51 in De Veaux,
Velleman & Bock, (5th edition) should be less confusing. In part (a) of this problem
we are effectively asked about one pregnancy (n = 1). OK, it looks like we’re asked
5.1. Sampling Distribution of a Sample Mean 5.11

about all pregnancies but think of this question as asking what is the probability
of one pregnancy chosen at random being between 270 and 280 days and we get
the same answer as the original problem. Then in parts (c) and (d) the questions
involve n = 60 pregnancies. We can therefore use formula (5.1) in parts (a), (c) or
(d) provided we put in the right value of n! While we’re at it, part (b) is found using
the unstandardising formula
y =µ+z×σ (5.3)
which we’ve seen before as (5.2) rearranged. Better still, by rearranging (5.1) we get
σ
y =µ+z× √ (5.4)
n

which is the more general formula because it reduces to (5.3) when n = 1.

5.1.5 Reporting Statistics


What we have seen in this section is that a statistic, such as a mean, has an error
associated with it because the true value of the mean or whatever we’re interested
in could be different to the statistic we’re using to estimate it. For example, even
though the mean gestation period of human pregnancies is 266 days, we shouldn’t be
surprised if a sample of 10 pregnancies gave a mean of say 264 days. (250 days though
would be surprising given the standard deviation of 16 days—why is this?). 264 days
has an error then of −2 days (negative because, as for residuals in regression, the
observed value is below the expected value).

Usually in practice, we don’t know what the mean of the model is so we don’t know
what the error is. But we can still figure out the standard error! That’s very nice.
So, for the pregnancy example, even though we may not know that the model or true
mean is 266 days, or that the true standard deviation is 16 days, we can still calculate
the standard error from the data we have. Say, for example, the n = 10 pregnancies
had durations (in days) as follows:

269, 279, 245, 257, 259, 246, 270, 261, 279, 275

Then the sample mean is y = 264.0 days and the sample standard deviation is s =
12.47 days. The standard error (of the mean) is then
s 12.47
SE(Y ) = √ = √ = 3.94 days
n 10
and this tells us a lot about how far apart the sample mean y = 264.0 days and
(unknown) true mean are likely to be.
5.12 Module 5. Statistical Inference for One Mean

Figure 5.2: Mean rainfalls (with standard errors) for data in Example 1.1.

When it comes to reporting statistics such as means, it is appropriate (some of us


would say essential) to also report the value of the standard error and the sample
size to provide an idea of the precision of the statistic. In other words, we are letting
the reader know that the mean we have found, for example, is variable and that it
is likely to be different from the true mean. If we were to do the study again, we’d
likely find a different result. How much different? Well the standard error gives some
idea of that. Including the sample size provides a more complete summary of the
study for the reader and will permit a more reliable interpretation of the results and
comparison possibly with previous studies.

It’s common practice in reporting statistics to include the SE and n where appropriate.
For example, ‘The mean gestation period was 264.0 (±3.9) days (n = 10)’.

Included with this method of reporting should be a clear statement explaining that
the bracketed value is the standard error. A footnote at the first occurrence is often
used (for example, ‘statistics are reported (±SE)’, or an indication as appropriate in
the column headers of a table containing summary statistics).

Failing to give some idea of the precision of important summary statistics can be quite
misleading. If for example a report stated that 54.36% of respondents were against
the proposed change, the reader would be forgiven for believing that a majority of
the target population were against the change. The statement 54.4% (±7.5%) puts a
different perspective on matters. It suggests that a repeat of the study may well find
a minority of respondents against the change since 54.4 − 7.5% is below 50%.

It’s smart also to include standard errors in graphs representing statistics. An ex-
ample is shown in Figure 5.2. See the spss video Error Bars on the StudyDesk for
instructions on how to do this.
5.1. Sampling Distribution of a Sample Mean 5.13

5.1.6 Jargon and Other Things


Jargon is part and parcel of statistics. It is used for the purpose of avoiding ambiguity
and promoting conciseness.

Numerical characteristics of samples are called statistics. Numerical characteristics


of populations are called parameters.1 Summary measures used to describe samples
such as the mean Y , variance s2 and standard deviation s, are examples of statis-
tics. Measures such as the mean µ, variance σ 2 , standard deviation σ describing
populations are examples of parameters.

The distinction between statistics and parameters is fundamental in Statistics (that


is, Statistics, singular!). Statistics (plural) are to samples as parameters are to pop-
ulations. Much of what we do in the rest of this course involves using statistics such
as sample means to estimate parameters such as population means; for example, es-
timation of the mean income of all Australian workers from the mean income of a
sample of Australian workers.

Statistics are random variables. Parameters are not. Every sample of Australian
workers can provide a sample mean income, and these means will vary from one
sample to the next. The mean income of all Australian workers however is not a
variable but, at least at some point in time, is a fixed number that would remain
unknown unless we were to take a census at that time.

Probability or population distributions are characterised by parameters. The normal


distribution is characterised by two parameters, µ and σ (or σ 2 ). Once these two
parameters are known, the exact form of the distribution is known. We say that µ
and σ are the defining parameters of the normal distribution.

Statistics and defining parameters come together in sampling distributions. A sam-


pling distribution is the distribution of values taken by a statistic obtained from all
possible samples of a particular size from a population. We have dealt with the sam-
pling distribution of the sample mean in this section. The defining parameters of this
sampling distribution involve the important parameter µ. Hence the connection, vital
to a large component of the rest of this course, is made between the sample mean
and population mean.

A statistic such as a sample mean has got a standard deviation. We know that the

standard deviation of the sample mean Y is SD(Y ) = σ/ n where n is the sample
size. This formula is not much use to us though because we need to have values for
the parameter σ to use it, and we seldom know what this value is. What we do in
practice is replace the σ with the sample standard deviation s. Then we talk about

the standard error of the sample mean Y as SE(Y ) = s/ n. In general terms
1
Remember that samples and statistics both start with ‘s’; parameters and populations both start
with ‘p’.
5.14 Module 5. Statistical Inference for One Mean

then, a standard error is an estimate of the standard deviation of a statistic. You will
notice as you keep using spss that standard errors get reported for many statistics,
not just for means. Just remember they are estimates of the standard deviation of
the distribution of the statistic.

Usually, upper-case letters such as X and Y are used to refer to random variables
and lower-case letters such as x and y to refer to observed values of random variables.
When we talk about a random variable or a statistic in a random or sampling situation
we should use the upper-case version of the letter, e.g. Y . The lower-case version, y,
refers to values that are actually observed in a random or sampling situation. For
example: ‘The value of y from the experiment was 12.4’. Similarly, on the graph of the
distribution of X we use the upper-case X when talking about the distribution itself
and the lower-case x to describe particular values of the variable on the horizontal
axis. In practice, in this course, you will not be penalised for inappropriate use of
lower and upper-case letters.

Write down answers to the following questions to check your understanding of this
material:

ˆ What is a population?

ˆ What is a sample?

ˆ What is a parameter? Give an example.

ˆ What is a statistic? Give an example.

ˆ Write down the symbols used to describe a sample mean and a population mean.

ˆ What are defining parameters?

ˆ What are the defining parameters of the normal distribution?

ˆ What is a standard error?

ˆ What is the formula for the standard error of a sample mean?

ˆ In what sense does a law of diminishing returns apply to sampling?

ˆ What general type of variable can commonly be described by a normal model


as n gets large? Give an example.

ˆ When can we use the normal model for a mean?


5.1. Sampling Distribution of a Sample Mean 5.15

5.1.7 In summary
A number of concepts of importance to the rest of the course are discussed in this
section of the module.

Firstly, the sample we collect or the sample that we are presented with should be
thought of as only one of many possible samples that could have been obtained of
the same size from the population. Hence the sample mean or standard deviation
or whatever statistic we calculate from the data is really just a value of a variable.
Knowing about the distribution of the variable is the basis of the idea of generalising
from the data at hand to the world at large. We introduce this in more detail in the
next section.

The distribution of possible values of the statistic we are interested in is called a


sampling distribution. When the sample is big enough the normal model describes
the sampling distribution for a mean (that’s what the Central Limit Theorem is all
about). It turns out in fact that when the sample is large enough, practically all
statistics can be described by a normal model.

As sample size n increases, the variability of these statistics decreases. This variability
is measured by the standard error. The standard error gets smaller in proportion to
one over the square root of n. This means that the rate at which the SE gets smaller
slows down as n gets larger and so a law of diminishing returns applies.

Because the normal model applies to statistics of data (at least when the sample is not
too small), the mean and standard deviation are appropriate measures of centre and
spread for sampling distributions. Hence, even though the median and iqr may better
describe the centre and spread than the mean and standard deviation of a population
which is skewed, when our main interest is in sampling with a view towards drawing
conclusions, the mean and standard deviation are more important parameters than
the median or iqr. For this reason in the rest of this course we focus on the mean and
standard deviation rather than median and iqr as measures of centre and spread.

The ideas in this section are used in one way or another in the rest of the course. If
you are still struggling with them, don’t worry and don’t spend too much time on this
material. The relevance of this material will become clearer once you have actually
applied these concepts in conducting a range of inferential methods.
5.16 Module 5. Statistical Inference for One Mean

5.2 Statistical Inference for a Mean


In this module we have been introduced to the idea of statistical inference – general-
ising from data (samples) to the world at large (populations) by doing a hypothesis
test or using a confidence interval. We have seen the theoretical basis of statistical
inference through sampling distributions. Now we are ready to apply these ideas in
many situations.

Specifically, in this module, we look at a scenario involving one mean. From the Cen-
tral Limit Theorem, we know that the sampling distribution of the mean is approxi-
mately normal with mean µ and standard deviation √σn regardless of the distribution
of the population, if n is large. However, if the population standard deviation is not
known (which is generally the case), the next best thing is to use the sample standard
deviation s instead of σ. In doing so, the sampling distribution of the mean becomes
a t distribution provided the underlying population has an approximate normal dis-
tribution. This ‘nearly normal’ condition in relation to the population distribution is
important when the sample size is small, but not so critical for larger samples.

Before we undertake any statistical inference procedure, we need to be mindful of


certain conditions that need to be satisfied before doing so. Since we are using
sample information to infer something about a broader population, our sample needs
to be representative of that population. For statistcial inference about one mean
using a t procedure, this requires randomness in the selection of the sample from the
population, a nearly normal population distribution from which the sample is taken,
independence of each case from all others and the size of the sample being less than
10% of the population. In practice these conditions/assumptions should be checked
before undertaking statistical inference.

De Veaux, Velleman & Bock, (5th edition) covers the material discussed in this section
in parts of Chapters 17, 18 and 19. In addition to presenting inference for a mean
these chapters include aspects of inference for a proportion. This course does not
specifically cover inference for proportion although many of the concepts in these
chapters are common to both mean and proportion.

5.2.1 Estimating with Confidence for One Mean

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 17 (excluding
Section 4).
5.2. Statistical Inference for a Mean 5.17

After this reading write down answers to the following questions:

ˆ Describe the general form of a confidence interval.

ˆ State in non-technical terms what is meant by ‘95% confidence’ or other state-


ments of confidence in statistical reports.

ˆ What is the formula for the standard error of a sample mean?

ˆ What is the difference between the Student’s t-model and the normal model?

ˆ What is the formula for the standardised sample mean when σ is unknown?

ˆ In what ways is a t distribution similar to a standard normal distribution? In


what ways is it different?

ˆ What table do we use to find critical t values?

ˆ What assumptions are needed to use Student’s t-models?

ˆ What is meant by nearly normal conditions? When is normality important?


When is it less important?

ˆ How is the normality condition checked?

ˆ What is the formula for the one-sample t-interval for the mean?

ˆ What is the formula for the degrees of freedom of the Student’s t-model in this
context?

The idea of a confidence interval (CI) is a natural extension of the idea of a standard
error. In Section 5.1.5 we suggested (actually, more like insisted) that we include
standard errors when describing sample means. Provided the sample size is large
enough so that the normal model applies, a straightforward application of the 68-95-
99.7 rule tells us that y ± SE is really a 68% confidence interval, y ± 2 × SE is a 95%
confidence interval and y ± 3 × SE is a 99.7% confidence interval for µ.

To find a confidence interval for a mean we use the formula


s
y ± t∗n−1 √ (5.5)
n

We don’t usually talk about 68% or 99.7% CI’s though, although 95% CIs are popular.
The common confidence levels are given at the bottom of Table T. When using formula
(5.5), the degrees of freedom used to find the critical value t∗ in Table T is given by
n − 1. The degrees of freedom in other situations when t is used are given by other
formulas (you will encounter this in later modules). Performing a one-sample t-test
and constructing a confidence interval for a mean using the t distribution are basic
5.18 Module 5. Statistical Inference for One Mean

skills in Statistics. You should be able to perform the analysis both by hand and, if
the data is available, by using spss.

Notice that t and z are very similar—z is used if σ is known; t is used if σ is not
known. If s is replaced by σ in the above formula then the t becomes a z and Table Z
rather than Table T is used for the calculations (however, in practice we rarely know
σ). Don’t be put off by the ∗ on the t and z. It just indicates that these values are
‘important’ values that can be found in Table T. If you leave them off the formulas,
nobody will complain.

Example 5.3
Beanies (a fruit-flavoured gummy-textured sweet in the shape of a broad bean
seed) are sold in 180 gram bags. A random sample of 24 bags of Beanies were
collected and contents were weighed by the members of a tutorial group of
students, yielding the following data:

185, 182, 177, 180, 181, 182, 171, 182, 176, 182, 182, 176,
185, 175, 175, 176, 175, 186, 170, 182, 172, 185, 176, 176

Find a 95% confidence interval for the weight of Beanies in bags labeled as
containing 180 grams.
Solution:
From the data, the sample mean (y) is calculated to be 178.71 grams and the
sample standard deviation (s) 4.686 grams.
Using Formula 5.5, with df = n − 1 = 24 − 1 = 23 and thus t∗ = 2.069 from
Table T,
s 4.686
y ± t∗n−1 √ = 178.71 ± 2.069 × √
n 24
= 178.71 ± 1.979

With 95% confidence we can say that the population mean weight of 180 gram
bags of Beanies is between 176.73 grams and 180.69 grams.
Checking the assumptions are satisfied for formulating this confidence interval:

(a) Randomisation: we are told that students took a random sample of bags
of Beanies.
(b) Nearly Normal : we have a sample size of 24 which is reasonably large so
the ‘nearly normal’ condition is not as critical, however we can check the
reasonableness of this assumption by plotting the data in a histogram to
check for outliers or skewness. Figure 5.3 shows no outliers and minimal,
if any, skewness.
5.2. Statistical Inference for a Mean 5.19

Figure 5.3: Histogram of Weight of Beanies using SPSS.

(c) Independence: with random sampling, it is reasonable to expect that the


weight of one bag of Beanies is independent of the weight of any other
bag of Beanies.
(d) 10% Condition: it is reasonable to expect that there would be more than
240 bags of Beanies in the population.

Having said all this, note that in the assignments you only need to discuss
assumptions if asked explicitly to do so, however, in practice, it is vital to check
the validity of using the procedure by checking that the assumptions/conditions
have been satisfied.

Exercise 5.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 17.11, 17.13,
17.21, 17.29, 17.39, 17.55 (a)-(e), 17.59, 17.61.

5.2.2 Hypothesis Testing for One Mean


Now that the idea of drawing conclusions about the world at large using confidence
intervals has been introduced, we’ll look at the other important way of drawing con-
clusions about the world at large. This is the method of hypothesis testing, also
known as significance testing.
5.20 Module 5. Statistical Inference for One Mean

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapters 18 and 19
(note that these chapters also refer to hypothesis testing for a proportion
which is not required for this course; however it is worthwhile reading all
sections of these chapters as the broad concepts discussed refer to both
means and proportions).

After this reading write down answers to the following questions:

ˆ What are the two hypotheses in a hypothesis test called?

ˆ What is another name for a hypothesis test?

ˆ In general terms what does the null hypothesis say?

ˆ What notation is used for the null hypothesis?

ˆ What notation is used for the alternative hypothesis?

ˆ What is a two-sided alternative? What is a one-sided alternative?

ˆ How do we decide whether or not an alternative is one or two-sided?

ˆ What distribution does the test of one mean make use of?

ˆ What is the form of the null hypothesis in the one mean t-test?

ˆ What is the form of the alternative hypothesis in the one mean t-test?

ˆ What is the test statistic formula for the one-sample-t-test for the mean?

ˆ What is a P -value?

ˆ What does a small P -value indicate?

ˆ What conditions need to be satisfied for the one-mean t-test?

ˆ What is the formula for sample size needed to estimate a mean?

ˆ What is a Type I Error?

ˆ What is a Type II Error?

ˆ What is the difference between significance and importance?


5.2. Statistical Inference for a Mean 5.21

In performing a hypothesis test for a mean we use the test statistic


y−µ s
t= where SE(Y ) = √ (5.6)
SE(Y ) n

When using formulas (5.5) and (5.6), the degrees of freedom used to find t∗ in Table T
is given by n − 1. The degrees of freedom in other situations when t is used are given
by other formulas (you may encounter this in other courses). Performing a one-sample
t-test and constructing a confidence interval for a mean using the t distribution are
basic skills in Statistics. You should be able to perform the analysis both by hand
and, if the data is available, by using spss.

On some occasions it’s not necessary that a definitive decision be made. Then it’s
sufficient to report the P -value itself with some covering comment.

Interpreting the P -value


Understanding the concept of a P -value is essential in correctly interpreting the re-
sults of a hypothesis test. It is useful to remember that a P -value is a measure of
the plausibility of the null hypothesis. A small P -value is indicative that it would be
highly unlikely to obtain a sample result such as that produced by the observed sam-
ple, if the null hypothesis was true. So, the credibility of the null hypothesis is under
question when P -value is very small, and hence we may reject the null hypothesis in
favour of the alternative hypothesis. In summary, the smaller the P -value, the less
plausible is the null hypothesis. The larger it is, the more plausible. Don’t fall into
the trap of saying the null hypothesis is either right or wrong. Ultimately all we can
say is that the null hypothesis is plausible in which case it cannot be rejected, or it is
implausible in which case the null is rejected and thus the alternative is supported.

As a guide as to how to interpret the P -value when making a conclusion, there are
two situations depending on whether or not a level of significance is given with the
problem. The level of significance is a pre-assigned threshold for deciding whether to
reject a null hypothesis.

In practice, if a decision one way or the other is required, we set a criterion on the size
of the P -value before doing the study so that a definitive conclusion can be reached
at the end. This size is called the significance level and is denoted by α. If a level
of significance is given then it’s a matter of simply comparing the P -value with α.
Common levels of significance would set α as 0.001 (0.1%), 0.01 (1%), or 0.05 (5%),
depending on the context. If the P -value < α then we conclude that H0 is rejected
in favour of Ha at the α level of significance. If, for example, it was decided that
a significance level of 0.10 or 10% was appropriate before a test was run, a P -value
of 0.06 (6%) would allow us to reject a null hypothesis in favour of the alternative.
However, if the level of significance had been set at 5%, a P -value of 0.06 (6%) would
5.22 Module 5. Statistical Inference for One Mean

not allow us to reject a null hypothesis in favour of the alternative. It is essential that
the level of significance be set before the data is collected if a definitive conclusion
is required! If a level of significance is given then it’s a matter of simply comparing
the P -value with α. Of course, we then put our conclusion into the context of the
problem.

On some occasions it’s not necessary that a definitive decision be made. Then it’s
sufficient to report the P -value itself with some covering comment. In this case, write
down the conclusion in words as follows according to the size of the P -value:

ˆ P -value > 0.10 (or 10%): state ‘there is insuffcient evidence to support Ha ’.

ˆ P -value between 0.05 (5%) and 0.10 (10%): state ‘there is slight evidence to
support Ha ’.

ˆ P -value between 0.01 (1%) and 0.05 (5%): state ‘there is moderate evidence to
support Ha ’.

ˆ P -value between 0.001 (0.1%) and 0.01 (1%): state ‘there is strong evidence to
support Ha ’.

ˆ P -value < 0.001 (0.1%): state ‘there is very strong evidence to support Ha ’.

Of course, these conclusions should be made in the context of the problem. In other
words, Ha should be replaced by a contextual statement of the alternative hypothesis
in non-technical language.

Hypothesis testing can be presented in four steps as follows:

1. State the null and alternative hypotheses (Check all relevant assumptions of the
test have been met);

2. Assume the null hypothesis is true and calculate the value of the test statistic;

3. Determine the P -value, a measure of the plausibility of the null hypothesis;

4. Write the conclusion in context.

Example 5.4
Using the data in Example 5.3, students decide to test the hypothesis that,
on average, the manufacturers of Beanies are underfilling the 180 gram bags,
using a 5% level of significance (note that the level of significance should be
decided before the data is collected).
Solution:
Using the four steps to hypothesis testing:
5.2. Statistical Inference for a Mean 5.23

ˆ Step 1: State the null and alternative hypotheses (Check all relevant
assumptions of the test have been met).
H0 : µ = 180
Ha : µ < 180
where µ is the population mean weight of Beanies.
The conditions/assumptions for this question are the same as those for the
confidence interval and have been shown to be satisfied in Example 5.3.
ˆ Step 2: Assume the null hypothesis is true and calculate the value of the
test statistic (using Formula 5.6).

y−µ s
t= where SE(Y ) = √
SE(Y ) n
178.71 − 180 4.686
= where SE(Y ) = √
0.9565 24
= −1.349

ˆ Step 3: Determine the P -value, a measure of the plausibility of the null


hypothesis.
With df = 23 and t = −1.349, the one-tailed P -value from Table T is
between 0.05 and 0.10 (we ignore the negative sign for t when looking at
Table T; the negative sign just indicates the sample mean is less than the
hypothesised population mean).
ˆ Step 4: Write the conclusion in context.
With a P -value of between 0.05 and 0.10 there is only slight evidence to
support the notion that the manufacturers of Beanies are underfilling the
180 gram bags. With such a large P -value (i.e., only ‘slight evidence’),
it is possible that the observed difference between the sample mean and
hypothesised population mean is due to sampling variability or it may be
an indication of underfilling. Since there is only slight evidence to support
the notion of underfilling and we have a relatively small sample, it would
be advisable to obtain more information by taking a larger sample. We
should also consider the effect size. Is the difference we observe of practical
significance?

Using SPSS
Input the data in one column. Figure 5.4 shows a snapshot of the first 7 data
values for the sample of n = 24 bags of Beanies.
5.24 Module 5. Statistical Inference for One Mean

Figure 5.4: Input Data on Weight of Beanies into SPSS.

Produce Summary Statistics using the Explore procedure. In Figure 5.5 we


see that the Lower Bound and Upper Bound for a 95% Confidence Interval
for Mean agree with the values we produced in Example 5.3. This procedure
produces much more than we need so you have to be able to identify the relevant
components to perform the analysis by hand.

Figure 5.5: Summary Statistics on Weight of Beanies using SPSS.

Use the One-Sample T Test procedure to test the hypothesis H0 : µ = 180.


Notice in Figure 5.6 that the ‘Test Value’ is the hypothesised value stated in
the null hypothesis.
5.2. Statistical Inference for a Mean 5.25

Figure 5.6: One-sample T Test in SPSS.

Output from the One-Sample T Test. From Figure 5.7, the t-test statistic is
−1.350 which is slightly different to the −1.349 we obtained when doing the
hypothesis test by hand. This is due to round off error. We used values of the
mean and standard deviation rounded to 2 and 3 decimal places, respectively,
to calculate the t-test statistic, whereas spss uses many more decimal places
than this. The df, as expected, is the same. The ‘Sig.(2-tailed)’ is the P -value
for a two-tailed test. We have performed a one-tailed test (only interested in
underfilling), thus the P -value is 0.095, half of this Sig.(2-tailed) value of 0.190.
This is consistent with what we found before with 0.095 being between 0.05
and 0.10. To obtain the 95% confidence interval from this output we add the
null hypothesised value to each bound thus 180 − 3.27 gives a Lower Bound
of 176.73 and 180 + 0.69 gives an Upper Bound of 180.69 (in agreement with
what we obtained in Example 5.3).

Figure 5.7: Output from One Sample T Test in SPSS.


5.26 Module 5. Statistical Inference for One Mean

Exercise 5.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 18.11, 18.45,
18.47, 18.49, 18.53, 18.55.

Exercise 5.4
Watch the following spss video.

ˆ One Sample t-test

5.2.3 Sample size determination


To determine the minimum sample size required to estimate the unknown population
mean (µ) with a predetermined confidence level (1 − α) within a given margin of
error, say ME, use the formula
z ∗2 s2
n= (5.7)
ME2
where z∗ is the critical value of the standard normal distribution at the α level of
significance.

We would expect to have t∗ not z ∗ in this formula, but we don’t know n (we are
trying to find n) to be able to find the degrees of freedom and thus t∗ . However,
when the sample size n is large, t∗ and z ∗ will be about the same value. In this course
we will only use the sample size formula (5.7) when the answer we expect for n is
large, because then we can replace t by z.

The answer we get from this formula should be rounded up to the next whole number.
Either s (as shown) or σ can be used in this formula.

Example 5.5
Refer back to information given in Example 5.3. The students now wish to
estimate the population mean weight of 180 gram bags of Beanies to within
a margin of error of 1.5 grams with 95% confidence. What is the minimum
sample size required? (Use the sample standard deviation from the previous
sample to apply Formula 5.7).
5.2. Statistical Inference for a Mean 5.27

Solution:
The minimum required sample size is

z ∗2 s2
n=
ME2
1.962 × 5.1462
=
1.52
= 45.21

Rounding up 45.21 to the next integer number, we get the required sample size
to be 46.

Note: A quick way to find z ∗ values is to use the bottom row of Table T. The common
confidence levels are given at the bottom of Table T; 95% is there with a critical value
for z in the row immediately above it, namely 1.960 (which is pretty close to the 2
indicated by the 68-95-99.7 rule). We use the values z = 1.960 (95%), 1.645 (90%),
and 2.576 (99%) so often that they are worth remembering to save time in looking
them up.

5.2.4 More about Tests

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 19. The sec-
tion about P -values (Chapter 19, Section 1) is optional since it involves
notation (conditional probabilities) which was deliberately omitted from
the course. However, reading this section might give you a better under-
standing of P -values.

After this reading write down answers to the following questions:

ˆ What is a significance level and how is it used?

ˆ What symbol is used to denote the significance level?

ˆ What are the commonly used levels of significance?

ˆ What are the two types of decision errors called?

ˆ What is meant by a Type I error?


5.28 Module 5. Statistical Inference for One Mean

ˆ What is meant by a Type II error?

ˆ How are α and Type I errors related?

ˆ What is the power of a test?

ˆ In what way is a confidence interval and a two-sided hypothesis test equivalent?

ˆ What can be said about H0 : µ = µ0 at the 5% level of significance if µ0 lies


inside a 95% confidence interval and the test is two-sided?

ˆ Apart from the size of the effect being estimated, what else affects the P -value
in a test of significance?

ˆ How might the practical significance rather than the statistical significance of
an effect be judged?

De Veaux, Velleman & Bock, (5th edition) put emphasis on the null hypothesis in
setting up a hypothesis test and making conclusions from it. We prefer to give the
emphasis to the alternative hypothesis, often describing it as the research hypothe-
sis to indicate that it represents the hunch or conjecture that a researcher believes
represents reality. As such the alternative is decided on first. The null hypothesis
then represents an absence of the effect or condition described in the alternative hy-
pothesis. The reason for the text emphasis on the null hypothesis is that the test
procedure focuses heavily on the null hypothesis. The null hypothesis is assumed to
be true when doing a hypothesis test. The aim is to find whether or not the data is
consistent with the null hypothesis. We also accept conclusions which focus on the
alternative hypothesis and whether or not there is sufficient evidence provided by the
sample to support it.

Some hypotheses arise as a claim made by a manufacturer. An example of this type


is “bottles of our product contain no less than 375 ml.” We refer to such a statement
as a manufacturer’s claim. If we want to dispute such a claim our interest would be
in making the alternative hypothesis the negative or opposite of the manufacturer’s
claim. Hence the alternative for this example would be “bottles of this product
contain less than 375 ml.” In situations like this, in which our interest is in disputing
a manufacturer’s claim, the claim itself is the null hypothesis and the negative of
the claim is the alternative hypothesis. Although only a brief reference is made in
the textbook to the relationship between confidence intervals and hypothesis testing
(Example 19.5), an understanding of this relationship is useful in appreciating the
subtleties of hypothesis testing.

Some related definitions


A Type I Error is committed if a true null hypothesis is rejected due to suspect ev-
idence from a sample. In technical terms, it is the rejection of the null hypothesis
5.3. Closing Comments 5.29

when in fact the null hypothesis is true.


A Type II Error is committed when a false null hypothesis is not rejected due to
insufficient evidence from a sample. In technical terms, it the ‘acceptance’ of a null
hypothesis when in fact it is false.
The Power of a Test is the probability that the test correctly rejects a false null
hypothesis. The power is defined as 1 − β, where β = P(Type II Error).
Significance level of a Test is the probability of Type I error. Since it is a probability,
the significance level is a number between 0 and 1. It is represented by α and prese-
lected by the researcher (often at 5%).
A Critical Value is the value of a test statistic that divides the sampling distribution
of the test statistic into rejection and acceptance regions depending on a given sig-
nificance level.
The P -value is the probability of obtaining a test result at least as extreme as the
result actually observed, if the null hypothesis is correct. In general terms, the smaller
the P -value the stronger the evidence to reject the null hypothesis.

Exercise 5.5
Do De Veaux, Velleman & Bock, (5th edition), exercises 19.3, 19.11 (c),
19.12 (a), 19.13, 19.15, 19.27, 19.39.

5.3 Closing Comments


We’ve been introduced to the ideas of confidence intervals and hypothesis testing in
this module. The label applied to what we’ve been doing is statistical inference.
The rest of the course involves using the ideas of statistical inference in a variety of
scenarios. Hence we deal with inference concerning two means in Module 6, inference
concerning more than two means in Module 7, inference concerning contingency tables
in Module 8 and inference about regression parameters in Module 9. The only things
which are new are the situations and some formulas. All the basic principles and
steps are the same as we’ve seen in this module.

Note that a Quick Review and additional exercises for elements of Module 5 are given
after Chapter 19 in the text book (after the exercises).
5.30 Module 5. Statistical Inference for One Mean

5.4 Tutorial Module 5


The following section contains Tutorial 5 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1: How well do the students in a Statistics class perform?

It is known from student records that exam scores aggregated across the entire uni-
versity follow an approximately normal distribution with mean of 70 marks and a
standard deviation of 8 marks.

(a) Use this information to estimate the percentage of exam scores that are 78 or
higher.

(b) One of the statistics tutorial groups of 35 students achieves a mean exam score
of over 78 marks. The tutor claims outstanding performance but the lecturer
believes this result is not unusual. Assuming that the 35 students in the tutorial
group represent a random sample of university students, calculate the probability
that the mean exam score of this random sample of 35 exams is greater than 78
marks.

Question 2: Estimating a Mean using a Confidence Interval

I measured the weight of 10 Mars Bars and found the weights to be (all measured in
grams):

61.5, 62, 59, 60, 61, 63, 57, 62, 62, 60.5

Using this data calculate a 95% confidence interval for the population mean weight
of Mars Bars.

(a) Define the parameter that needs to be estimated.

(b) State and check any assumptions that need to be made.

(c) Give the value of the sample statistic.

(d) What is the critical value for this level of confidence (also give the degrees of
freedom)?

(e) Calculate the standard error of the statistic.

(f) Calculate the margin of error for the estimate.


5.4. Tutorial Module 5 5.31

(g) Give the confidence interval.

(h) Give a meaningful statement of the estimate that you have calculated.

Question 3: Investigate the following research question.

Do Mars Bars weigh what they say they weigh?

Using the data in Question 2, test if the mean weight of the Mars Bars is different
from the 60 grams stated by the manufacturers, at the 1% level of significance.

(a) Define the parameter of interest.

(b) State the hypotheses (define any symbols used).

(c) State any assumptions that need to be made.

(d) Assuming the null hypothesis is true, calculate the test statistic.

(e) What are the degrees of freedom.

(f) Give the P -value.

(g) Give a conclusion in the context of the question.

(h) In the context of this problem, what is a Type I error.

(i) In the context of this problem, what is a Type II error.

(j) If the population standard deviation is 2 grams, find the minimum sample size
required in estimating the known population mean with 95% confidence within
0.25 margin of error.

Question 4: Using SPSS in a test of one mean

Using the data in Question 2:

(a) Enter the data into spss.

(b) Using spss, calculate a 95% confidence interval for the population mean weight
of Mars Bars.

(c) Using spss, test if the mean weight of the Mars Bars is different from the 60
grams stated by the manufacturers, at the 1% level of significance.

(d) Compare your answers with those obtained in Questions 2 and 3.


5.32 Module 5. Statistical Inference for One Mean

Question 5: Errors

Discuss the meaning of Type I and Type II error in the context of the following
hypotheses:
H0 : A person is not guilty of a murder
H1 : A person is guilty of a murder

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

5.5 Module 5 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ recognise that a statistic will take different values when sampling is repeated;

ˆ describe the sampling distribution model of a mean;

ˆ state and interpret the Central Limit Theorem (CLT);

ˆ interpret the meaning of a standard error in general and for a sample mean;

ˆ describe the general form of a confidence interval;

ˆ state in non-technical terms what is meant by ‘95% confidence’ or other state-


ments of confidence in statistical reports;

ˆ determine, by hand and using spss, a one-mean t confidence interval;

ˆ carry out ,by hand and using spss, a hypothesis test about one mean using a t
distribution;

ˆ state the assumptions required for a t procedure;

ˆ estimate the P -value for a one or two-sided t-test for a mean;

ˆ assess statistical significance at standard levels using the P -value;

ˆ estimate the size of a sample required to estimate a mean to within a specified


margin of error at a specified level of confidence;

ˆ define the difference between statistical significance and practical importance;


and

ˆ state the meaning of power, Type I and Type II errors.


Module 6
Statistical
Inference for
Comparing
Means
6.2 Module 6. Statistical Inference for Comparing Means

Contents
6.1 Inference for Paired Samples . . . . . . . . . . . . . . . . . . . . 6.4
6.2 Independent Groups . . . . . . . . . . . . . . . . . . . . . . . . . . 6.11
6.3 Parametric and Nonparametric Tests . . . . . . . . . . . . . . . 6.18
6.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.20
6.5 Tutorial Module 6 . . . . . . . . . . . . . . . . . . . . . . . . . . . 6.21
6.6 Module 6 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 6.23
6.3

Module Objectives
On successful completion of this module students should be able to:

ˆ determine by hand and using spss a paired-samples t confidence interval;


ˆ carry out by hand and using spss a paired-samples t-test;
ˆ state the assumptions for and check the appropriateness of the matched pairs t
procedures;
ˆ distinguish between matched pairs data and independent samples data;
ˆ determine by hand and using spss a conservative two-independent-samples t
confidence interval;
ˆ carry out by hand and using spss a conservative two-independent-samples t-
test;
ˆ state the assumptions for and check the appropriateness of the two-independent-
samples t procedures;
ˆ recognise there is often a choice of inferential procedures for a given problem;
and
ˆ describe differences between nonparametric and parametric procedures.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
This module is really just a continuation of the previous one. Now that we’ve seen
the idea of generalising from data to the world at large by using a confidence interval
or doing a hypothesis test, we are ready to apply these ideas in many situations.
Specifically in this module we look at cases involving two groups measured on the
same variable.

De Veaux, Velleman & Bock, (5th edition) covers this material in parts of Chapters 20
and 21.
6.4 Module 6. Statistical Inference for Comparing Means

6.1 Inference for Paired Samples

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 21.

After this reading, write down answers to the following questions:

ˆ How do paired samples arise in an observational study? In an experiment?

ˆ What symbol is used to denote – the mean difference in the data; the mean
difference in the population?

ˆ What symbol is used to denote the standard deviation of the differences?

ˆ What assumptions are required for a paired t procedure?

ˆ What is the null hypothesis?

ˆ What is the formula for the degrees of freedom in a paired t procedure?

There is essentially very little new work in this reading. The idea of blocking and of
matched pairs was briefly discussed in Module 4.

From the analysis point of view, once we can perform a one-sample t-test or produce
a one-sample t confidence interval we can analyse matched pairs data. All we must do
as the first step is calculate the directed differences between the pairs of observations.
‘Directed differences’ means we subtract one member of each pair from the other
member in the same direction for all pairs.)

There are really only the two formulas we need to be able to use and they are just
rewrites of (5.5) and (5.6) as follows:

ˆ To find a confidence interval for the mean difference use the formula
sd
d ± t∗n−1 √ (6.1)
n

ˆ To perform a hypothesis test for the mean difference use the test statistic

d
t= √ (6.2)
sd / n
6.1. Inference for Paired Samples 6.5

Comparing (6.2) with the formula given in Chapter 21, Section 2 of De Veaux, Velle-
man & Bock, (5th edition) (The paired t-test), we see that we have put ∆0 = 0. This
will be the case for all the examples and problems we deal with in this course.

Some problems may arise even though the method is not really new. Firstly we need
to be able to recognise a paired samples experiment. Secondly, we need to take care in
stating the hypothesis or confidence interval of interest on the basis of the information
given.

Recognising Matched Pairs Data


A matched pairs design consists of two samples of subjects or experimental units
matched according to some attribute. Such a design, consisting of n pairs of experi-
mental units, can be represented diagrammatically as follows where each  represents
an experimental course.
Pair 1 2 3 4 ... n
Sample 1     ... 
Sample 2     ... 

The attribute used for matching may be age (in which case each pair of experimental
units is of a similar age), weight, height, ability, status, condition, and so on. Often
the same experimental units, whether they are human beings, animals, plants or
inanimate objects, are used in both samples and are therefore obviously matched.
Typically such experiments are of the before-after kind consisting of measurements
taken before and after the application of or exposure to some condition. For example,
the heart rates of subjects measured before and after performing a physical task, the
weights of calves measured before and after being put on a new diet, or the state of
paint coatings measured before and after exposure to a harsh environmental condition.

The expectation is that, as a result of this matching, the measurements in each pair
will be more alike than they would be if the experimental units in both samples were
chosen completely independently of each other. We would therefore expect the two
samples to be positively correlated or at least have some positive association. A
scatterplot of the measurements in sample 1 against the measurements in sample 2
should reveal such an association. Because of this, matched samples are sometimes
called correlated or dependent samples.
6.6 Module 6. Statistical Inference for Comparing Means

Using SPSS

Here’s an example of a paired samples t test for an experimental study. It is taken


from Moore, The Basic Practice of Statistics.

Example 6.1
We hear that listening to Mozart improves student’s performance on tests.
Perhaps pleasant odours have a similar effect. To test this idea, 21 subjects
worked a paper-and-pencil maze while wearing a mask. The mask was either
unscented or carried a floral scent. The response variable is their average time
on three trials. Each subject worked the maze with both masks, in random
order. The randomisation is important because subjects tend to improve their
times as they work a maze repeatedly. Table 6.1 gives the subjects’ average
times with both masks.

Table 6.1: Floral scents and learning data for Example 6.1.

Unscented Scented
Subject (seconds) (seconds) Difference
1 30.60 37.97 −7.37
2 48.43 51.57 −3.14
3 60.77 56.67 4.10
4 36.07 40.47 −4.40
5 68.47 49.00 19.47
6 32.43 43.23 −10.80
7 43.70 44.57 −0.87
8 37.10 28.40 8.70
9 31.17 28.23 2.94
10 51.23 68.47 −17.24
11 65.40 51.10 14.30
12 58.93 83.50 −24.57
13 54.47 38.30 16.17
14 43.53 51.37 −7.84
15 37.93 29.33 8.60
16 43.50 54.27 −10.77
17 87.70 62.73 24.97
18 53.53 58.00 −4.47
19 64.30 52.40 11.90
20 47.37 53.63 −6.26
21 53.67 47.00 6.67
6.1. Inference for Paired Samples 6.7

This is an example of a matched pairs experiment because each of 21 subjects


is measured under two treatments. To analyse these data we need to establish
the research hypothesis of interest. The key here is to note that the floral scents
are expected to improve the skills of the 21 subjects and decrease their average
time to complete the maze. Hence, if µd is the mean decrease in time required
to complete the maze as a result of the floral scent, the alternative hypothesis
is that µd > 0. We estimate µd by calculating the mean of the differences, time
unscented minus time scented. These time differences are labelled ‘Difference’
in Table 6.1. The mean d and standard deviation sd of the differences are used
in the hypothesis test and confidence interval calculations.
Type the data into spss as shown in the left-hand panel of Figure 6.1. Each
row contains a pair of measurements for each subject. Create a new variable of
differences (using the ‘Transform/Compute’ procedure—see Figure 6.2). The
Data View window should look like the right-hand panel of Figure 6.1.
Before doing a t procedure, check the assumptions.
Paired data assumption: The data are paired by subject.
Randomisation condition: Assume the subjects are representative of all stu-
dents of interest.
Normality assumption: The histogram of differences is roughly unimodal and
symmetric (see Figure 6.3).
10% condition: The 21 subjects represent less than 10% of all students of in-
terest.
Independence: one subject’s responses are independent of another subject’s
responses.
Call up the paired samples t procedure under ‘Analyze/Compare Means/Paired
Samples T Test’ (see Figure 6.4). The output from this procedure is shown in
Figure 6.5.
A 95% confidence interval for µd is (−4.7551, 6.6684) from Figure 6.5. In other
words, the mean decrease in time required to complete the maze as a result of
the floral scent is between −4.8 sec and 6.7 sec, with 95% confidence.
Calculating this confidence interval by hand we find: d = 0.957 sec, sd =
12.548 sec, n = 21. Using (6.1), a 95% confidence interval for µd is
sd 12.548
d ± t∗n−1 √ = 0.957 ± 2.086 √
n 21
= 0.957 ± 5.712
= (−4.75, 6.67)

as found from spss.


To test H0 : µd = 0 against HA : µd > 0 we notice from the 95% confidence
interval that there’s not very good evidence in support of the alternative. This
6.8 Module 6. Statistical Inference for Comparing Means

Figure 6.1: Data for Example 6.1 (i) as entered into spss and (ii) after computing
differences.

is confirmed by the spss output in Figure 6.5. The two-sided P -value is 0.730.
Because Ha is one-sided we need to halve the two-sided P -value to give the one-
sided P -value of 0.365 or 37%. We conclude that there is insufficient evidence
to support the claim that floral scents improve performance, i.e., the sample
mean difference in times of 0.957 sec is very likely due to sampling variation.
The mean improvement in performance is small, only about 1 second over the
50 or so seconds that subjects took wearing the unscented mask. This small
improvement is not statistically significant at even a 25% level of significance.
Doing this test by hand we use the test statistic (6.2). This gives

d
t= √
sd / n
0.957
= √
12.548/ 21
= 0.35

with df = 21. Table T shows that 0.35 is less than the critical t value with one-
sided probability of 0.10. The P -value is therefore greater than 0.10, leading
to the same conclusion as given above.
6.1. Inference for Paired Samples 6.9

Figure 6.2: spss dialog box for computing differences.

Figure 6.3: Histogram of differences for Example 6.1.


6.10 Module 6. Statistical Inference for Comparing Means

Figure 6.4: Dialog box for the Paired Samples procedure in spss.

Figure 6.5: spss output from the Paired Samples procedure.


6.2. Independent Groups 6.11

Exercise 6.1
Watch the following spss video.

ˆ Two Related Samples t-test (including checking the assumptions)

Exercise 6.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 21.23 (a) & (b)
(use spss), 21.29 (use spss), 21.31 (by hand and using spss).

6.2 Independent Groups


The data from independent samples are dealt with differently from paired samples. If
the wrong method is applied in assignments, the penalties may be quite severe. We
therefore need to be clear about the difference between these two types of designs.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 20 (Sec-
tions 4 & 5).

After this reading, write down answers to the following questions:

ˆ What type of plot is recommended to display the data from the two independent
samples?

ˆ What assumptions are needed to use Student’s t-models with two independent
samples?

ˆ What is the formula for the two-independent-samples t confidence interval for


the difference in means?

ˆ Write down the formula for the two-independent-samples t-test for the difference
in means.

ˆ How many degrees of freedom are used in Student’s t-models with two indepen-
dent samples following the conservative approach?
6.12 Module 6. Statistical Inference for Comparing Means

This reading essentially covers two procedures:

ˆ Finding a confidence interval for the difference µ1 −µ2 between two means using
the formula s
s21 s2
(y 1 − y 2 ) ± t∗df + 2 (6.3)
n1 n2
with df given by one less than the smaller of n1 and n2 , that is, smaller of n1 − 1
and n2 − 1.

ˆ Performing a hypothesis test for the difference between two means using the
test statistic
y −y
t = q 12 2 2 (6.4)
s1 s2
n1 + n2
with df given by one less than the smaller of n1 and n2 , that is, smaller of n1 − 1
and n2 − 1. (This formula for the df is what the book calls the ‘easy rule’.)

Now that we are dealing with two samples, which we have denoted by the subscripts
1 and 2, we should make sure we state which sample is which in a problem. So, for
example, we might define µ1 as the mean height of males and µ2 as the mean height
of females (or, better still, use µm for males and µf for females). Also, if we are asked
for a confidence interval for the difference in means, we must state whether it’s for
µ1 − µ2 or for µ2 − µ1 , using (6.3) appropriately of course.

What about the Assumptions?


In practice we should always check all the assumptions before carrying out any pro-
cedure. For a t procedure this amounts to randomness for each sample, normality for
each population from which the samples are drawn, independence between the two
samples (if there are two samples) and the 10% condition.

The randomness assumption usually relies on us being told or knowing that each
sample was either a srs (or could be treated as such) if the study was observational,
or involved randomisation if the study was an experiment. In practice we would
rely on information concerning the data collection process, remembering that simple
random sampling in an observational study and randomisation in an experiment are
descriptions of the protocol of the study, not the data itself.

Similarly, independence of two samples must often be taken on faith. More about this
in Section 6.3 though, where studies in which the two samples are not independent
are considered. The 10% condition is seldom an issue.
6.2. Independent Groups 6.13

Checking normality requires judgement, and opinions may vary. The discussion on
assumptions and conditions in Section 4 (Chapter 20) of De Veaux, Velleman & Bock,
(5th edition), should be noted. The basic gist is that for t procedures, the importance
of normality diminishes as sample size increases and for n larger than 40 (often 25
is enough as suggested in Section 5.1.3), the assumption can be pretty well ignored.
Although looking at boxplots and normal probability plots (P-P plots) is helpful,
with small samples it is usually not possible to decide from the data alone whether
or not normality is acceptable. We must rely on previous knowledge concerning the
distribution of the variable under consideration and should include a statement to
this effect in the conclusion, or, if considerable doubt exists about the validity of the
normality assumption, seek alternative analyses which do not rely on normality. Such
methods are called nonparametric procedures (see Section 6.3 for more details).

Having said all this, it’s worth repeating that in the assignments you only need to
discuss assumptions if asked explicitly to do so. For example, if a two-independent-
samples t-interval is required and (6.3) is appropriately applied, no marks will be
lost for not mentioning or checking assumptions unless specifically requested. In the
practice of Statistics however, whether it be in the workplace, in other
university studies or elsewhere, it is unwise and potentially dangerous to
make use of any statistical procedure without checking the assumptions.

Using SPSS

Here spss is applied to an exercise from the text. The data for this exercise can be
found by looking at 20.66 (Chapter 20, Question 66).

Example 6.2
Read Question 66 in Chapter 20 of De Veaux, Velleman & Bock, (5th edition).
The data as provided is in two columns, part of which is shown in the left-hand
panel of Figure 6.6. The first step is to reformat the Data View so all the skull
measurements are in just one column. A simple cut and paste achieves that.
Now a second column is inserted with a coding to represent the date of each
of the skull measurements. In the right-hand panel of Figure 6.6 part of the
reformatted data is shown with 1 representing 4000 B.C.E and 2 representing
200 B.C.E. Notice how the data is now in standard format with each of the
60 rows representing an individual skull and the two columns representing the
variables ‘breadth’ and ‘date’.

(a) Randomisation assumption: We assume the skulls measured are represen-


tative of all Egyptians at the time.
6.14 Module 6. Statistical Inference for Comparing Means

Normality assumption: The histograms/dotplots/boxplots of the two sam-


ples (see Figure 6.7 are unimodal and approximately symmetric and with
30 in each group give no concerns about this assumption. Also included
in this figure for good measure are the Q-Q plots for the two groups. (Any
of these plots demonstrate the adequacy of this assumption but it takes
little effort to produce them all!)
Independent groups assumption: The skull breadth of Egyptians in 4000
B.C.E. are independent of those in 200 B.C.E.
We conclude that two-independent-samples inference using t procedures
is appropriate.
(b) The ‘Analyze/Compare Means/Independent-Samples T Test’ procedure in
spss produces the dialog box shown in Figure 6.8. Two test results and two
confidence intervals are reported (see Figure 6.9). There is little difference
between them but the second row labelled ‘Equal variances not assumed’
is relevant for us. A 95% confidence interval for µ4K − µ200 is (1.88, 6.66)
where µ200 is the 200 B.C.E. mean and µ4K the 4000 B.C.E. mean. We
are therefore 95% confident that Egyptian males in 200 B.C.E. had a
mean skull breadth between 1.88 and 6.66 mm larger than the mean skull
breadth of Egyptian males in 4000 B.C.E. (assuming the measurements
are in mm).
(c) We are testing H0 : µ4K −µ200 = 0 against HA : µ4K −µ200 6= 0. From the
second row of the spss output, for this test t = 3.58 with 54.973 df. The
P -value (two-sided) is 0.001 which is strong evidence of a change in mean
skull breadth over the time period, i.e., the difference in sample mean
skull breadths of 4.266 mm is very unlikely to be due to chance/sampling
variation. We could make the same conclusion by observing that the
confidence interval in part (b) is completely above zero.

What if we did parts (b) and (c) of this example on a calculator?

(b) From formula (6.3) we get


s
s24K s2
(y 4K − y 200 ) ± t∗df + 200
n4K n200
r
4.038462 5.129252
= (135.633 − 131.367) ± 2.045 +
30 30
= 4.266 ± 2.437
= (1.83, 6.70)

where the easy rule for df has been used giving df = 29 and t∗29 = 2.045
from Table T. This is a little wider than what spss produces because the
easy rule for df is conservative.
6.2. Independent Groups 6.15

Figure 6.6: spss data from 22.66 before and after reformatting.

(b) From formula (6.4) we get for the test statistic


y − y 200
t = q 4K
s24K s2200
n4K + n200
135.633 − 131.367
=q
4.038462 2
30 + 5.12925
30
4.266
=
1.1919
= 3.58

which is the same as that given by spss. We can estimate the P -value
from Table T as follows. We have df = 29 and a two-sided test. From
Table T we see that 3.579 is larger than the largest entry (2.756) for 29 df
and therefore, because 2.756 corresponds to a (two-sided) P -value of 0.01,
3.58 corresponds to a (two-sided) P -value less than 0.01. Our conclusion
is the same as from the spss output.

spss needs the data to perform a procedure. This means that problems in which only
summary statistics (means, standard deviations, sample size, etc) are given cannot
be done using spss and must be done by hand using a calculator.

In assignment answers, when doing a question by hand by applying formulas (6.3)


or (6.4), use the ‘easy’ df formula. Although the book makes the point that the ’easy’
df method is conservative (i.e., the margin of error of a CI will be larger and P -value
larger), being conservative in statistics is usually a good thing because quite often
there is uncertainty about the validity of the assumptions. spss automatically applies
the more complex formula for df and as such does not usually match the ‘easy’ df
and is not necessarily a whole number.
6.16 Module 6. Statistical Inference for Comparing Means

Figure 6.7: Various graphical checks of the normality assumption in spss.


6.2. Independent Groups 6.17

Figure 6.8: Dialog box for the ‘Independent-Samples T Test’ procedure in spss.

Figure 6.9: Output from the ‘Independent-Samples T Test’ procedure in spss.


6.18 Module 6. Statistical Inference for Comparing Means

Exercise 6.3
Watch the following spss videos. After watching each video try to
replicate each one.

ˆ Two Independent Samples t-test

Exercise 6.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 21.9, 20.57,
20.59, 20.65 (use spss), 20.67 (by hand and using spss), 20.71, 20.73.

6.3 Parametric and Nonparametric Tests


In Statistics all the procedures we use rely on assumptions, and some procedures
need more assumptions than others. One of the things we’ve been emphasising is
that assumptions should be checked. Not much has been said about what to do if
assumptions fail.

Sometimes the best thing to do would be to discard the data altogether and do no
analysis at all! This would be the case if we don’t believe the data is representative
of the population or model we’re wanting to describe. It might be that a survey has
been run with voluntary respondents for example—like those ones TV channels are
keen on promoting. Unless there’s no reason to believe a relationship exists between
the variable(s) of interest and whether or not a person is likely to respond to such
a survey, the data from these types of surveys are of no value (except of course
financially to the TV and telephone companies!).

For example, it might be possible to argue that the height of participants is unlikely
to be associated with whether or not a person takes part, so these heights might be
representative of a wider audience. But height is hardly going to be the focus of
attention in such a survey! Interest is likely to be on attitudes, opinions, perceptions,
popularity, etc., and these are exactly the variables that will be associated with the
likelihood of a person responding. Typically, for example, people who feel strongly
about an issue respond, whereas those who are indifferent do not. Also, of course,
there’s an undercoverage issue in that those who are most likely to be exposed to the
survey (i.e., watching that TV channel at that time) may not be representative of the
population of interest. The upshot is that the statistics reported from these surveys
are essentially worthless as a gauge on public opinion or the like.
6.3. Parametric and Nonparametric Tests 6.19

Assuming however that representative samples are obtained in observational studies


and randomisation is applied appropriately in experiments, how do we decide how
to analyse the data? Well, that’s where having a choice of procedures is useful.
If we look through spss we see that there are lots of techniques available. Under
‘Analyze/Nonparametric Tests’ for example, there are a number of choices in the
Settings tab under ‘Independent Samples’ and ‘Related Samples’, some of which are
alternatives to t procedures.

The point is that if an assumption like normality for a t procedure is looking a bit
dodgy, other procedures are available. There are lots of them. We still assume
a representative or random sample from the population of interest and the data
points are independent of those for any other in the population. However, we are
not restricted by the assumption that the values in the population have a normal
distribution.

In general nonparametric tests are not as powerful as their parametric counterparts.


For quantitative data, nonparametric tests seldom make best use of the data. And
data is a valuable commodity, often gathered at great cost. Best then to analyse data
in the best way available, which in statistical jargon means using the most powerful
methods available—methods that are more likely to reject the null hypothesis when
it should be rejected, i.e., correctly reject it when it is in fact false.

As a rule, t procedures which rely on normality are the most powerful available for the
types of questions they answer. For quantitative data there is a strong compulsion
to use procedures relying on normality. Sometimes, as a result they get used when
they shouldn’t, when normality fails or when the data is not even quantitative.

How do we know when the normality assumption fails? We usually don’t have access
to information about the whole population, so we can only use our sample data to
check the normality assumption, by plotting the data. t procedures are fairly robust
when sample sizes are large, i.e., when the sample size is large, the sampling distri-
bution of mean differences will be approximately normal, even if the distribution of
differences from the sample is not quite normal. Provided the distribution of the data
is not too skewed nor contains outliers, the t procedure may be appropriate. However,
small sample size may increase vulnerability to violations of the normality assump-
tion because the sampling distribution of the mean differences may not necessarily
be normal in these cases. In summary, with quantitative data where a t procedure
would usually be preferred, we take the conservative/cautious approach that, if n is
small, it is safer to use a nonparametric test.

The word ‘nonparametric’ is used to describe inferential procedures which have weak
assumptions, none as strong as normality. These methods don’t usually involve means
and standard deviations but rely on counts or ranks, and so can work on data that is
not quantitative. Hence ordinal or categorical data is dealt with using nonparametric
methods.
6.20 Module 6. Statistical Inference for Comparing Means

Procedures which assume normality (or some other assumption about the exact shape
of a distribution) are called parametric. The t procedures are classified as being
parametric.

Obviously in the time available for this course we can only look at some of the many
techniques. Although practical guidelines are given for using t procedures in the
presence of outliers and non-normality, considerable judgement is needed in applying
them. It is possible that two analysts may disagree on the appropriateness of a t
procedure for a given set of data.

The principle of conservatism however usually results in similar conclusions being


drawn. In the presence of outliers this means repeating an analysis with and with-
out outliers included and reporting the most conservative result as described in the
Reading. If skewness or other non-normality appears to be a problem and there is
some doubt about the validity of a t procedure, other procedures are available to
corroborate a particular inference. However, remember that the results of an inferen-
tial procedure should come as no surprise. A well-chosen plot of the data will often
reveal the likely conclusion in advance of the inferential analysis. While these other
procedures are outside the scope of this course, a diagram illustrating nonparametric
tests equivalent to parametric tests is given as a reference in Section 10.4.

6.4 Closing Comments


We have further extended the ideas about confidence intervals and hypothesis testing
in this module. The only things which are new are the situations and some formulas.
All the basic principles in this module are the same as we’ve seen in the previous
module. The next module extends these same concepts to more complex situations.
6.5. Tutorial Module 6 6.21

6.5 Tutorial Module 6


The following section contains Tutorial 6 - a good summary of the work learnt in this
module and most importantly, testing your knowledge.

Question 1: Investigate the following research question

Do you get lower fuel consumption (i.e. fewer litres/100 km) using premium unleaded
petrol?

Many drivers of cars that can run on regular unleaded petrol actually buy premium
in the belief that they will use less fuel. To test that belief, 10 cars were tested in
which all the cars regularly run on regular unleaded petrol. Each car is filled first with
either regular or premium unleaded petrol, decided by a coin toss, and the number of
kilometres for that tank-full recorded. Then the fuel consumption in litres/100km is
recorded again for the same car for a tank-full of the other kind of unleaded petrol.
The drivers do not know about the experiment. The results are as follows (litres/100
km):

Car 1 2 3 4 5 6 7 8 9 10
Regular 14.7 11.8 11.2 10.7 10.2 10.7 8.7 9.4 8.7 8.4
Premium 12.4 10.7 9.8 9.8 9.4 9.4 9.1 9.1 8.4 7.4
Difference (R-P)

(a) Perform a hypothesis test to see if cars get on average lower fuel consumption
with premium unleaded petrol.

Step 1: Write the hypotheses (define any symbols used).


Step 1a: What assumptions are needed to perform this test?
Step 2: Calculate the test statistic:

ˆ First calculate the differences between the two samples (define the way the
difference was obtained);
ˆ Second calculate the mean of the differences;
ˆ Third calculate the standard deviation of the differences;
ˆ Fourthly calculate the standard error of the mean of the differences;
ˆ Fifthly give the degrees of freedom, df;
ˆ Finally calculate the test statistic.

Step 3: Calculate the P -value.


Step 4: Write a conclusion in context.
6.22 Module 6. Statistical Inference for Comparing Means

(b) Find a 95% confidence interval for the population mean of differences in fuel
consumption.

(c) Enter the data into spss and repeat the hypothesis test. Compare the test statis-
tic, degrees of freedom and P -value with those obtained in Part (a).

Question 2: Investigate the following research question


Do Mars and Snickers Bars (both labeled as 60g) differ in their weights?

Independent random samples of Mars and Snickers Bars were weighed (in grams) to
test if the population mean weights of the two bars differ.

1 2 3 4 5 6 7 8 9 10
Mars 61 62 63 60 59 61 62 63 64 58
Snickers 57 58 59 61 60 59 62 59 58 57

(a) Do the population mean weights of the two bars differ?


Step 1: Write the hypotheses (define any symbols used).
Step 1a: What assumptions are needed to perform this test?
Step 2: Calculate the test statistic:

ˆ First calculate the sample mean for each sample;


ˆ Second calculate the sample standard deviation for each sample;
ˆ Third calculate the sample statistic for the test;
ˆ Fourth calculate the standard error of this statistic;
ˆ Fifth calculate the degrees of freedom, df;
ˆ Finally calculate the test statistic.

Step 3: Calculate the P -value.


Step 4: Write a conclusion at the 5% level of significance.

(b) Find a 95% confidence interval for the difference of the two population means.

(c) Enter the data into spss and repeat the hypothesis test. Compare the test statis-
tic, degrees of freedom and P -value with those obtained in Part (a).

Question 3
“Absence of evidence is not the same as evidence of absence”. Discuss this statement
in relation to the concept of hypothesis testing.

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
6.6. Module 6 Checklist 6.23

6.6 Module 6 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ recognise a paired samples design;

ˆ denote the mean and standard deviation of the differences in a paired t proce-
dure;

ˆ state the assumptions required for a paired t procedure;

ˆ dtermine a confidence interval for the population mean of the differences of


paired data;

ˆ perform a hypothesis test for the mean difference of paired data (by hand and
using spss);

ˆ determine the degrees of freedom in a paired t procedure;

ˆ use a conservative two-independent-samples t procedure to obtain a confidence


interval at a stated level of confidence for the difference between two means (by
hand and using spss);

ˆ perform a conservative two-independent-samples t-test for the hypothesis that


two populations have equal means against either a one-sided or a two-sided
alternative (by hand and using spss);

ˆ state the assumptions required for a two-independent-samples t procedure;

ˆ recognise a two-independent-samples design and be able to distinguish it from


one sample and matched pairs designs;

ˆ recognise when it is appropriate to use a nonparametric test rather than a


parametric test.
Module 7
Analysis of
Variance
(ANOVA)
7.2 Module 7. Analysis of Variance (ANOVA)

Contents
7.1 One-way Analysis of Variance (ANOVA) . . . . . . . . . . . . . 7.3
7.2 One-way ANOVA and spss . . . . . . . . . . . . . . . . . . . . . . 7.9
7.3 Two-way Factorial ANOVA . . . . . . . . . . . . . . . . . . . . . 7.15
7.4 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.22
7.5 Tutorial Module 7 . . . . . . . . . . . . . . . . . . . . . . . . . . . 7.24
7.6 Module 7 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 7.28
7.1. One-way Analysis of Variance (ANOVA) 7.3

Module Objectives
On successful completion of this module students should be able to:

ˆ recognise when to use an ANOVA to compare means among groups;

ˆ state the assumptions associated with performing ANOVA analysis;

ˆ carry out a one-way ANOVA in spss;

ˆ read an ANOVA table, including identifying df, MSS and the F -statistic;

ˆ interpret the output for one-way and a two-way ANOVA and subsequent pair-
wise t-tests for significant effects.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
In this module we look at situations where we want to compare the means of three
or more groups.

We cover Chapters 25 and 26 of De Veaux, Velleman & Bock, (5th edition) in this
module.

7.1 One-way Analysis of Variance (ANOVA)


In Module 6, we used an independent samples t-test to compare whether two popula-
tions had equal means. It would be expected that there would be some difference in
the mean values just because they are calculated from two different samples of data.
By using a t-test we can determine if the difference we find is larger than what we
could expect due to just random natural variation between individuals and, therefore,
could indicate that the samples are drawn from two different populations.
7.4 Module 7. Analysis of Variance (ANOVA)

In many situations we may be interested in comparing the means of three or more


groups. Analysis of Variance (ANOVA) is a statistical method that allows us to
compare multiple group sample means to determine if they come from the same or
different underlying populations.

If we are comparing group means why is the test called Analysis of Variance?

Have a look at Figures 25.2 and 25.3 in the De Veaux, Velleman & Bock, (5th edi-
tion). Each boxplot represents a different group (sometimes called ‘treatments’). The
ANOVA compares the variances between the group means to the variances within the
groups.

In Figure 25.2 the over-lapping spread of the boxplots tells us that the variances
within the groups are larger than the variance among the group means. It would be
very difficult for us to tell whether any differences among the group means (31, 36,
38 and 31 respectively) are due to the groups representing different populations, or
whether the means vary just because of natural sampling variation.

In Figure 25.3 however, the variance within the groups is much smaller than the
variance among (or between) the group means. This could mean that the differences
among the means are more likely due to them representing different populations
rather than differences due to sampling.

In Figure 25.3 knowing the mean of each group would help us distinguish the obser-
vations in one group from the observations in another group.

In Figure 25.2 however, the mean of each group does not help us distinguish the
observations in one group from the observations in another group.

ANOVA is a significance test of a null hypothesis that states that all population
means are equal. This means the null hypothesis asserts that the groups represent
samples from the same overall population (or that we do not have enough evidence in
our samples to suggest otherwise). The alternative hypothesis is that the population
means are not all equal.

This alternative hypothesis does not necessarily mean that all the means are signifi-
cantly different to each other. It means only that at least one of the means is different,
indicating that at least one of the groups is sampled from a different underlying pop-
ulation. This also means that if we reject the ANOVA null hypothesis we will need
to perform additional analysis using t-tests, or similar pairwise tests, to determine
which groups are different to each other.

If we have k groups, then each group has a corresponding population of subjects, and
the means of the response variable for the k groups can be denoted by µ1 , µ2 , ..., µk .
Since we are trying to compare means of several groups, the null hypothesis, as always,
assumes that there is no difference between the groups we are investigating, that is:
H0 : µ1 = µ2 = ... = µk .
7.1. One-way Analysis of Variance (ANOVA) 7.5

The alternative hypothesis then becomes:

Ha : at least one of the population means is different.

Note that no matter what the specific context of the data we are analysing, we use
an ANOVA when we want to investigate and compare the group means and we use
the variances within and between the groups to determine whether we should accept
or reject the null hypothesis.

In order to statistically compare multiple means in the ANOVA we need to use a new
sampling distribution model, the F distribution model. The associated F -test allows
us to compare the variation (differences) between the means of the groups with the
variation within the groups. When the differences or variation between the means are
large compared with the variation within the groups, we reject the null hypothesis
and conclude that at least one mean is different to the others. Please note that in this
course, you will not be required to look up the F -tables. Instead, you will interpret
spss output.

Remember, as previously mentioned, although we are comparing means, the method


is called Analysis of Variance because the test statistic we calculate uses evidence
about two types of variability: the variability between groups and the variability
within groups.

It can be helpful to understand a bit more about how we calculate the within and
between variance components. When our groups have similar variance (σ 2 ) measures
we can pool all the within group variance estimates together to get an overall esti-
mate of the σ 2 . Because it is a pooled variance, it is denoted s2p . This quantity is
traditionally called the Error Mean Square, or the Within Mean Square and
is denoted by MSE . The error mean square is an estimate of the error or residual
variance and represents the amount of random variation within treatments/groups.

Remember though, that we also have a separate estimate of σ 2 from the variation
between the means of the treatments/groups we are investigating. This quantity is
called the Treatment Mean Square, or the Between Treatment Mean Square
and is denoted by MST .

If the null hypothesis is true, then the group means are equal and both MST and
MSE estimate σ 2 . Their ratio then, should be close to 1.0. If the null hypothesis is
false, then the MST will be larger because the group means are not equal. The MSE
is a pooled estimate in which the variation within each group is found around its
own group mean, so differences between means won’t inflate MSE . This makes the
ratio MST /MSE perfect for testing the null hypothesis. When the null hypothesis is
true, the ratio should be near 1. If the group means really ARE different, then the
numerator MST will tend to be larger than the denominator MSE and the ratio will
tend to be bigger than 1.
7.6 Module 7. Analysis of Variance (ANOVA)

The ratio of the treatment mean square MST to the error mean square MSE is
called the variance ratio or the F -statistic or the F -ratio. Under H0 , given all
assumptions are true, this variance ratio has an F-distribution.

Even when the null hypothesis is true, the variance ratio will rarely equal 1 exactly
due to natural sampling variability. So how can we tell when the ratio is big enough to
reject the null hypothesis? The answer lies with the F -distribution – the distribution
of MST /MSE (the F -statistic), and by comparing this statistic calculated from our
data, with the appropriate F -distribution we can obtain a P -value.

The resulting F -test is one-tailed because any differences in the means make the
F -statistic larger. Larger differences in the treatment effects lead to the means being
more variable, making MST bigger. This makes F increase, so the test is significant
if the F -statistic is large enough. In practice, larger F -statistic values have smaller
P -values, because there is a smaller probability of getting a large F -statistic if H0 is
true.

F -models depend on two degrees of freedom parameters: the degrees of freedom


come from the two variance estimates. The Treatment Mean Square, MST , will have
k groups and, therefore, this variance has k − 1 degrees of freedom. The Error Mean
Square, MSE is the pooled estimate of the variance within the groups. If there are n
observations in each group, we get n − 1 degrees of freedom from each group, for a
total of k(n − 1) degrees of freedom. Alternatively, we can use the total sample size
minus the number of groups, N − k.

Therefore overall we say that the F -statistic is:


Between group variability Treatment Mean Square M ST
F = = =
Within group variability Error Mean Square M SE
with k − 1 and N − k degrees of freedom.

Assumptions for ANOVA


Before doing an ANOVA, we should consider the assumptions.

Independence Assumption: The groups must be independent of each other. The data
within each treatment must be independent as well, i.e. each case must be independent
of all others.

Randomisation Condition: Ensure the data was collected randomly within each group
or that appropriate randomisation was applied when allocating cases to treatments/
groups.
7.1. One-way Analysis of Variance (ANOVA) 7.7

Equal Variance Assumption: The ANOVA requires that the variances of the treat-
ments/groups be equal (or equal enough), i.e. homogeneity of variance.

Normal Population Assumption: Check Normality with a histogram. Check for out-
liers in the boxplots of the values for each group.

The ANOVA table


The ANOVA table can be constructed by hand, or a partially completed ANOVA table
can be finalised when you understand the relationships between different elements of
the table. This is because the format for a one-way ANOVA table is always the same:

Source of Degrees of Sum of Mean Square F-statistic


Variation Freedom Squares

SST M ST
Between k−1 SST M ST =
k−1 M SE

SSE
Within N −k SSE M SE =
N −k
Total N −1 TSS

ˆ as long as we know the number of groups k and the total sample size N we
can calculate all df components. The Total df should also equal the sum of the
Between and Within df.

ˆ SST and SSE are calculated from the data.

ˆ the TSS should be equal to the SST + SSE.

ˆ MST and MSE are based on their respective SS and df.

ˆ the F-statistic is always the MST divided by MSE.

Example 7.1
An experiment to determine the effect of different post-surgery care plans
on the swelling in postsurgical patients recorded the hand volume changes for
7.8 Module 7. Analysis of Variance (ANOVA)

patients who had been randomly assigned to one of the care plans (treatments).
The ANOVA for the data is as follows:

Source of Degrees of Sum of Mean F-statistic P-value


Variation Freedom Squares Square

Treatment 2 716.17 358.08 7.41 0.0014


Error 56 2704.38 48.29
Total 58 3420.54

We are not given a lot of information here about the experiment, but there is
quite a bit we can figure out from the ANOVA table just by knowing how each
part is calculated.
From the ANOVA table we can see that there must have been a total sample
size of 59 because N − 1 = 58 and there must have been 3 treatments because
k − 1 = 2 (but we don’t know what the treatments were). We can also see that
the TSS (Total Sum of Squares) of 3420.54 is the sum of the SST (Treatment
Sum of Squares) and SSE (Error Sum of Squares), i.e. 716.17 + 2704.38 =
3240.54. The ratio of the Mean Square values (358.08/48.29) equals the F -
statistic of 7.41.

What does the ANOVA say about the results of the experiment? Specifically,
what does it say about the null hypothesis?
The F -statistic of 7.41 has a P -value that is quite small (0.0014) and much
less than α = 0.05. We can reject the null hypothesis that the mean change in
hand volume is the same for all three treatments. In order to determine which
treatments were different to each other we would need to perform pairwise t-
tests or similar. [Note, you may see P -value, p-value or Sig. used in the text
book, spss output and these notes – they are all referring to the same thing].

If the ANOVA indicates the treatment effect is significant and we will need to
use t-tests to determine which treatments are different to each other, why do
we bother with the ANOVA? Why not just do the t-tests?
The ANOVA is a more statistically ‘powerful’ test. This means that it is less
likely to find a significant result by mistake. In any test we perform, we accept
that there is a certain probability (α) that we are rejecting H0 by mistake
(remember Type I error).
Using a single test (ANOVA) rather than multiple tests allows us to control
the probability of a Type I error, at least at the initial stages of our analysis.
With a significance level of α = 0.05 in the F -test, the probability of incorrectly
rejecting a true H0 is 0.05. When we do a separate t-test for each pair of means,
7.2. One-way ANOVA and spss 7.9

a Type I error probability applies for each comparison and for three treatments
that would need three t-tests to compare all pairs of treatments.
In the example above with 3 treatment groups there would be 3 pairwise com-
parison tests that we would need to perform (treatment A v B, A v C and B
v C). In that case, we are not controlling the overall Type I error rate for all
the comparisons.
If our ANOVA is not significant we should not proceed with applying t-tests or
similar. If we performed the t-tests without first checking the ANOVA there is
a chance that at least one of the pairwise t-tests would result in a Type I error.
Performing the ANOVA first protects us from these potential mistakes.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 25.

Following this reading, keep in mind these caveats in relation to this course:

ˆ you will not be required to look up the F -tables; instead you will be required
to interpret spss output.

ˆ you will not be required to calculate Sums of Squares by hand, however you
will need to understand how to calculate the df, MSS and F-statistic in the
ANOVA table based on other information provided (i.e. by understanding the
relationships between these values as shown above).

7.2 One-way ANOVA and SPSS


Example 7.2
The eight steps below show you how to analyse your data using a one-way
ANOVA in spss. At the end of these eight steps, we will interpret the results
from this test.
The effect of a new forest harvesting system on the damage to adjacent trees
is of interest to forestry managers. There are 4 Harvest system treatments: no
guidance system (nil); current system 1 (CS1); current system 2 (CS2) and a
new system (new). The measurement taken is the percentage of trees within
two metres of the cut down tree, which have some form of damage. The spss
file forestry [Link] is available on the Study Desk.
7.10 Module 7. Analysis of Variance (ANOVA)

harvest system individual measurements


nil 78.65 95.67 78.52 97.74 79.57
CS1 62.81 54.69 45.64 52.43 71.66
CS2 45.83 36.58 59.92 42.25 45.05
new 15.89 35.01 38.38 19.82 40.93

1. Click Analyze > Compare Means > One-Way ANOVA

2. You will be presented with the One-Way ANOVA dialogue box:

3. Transfer the dependent variable, percent damage, into the Dependent List:
box and the independent variable, Harvest system, into the Factor: box
using the appropriate right arrow buttons (or drag-and-drop the variables
into the boxes), as shown below:
7.2. One-way ANOVA and spss 7.11

4. Click on the Post hoc button. If the ANOVA is significant and we reject H0
then we will need to perform post-hoc pairwise tests to determine which
treatments are different to each other. There are a range of different tests
we can choose from. We will choose Tukey’s test as it includes a ‘correction
factor’ to account for the multiple tests it performs using the same group
data. Tick the Tukey checkbox as shown below:

5. Click on the Continue button.


6. Click on the Options button. Tick the Descriptive checkbox in the ‘Statis-
tics’ area, as shown below:

7. Click on the Continue button.


8. Click on the OK button.
7.12 Module 7. Analysis of Variance (ANOVA)

SPSS Output
The Descriptives table (see below) provides some very useful descriptive
statistics, including the mean, standard deviation and 95% confidence inter-
vals for the dependent variable (percent damage) for each separate group (nil,
CS1, CS2 and new), as well as when all groups are combined (Total). These
values are useful when you need to describe your data.

The ANOVA table shows the output of the ANOVA analysis and whether
there is a statistically significant difference between our group means. The
hypotheses for this test would be:
H0 : µnil = µCS1 = µCS2 = µnew .
Ha : at least two of the population means are unequal.
where µ is the mean percentage of trees within two metres of the cut down tree
which have some form of damage.
The F -statistic is quite large at 27.920. Remember that if H0 is true and there
is no statistical difference between the group means then we would expect
the F -statistic to be close to 1 (because F is the ratio of the between group
and within group MSS and if they are the same then F =1) We can see that
the significance value (P -value) is 0.000 (i.e. p < 0.001), which is below 0.05
and, therefore, there is a statistically significant difference in the mean percent
damage to neighbouring trees resulting from the different harvesting systems.

This is great to know, but we do not know which of the specific groups differed.
7.2. One-way ANOVA and spss 7.13

We can find this out in the Multiple Comparisons table which contains the
results of the Tukey post hoc tests.
The table below, Multiple Comparisons, shows which groups differed from
each other. The Tukey post hoc test is generally the preferred test for conduct-
ing post hoc tests from a one-way ANOVA, but there are many others. The
Sig. column gives us the P -value.
The no guidance (nil) group compared to current system 1 (CS1), current
system 2 (CS2) and new system (new) has P -values of 0.002, 0.00, and 0.00
respectively. As each of these P -values is < 0.05 we would reject H0 for all
three tests and conclude that the mean percent damage using the nil treatment
is significantly different to all other treatments. There is also a significant
difference (i.e. p < 0.05) in percent damage between the CS1 treatment and
the new treatment where p = 0.003. However, CS2 is not significantly different
to CS1 or the new system (p > 0.05 in both comparisons).
By looking at the Mean Difference column we can also see that the significant
difference between new and nil (no guidance) of −56.02 is based on the cal-
culation of mean percent damage in the new group minus the mean percent
damage in the nil group. As this is a negative difference, the percent damage
in the new group must have been less than in the nil group.
Our conclusion, therefore, is not only that there was a significant difference
between the new and nil groups, but that there was significantly less damage
in the new system group. The difference between the new and CS1 groups was
also negative, but less so (difference = −27.44). Therefore, the new system
did perform significantly better than CS1 (significantly less damage using the
new system), however, there was not as much improvement as found when
compared to the nil group.
7.14 Module 7. Analysis of Variance (ANOVA)

When reporting the ANOVA results we need to provide the Between and
Within group df, the F-statistic and the P -value. When reporting the re-
sults of the post-hoc Tukey tests we should provide the P -values and also some
summary information describing each group – usually the mean and standard
deviation (you can find these in the Descriptives Table in the spss output
above).
Based on the results above, you could report the results of the study as follows:
There was a statistically significant difference between groups as determined by
one-way ANOVA (F (3, 16) = 37.92, p < 0.001). A Tukey post hoc test revealed
that the percent damage was significantly lower for the new system (30.01 ±
11.37 percent) compared to the no guidance system (86.03 ± 9.78 percent)
and the current system 1 (57.45 ± 10.04 percent). There was no statistically
significant difference (p = 0.095) between the new system and current system
2 (45.93 ± 8.62 percent).

Summary of One-way ANOVA


1. Assumptions: independence, random sampling across groups and randomi-
sation of allocation of cases to treatments, equal variance, normal population
distribution.
2. Hypotheses:
H0 : µ1 = µ2 = ... = µk .
Ha : at least two of the population means are unequal.

3. Test statistic:
Between group variability Treatment Mean Square M ST
F = = =
Within group variability Error Mean Square M SE
F sampling distribution has df = number of groups −1 = k − 1 and N − k (total
sample size - number of groups).
4. P -value: Right tail probability of the above observed F -value.
5. Conclusion: Interpret in context. If decision is needed, reject H0 if P -value ≤
significance level (such as 0.05).

Exercise 7.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 25.6, 25.11,
25.13 - 25.22. Use spss where asked to perform an ANOVA.
7.3. Two-way Factorial ANOVA 7.15

7.3 Two-way Factorial ANOVA


In this course we will only cover Two-way factorial ANOVA conceptually – you will
not be required to calculate parts of the factorial ANOVA table or analyse data in
spss using a factorial ANOVA. If you are interested in understanding this method
more deeply, Factorial ANOVA is covered in some detail in Chapter 26 of De Veaux,
Velleman & Bock, (5th edition).

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 26.

One-way ANOVA is a bivariate (two-variable) method. It analyses the relationships


between the mean of a dependent variable and the groups of a factor (the independent
variable). ANOVA can be extended to analyse two or more factors (we will consider
two factors). With multiple factors, the analysis is factorial.

It is important to make sure we understand the terminology we are using as there are
often multiple words used to describe one part of the analysis.

Dependent variable or response variable – this is the thing that is being measured on
all cases regardless of treatment/group. It should be a quantitative or scale variable
so that we can sensibly calculate a mean, variance, etc. You should only have one
dependent variable in both One-way ANOVA and Factorial ANOVA. In Example 7.1
the response variable was ‘hand volume’. In Example 7.2 the response variable was
‘percentage of neighbouring trees damaged’.

Independent variable or factor – in an ANOVA the independent variable or factor


must be a categorical variable with two or more groups/categories.

ˆ In a One-way ANOVA there can only be one independent variable or factor. In


Example 7.1 the factor was ‘post surgical treatment group’. In Example 7.2 the
factor was ‘harvest system’.
ˆ in a Two-way factorial ANOVA there must be 2 factors. Each factor can have
a different number of groups. Each case that is measured must belong to only
one group in each factor (more on this later).

Groups, levels, treatments, interaction effects –

ˆ In a One-way ANOVA the groups in the factor are often called treatments or
levels. In Example 7.1 there were 3 groups/levels/treatments in the post surgical
treatment factor. In Example 7.2 there were 4 groups/levels/treatments in the
harvest system factor.
7.16 Module 7. Analysis of Variance (ANOVA)

ˆ In a Two-way factorial ANOVA the same meaning applies, however ‘treatments’


can also refer to the combinations of factor groups/levels between the two fac-
tors. These combinations are called the interaction effects.

For the rest of this module we will refer to the groups in a single factor as ‘levels’ and
the combinations of levels as ‘treatments’. Based on the example in the table below,
Factor 1 has 3 levels (A, B, C), Factor 2 has 4 levels (D, E, F, G). Each case belongs
to one level of Factor 1 and one level of Factor 2 so there are a total of 12 possible
factorial treatments (see table below). When using an ANOVA we are interested in
knowing if the variation in our response variable can be explained by our factor(s)
and/or the interaction between them.

Factor 1 Factor 2 Factorial Treatments


A D AD
A E AE
A F AF
A G AG
B D BD
B E BE
B F BF
B G BG
C D CD
C E CE
C F CF
C G CG

Example 7.3

A controlled laboratory experiment has been carried out to look at the effects of
four different chemicals on the control of a noxious weed by limiting its growth.
Since it was unknown whether the chemical effect would depend on the stage of
development of the plant at the time of application, some plants were treated
in an early stage of growth and others later in their growing process. Each
chemical was applied to six different experimental plants, three in early growth
stage and three in late growth, giving a total of 24 experimental plants. Each
chemical was applied at the concentration recommended by the manufacturers.
For confidentiality purposes, the chemicals are known simply as A, B, C and
D. The variable measured was dry matter production at six weeks of age. The
results for the 24 individuals are given below.
7.3. Two-way Factorial ANOVA 7.17

Chemical
Time of application
A B C D
5.9 2.6 4.7 2.0
Early 4.7 3.7 4.3 2.4
5.6 3.4 3.8 1.9
2.5 2.7 3.1 5.9
Late 0.6 2.3 4.1 4.6
1.0 2.9 2.8 5.5

Factor 1 ‘Time of Application’ has 2 levels (Early and Late), Factor 2 ‘Chemical’
has 4 levels (A, B, C, and D) and there are 2x4=8 treatments (e.g. Early-A,
Early-B, Early-C.....Late-C, Late-D). In each of the 8 treatments 3 different
plants where used and the dry matter production at 6 weeks recorded. The
values for each plant are shown in the table above. In total there are 3x8=24
plants. Earlier in this Module we mentioned that all cases that are measured
must belong to only one group/level in each factor. In this example the plants
are the cases and each plant was treated with one Chemical applied at one Time.
No plant was measured more than once and so each plant can be described as
an ‘independent experimental unit’.
Our total sample size N is the total number of independent experimental units
(N=24). Within each of the 8 treatments we have 3 ‘replicates’.
We can calculate the mean of each of the 8 sets of 3 measurements and plot
these mean values:
7.18 Module 7. Analysis of Variance (ANOVA)

The differences between the effect of the chemicals on dry matter production
clearly depends on whether the application occurred during the early or late
stage of growth. If application is in the early growth stage, chemical D appears
to be more effective as it produces much less dry matter of weed; but if ap-
plication is in the later growth stage it is chemical A which is more effective,
and chemical D is the worst of the chemicals (weeds treated with D Late have
a high average production of dry matter). Therefore, it looks like the effect of
chemicals interacts with the time of application — the effect of the chemicals
(which works best) depends on which time of application is being considered.
In the factorial model, instead of the single ‘Treatment’ term we used in the
one-way ANOVA, we now need two terms in the ANOVA table — ‘main effect’
terms and ‘interaction term(s)’.

Source of Degrees of Sum of Mean Square F-ratio


Variation Freedom Squares

SSA M SA
Factor A a−1 SSA M SA =
a−1 M SE
SSB M SB
Factor B b−1 SSB M SB =
b−1 M SE
SSAB M SAB
AB Interaction (a − 1)(b − 1) SSAB M SAB =
(a − 1)(b − 1) M SE

SSE
Error N − ab SSE M SE =
N − ab
Total N −1 TSS

In the ANOVA table above a is the number of levels in Factor A, b is the


number of levels in Factor B and N is the total sample size. Each of the SS
components is calculated from the data. Given the relationships between df
and SS to find the MS components, and also the relationship between Factor,
Interaction and Error MS components to find the F-ratios, once the SS and df
are known all other parts of the table can be easily calculated.
A main effect line provides a test of the equality of means for the particular
factor, where the mean for each level is computed across all levels of the other
factors. It describes the effect of the factor when it is averaged across all levels
of the other factors. In Example 7.3, ‘Chemical’ and ‘Time of application’ are
the 2 main effects.
7.3. Two-way Factorial ANOVA 7.19

H0 : Mean dry matter is equal for Chemicals A, B , C and D, regardless of level


of Time of application
H0 : Mean dry matter is equal for Early and Late Times of application, regard-
less of level of Chemical
An interaction term tests whether or not the two factors interact with each
other; are the differences between the levels of one factor the same regardless
of which level of the other factor is considered OR do the differences change,
depending on which level of the other factor is being considered?
H0 : there is no interaction between Chemical and Time of application
As with One-way ANOVA, the F -tests of the null hypotheses in Two-way
ANOVA assume that:

ˆ the population distribution for each group is normal


ˆ the population standard deviations are equal
ˆ the data result from random sampling or a randomised experiment.

Returning again to Example 7.3, the Two-way Factorial analysis output is


shown below.

Source of df Sum of Mean F-ratio P-value


Variation Squares Square

Time 1 2.042 2.042 5.463 .033


Chemical 3 2.788 0.929 2.487 .098
Time * Chemical 3 39.888 13.296 35.575 .000
Error 16 5.980 0.374
Total 23 50.6983

When interpreting this output we must consider the significance of the inter-
action term first. In this example the Time x Chemical interaction has a
P -value < 0.001 which is less than the standard significance level of 0.05 and
therefore we reject H0 . We conclude that there is a significant interaction be-
tween the levels of Chemical and Time – the variation in dry matter produced
depends on the interaction between the two factors. Recalling the earlier plot
showing the relationship between Time and Chemical this seems reasonable.
When the interaction term is significant, the main effects would need to be in-
terpreted very cautiously because the significant interaction term has already
alerted us that the main effects depend on each other – interpreting them inde-
pendently must be done only with strong justification for why their individual
effects on the dependent variable can be considered when we already know they
7.20 Module 7. Analysis of Variance (ANOVA)

influence each other in how they affect the dependent variable. If the inter-
action is significant then pairwise tests between all treatment levels would be
needed to determine which treatments are different to each other.
If we imagine for a moment that the interaction term in our example had not
been significant, then we could confidently interpret each of the main effects.
In this example there is a significant difference in mean dry matter between
levels of Time (p < 0.05), however there are no significant differences among
the levels of Chemical (p > 0.05).
We can look at the plots of each main effect to see if these results make sense
(continuing to imagine for a moment that the interaction term was not signifi-
cant). The plot below shows all 12 measures of dry matter for Early and Late
Time of application. The line represents the change in mean values between
the groups. Although there is a lot of overlap in dry matter values between the
two levels, because we had a relatively large number of replicates in each level
(n=12) the ANOVA was able to determine that the observed difference in the
means is unlikely to be due to sampling variation alone, and more likely due
to actual differences in average dry matter production between the two levels.

The plot below shows the 6 replicates within each of the 4 levels of Chemical.
Although the change in mean values across the 4 levels looks similar in magni-
tude to that seen for Time of application, in this analysis there are more levels
to compare and fewer replicates in each level. Therefore, the ANOVA could
not distinguish this variation in means from the variation that might occur just
by chance due to our sampling within each level. While the means are different
in each level, we do not have enough evidence to reject H0 .
7.3. Two-way Factorial ANOVA 7.21

Our interpretation is complicated by the fact that the interaction term was
significant in this example and so we need to look at pairwise t-tests between the
factor level combinations. The spss output is too large to include here, but an
example of the first part is shown below. The first and second columns show the
combinations of Time and Chemical being compared. Row 1 is Time1 (Early)
and Chemical1 (A) compared to Time1 Chemical2 (B) – the mean difference in
dry matter production is 2.1667 and this is a significant difference (sig=0.000,
which is much less than 0.05), so reject the null hypothesis. Time1Chemical1 is
significantly different to all other combinations of Time and Chemical, except
for Time2Chemical4 which has a mean difference of 0.0667 and significance of
p=0.870 (which is larger than 0.05), so we cannot reject the null hypothesis.
In the next section of the first column we can see Time1Chemical2 and it
is compared to all other treatment combinations listed in column 2. For all
comparisons where the Sig. value is < 0.05 we reject H0 and conclude that the
treatment means are significantly different. Where the Sig. values are > 0.05
we do not have enough evidence to reject H0 and so conclude that there is no
statistical difference.
7.22 Module 7. Analysis of Variance (ANOVA)

Exercise 7.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 26.5, 26.7, 26.13.

7.4 Closing Comments


In this module, we have learned to recognise when to use an Analysis of Variance
(ANOVA) to compare the means of several groups.

The One-way ANOVA is a natural way to extend the t-test for testing the means of
two groups to compare several groups when those groups represent different levels of
a single factor.
7.4. Closing Comments 7.23

Two-way factorial ANOVA compares means across groups when those categories form
levels of two main factors. Factors can also interact with each other and the factorial
ANOVA can be used to try to understand the influence of both main effects and the
interaction effect on the variation seen in the dependent variable.
7.24 Module 7. Analysis of Variance (ANOVA)

7.5 Tutorial Module 7


The following section contains Tutorial 7 – a good summary of the work learnt in
this module and most importantly, testing your knowledge.

Question 1: Transport Safety

Suppose the National Transportation Safety Board (NTSB) wants to examine the
safety of compact cars, midsize cars, and full-size cars. It collects data on the pressure
applied to the driver’s head during a crash test and runs three separate crash tests
for each of the car sizes. The SST and the SSE , were found to be 86049.56 and
10254 respectively. Use this information to answer the following:

(a) State the response variable, the factor and the treatments.

(b) Is this data appropriate for an ANOVA? What type?

(c) State the null and alternative hypotheses.

(d) What are the degrees of freedom for the treatment, error and total ANOVA table
terms?

(e) What are the M ST and M SE values? (try drawing up the ANOVA table to help
you find these values)

(f) Calculate the appropriate test statistic.

Question 2: Performance of New Design

To study the performance of a newly-designed light plane it was timed (in mins) over
a marked course under three wind conditions, calm, moderate and windy. Use the
following output to state the hypotheses and interpret the results:

Source of Degrees of Sum of Mean F-ratio P-value


Variation Freedom Squares Square

Condition 2 79.7 39.8 3.19 0.078


Error 12 149.9 12.5
Total 14 229.6
7.5. Tutorial Module 7 7.25

Question 3: Electrical Outages

In investigating electrical outages at three different locations, the total outages (in
mins) per month for five months were examined. The following output was obtained:

Source of Degrees of Sum of Mean F-ratio P-value


Variation Freedom Squares Square

Location 2 443.33 221.667 6.70 0.020


Month 4 331.33 82.833 2.50 0.125
Error 8 264.67 33.083
Total 14 1039.33

(a) What is the P -value to answer the question ‘Do the data provide evidence of
different mean levels of outage at the different locations?’

(b) What is the answer to the question ‘Do the data provide evidence of different
mean levels of outage at the different locations?’

(c) What is missing from this Two-way ANOVA?

(d) A person trying to answer the questions above carried out two one-way ANOVA’s,
the output of one of which is given below. Fill in the indicated gaps in the output
below.

Source of Degrees of Sum of Mean F-ratio P-value


Variation Freedom Squares Square

Location 2 ———- ———- ———- 0.036


Error 12 ———- ———-
Total 14 1039.33

Question 4: Arrival Times

The time difference between actual and scheduled arrival times of trains at a large
city station are recorded for three days in each of two weeks, and an analysis included
the following output.
7.26 Module 7. Analysis of Variance (ANOVA)

Source of Degrees of Sum of Mean F-ratio P-value


Variation Freedom Squares Square

Day 2 1168.7 584.3 2.77 0.067


Week 1 20.8 20.8 0.10 0.754
Day*Week 2 1673.7 836.9 3.97 0.021
Error 126 26586.0 211.0
Total 131 29449.2

(a) What does the P -value of 0.021 tell us?

(b) What does the P -value of 0.754 tell us?

(c) What does the P -value of 0.067 tell us?

Question 5: Strength of Plywood

The dry shear strength of birch plywood bonded with different resin glues was studied
with a completely randomised designed experiment. The data from the experiment
is as follows:

Glue A Glue C Glue F


102 100 220
58 102 243
45 80 189
79 119 176
68 176
63
117
Total 532 401 1004

A partially completed ANOVA table shows:

Source of Degrees of Sum of Mean F-statistic P-value


Variation Freedom Squares Square

Treatment 47737 0.000


Error
Total 15 55902
7.5. Tutorial Module 7 7.27

(a) For the shear strength plywood data, what are the treatment and error degrees
of freedom?

(b) For the shear strength plywood data, what are the SSE , M ST and M SE ?

(c) What is the F-Statistic?

(d) What would it mean if the F-statistic = 1.0?

Question 6: Effectiveness of Training Programs

An experiment was conducted to compare the effectiveness of three programs, A, B


and C, in training technicians to use environmental monitoring equipment. Fifteen
technicians were randomly assigned to the programs, five to each. After completion
of the courses, each person was required to carry out four tasks and their average
success rate was recorded:

Training Program Average Success Rate (%)


A 59 64 57 62 62
B 52 58 54 60 59
C 58 65 71 63 64

(a) Enter the data into spss. Make sure you set up your Data View and Variable
View correctly.

(b) Perform a One-way ANOVA in spss. In your output include Descriptives and
Post-hoc Tukey tests.

(c) State the null and alternative hypotheses, and interpret your analysis.

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
7.28 Module 7. Analysis of Variance (ANOVA)

7.6 Module 7 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ recognise when it is appropriate to apply an ANOVA method of analysis;

ˆ identify the assumptions to be checked before performing ANOVA analysis;

ˆ carry out a one-way ANOVA in spss;

ˆ complete and read an ANOVA table, including identifying df, MSS and the
F -statistic.

ˆ interpret the spss output for one-way and a two-way ANOVA and subsequent
pairwise t-tests for significant effects.
Module 8
Association
between
Categorical
Variables
8.2 Module 8. Association between Categorical Variables

Contents
8.1 Contingency Table . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.4
8.2 Associated or Not? . . . . . . . . . . . . . . . . . . . . . . . . . . 8.7
8.3 The Chi-Square Test of Independence . . . . . . . . . . . . . . . 8.10
8.4 The Chi-Square Goodness-of-Fit Test . . . . . . . . . . . . . . . 8.23
8.5 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.26
8.6 Tutorial Module 8 . . . . . . . . . . . . . . . . . . . . . . . . . . . 8.28
8.7 Module 8 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 8.34
8.3

Module Objectives
On successful completion of this module students should be able to:

ˆ identify association between two categorical variables;

ˆ interpret and construct, using spss, a stacked bar chart for two categorical
variables;

ˆ interpret and construct by hand and using spss the joint, marginal and condi-
tional distributions from a contingency table;

ˆ carry out, by hand and using spss, a chi-square test of independence and inter-
pret the results;

ˆ carry out, by hand and using spss, a chi-square test of goodness-of-fit and
interpret the results;

ˆ calculate the standardised cell residuals, both by hand and by using spss, after
a chi-square test and interpret them;

ˆ state and assess the assumptions and conditions necessary for the validity of a
chi-square test; and

ˆ determine when it is appropriate to apply each chi-square test.

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
A variable is categorical if it describes some specific attribute or characteristic of
individuals that can only be categorised, e.g., hair colour (with two ‘values’ dark
and fair), or smoking habit (with three ‘values’ never smoked, past smoker, current
smoker) as responses in a survey.

A contingency table displays the relationship between two categorical variables in a


two-way table of counts which displays frequencies for each combination of the values
8.4 Module 8. Association between Categorical Variables

of two categorical variables. spss calls it a ‘cross-tabulation’ or ‘cross-tab’ for short.


These counts in the cells of the table can be converted into percentages in three
different ways: by columns, by rows, or by the overall total.

The ‘by rows’ or ‘by columns’ percentages produce two sets of conditional distribu-
tions. If a set of conditional distributions are different, we argue that an association
exists between the two categorical variables. The problem with looking for a differ-
ence of course is to decide how different is different. In this module we answer that
question.

De Veaux, Velleman & Bock, (5th edition) cover more than we need for
this topic. Also it presents the material in a different order to the way we
do it. Consequently it’s important to follow closely what’s written below
in this module. Reference is made to the textbook, but only to reinforce
what’s written here, not to replace it.

8.1 Contingency Table

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 3, Sections
1-3.

After this reading write down answers to the following questions:

ˆ What is a contingency or two-way table? When is it used?

ˆ In a contingency table, what is a cell?

ˆ What is meant by a marginal distribution?

ˆ What is meant by a conditional distribution?

ˆ How can we tell if there is an association between two categorical variables?


How can we tell if there is independence or no association?

ˆ How can this information be displayed graphically?

The idea of describing the association between two categorical variables by comparing
conditional distributions is important. Here’s an example from a study of statistics
students.
8.1. Contingency Table 8.5

Example 8.1
Data from a survey of statistics students several years ago gave the following
results:

Student does not Student has job


have job
Student doesn’t 161 73
smoke
Student smokes 51 22

(a) Describe and interpret the joint distribution, the marginal distributions,
and the conditional distributions of smoking and job possession.
(b) Does whether or not a student has a job explain their smoking habit?

Solution
Firstly we need to calculate the marginal and grand totals as follows:

Student does not Student has job


have job
Student doesn’t 161 73 234
smoke
Student smokes 51 22 73
212 95 307

(a) Joint and marginal distributions


These are obtained by dividing each frequency by the grand total giving:

Student does not Student has job


have job
Student doesn’t 52.4% 23.8% 76.2%
smoke
Student smokes 16.6% 7.2% 23.8%
69.1% 30.9% 100%

The four cell percentages adding to 100% comprise the joint distribution.
The 52.4%, for example, may be interpreted as 52.4% of the statistics
students do not have jobs and do not smoke.
There are two marginal distributions, one for whether or not a student
smokes and one for whether or not a student has a job. The 76.2% and
23.8% comprise the first and the 69.1% and 30.9% the second. It is usual to
8.6 Module 8. Association between Categorical Variables

call them marginal distributions when determined from joint information


concerning two (or more) variables. The 23.8%, for example, may be
interpreted as 23.8% of statistics students smoke.

Conditional distributions
There are four conditional distributions in this example, one for each row
and one for each column of the table. The first is obtained by dividing
each cell frequency by the total in its row as shown in the first table below.
The second table below is obtained by dividing each cell frequency by the
total in its column.

Student does not Student has job


have job
Student doesn’t 68.8% 31.2% 100%
smoke
Student smokes 69.9% 30.1% 100%

Student does not Student has job


have job
Student doesn’t 75.9% 76.8%
smoke
Student smokes 24.1% 23.2%
100% 100%
In the first table, the first row may be described as the distribution of
whether or not a student has a job given that the student doesn’t smoke.
Similarly, the second row may be described as the distribution of whether
or not a student has a job given that the student does smoke. Hence, for
example, the 30.1% may be interpreted as the percentage of students who
have a job out of those who smoke.
The two columns in second table similarly define conditional distributions.
Both describe whether or not a student smokes; the first conditional on
the student not having a job, the second on the student having a job.
Hence, for example, the 23.2% may be interpreted as the percentage of
students who smoke out of those who have a job.
(b) If a student having a job helps explain their smoking habit (or, equiv-
alently, if a student smokes helps explain whether or not a student has
a job) then we say the two variables concerning jobs and smoking are
associated. Such an association is indicated by a difference between the
conditional distributions, either in the first table, or equivalently in the
second table displaying conditional distributions. In this instance, com-
paring 68.8% with 69.9% and 31.2% with 30.1% (or equivalently 75.9%
8.2. Associated or Not? 8.7

with 76.8% and 24.1% with 23.2%) we would conclude the differences are
too small to substantiate the existence of an association.

Exercise 8.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 3.3, 3.5, 3.9,
3.15, 3.29.

8.2 Associated or Not?


We answered this question subjectively in Example 8.1 by just eye-balling the condi-
tional distributions and arguing they didn’t look much different and therefore there
didn’t appear to be an association between having a job and smoking habit.

The main aim of this module is to do a little better than this by using the ideas of
inference to come up with an objective way of deciding whether or not there’s an
association between two categorical variables.

So what do we mean by association?

We all recognise that hair colour and eye colour are associated because dark-haired
people tend to have dark coloured eyes and fair-haired people tend to have blue or
light coloured eyes. Sure, there are exceptions to this ‘rule’, but ‘on average’ that’s
the state of affairs. In other words, the colour of your hair gives some idea of the
colour of your eyes (and the colour of your eyes gives some idea about the colour of
your hair).

What we mean then when we say two categorical variables are associated is that
there is some dependence between them in the sense that knowing the value of one
helps decide the value of the other, or some values of one variable tend to go with
particular values of the other. So, knowing a person is blue eyed increases the chance
the person is fair haired but doesn’t help much in deciding if the person is male or
female, because hair and eye colour are associated whereas eye colour and gender are
not.

Note the language. When two variables are associated we say they are dependent on
each other. And when two variables are not associated we say they are independent
of each other. The words ‘associated’ and ‘dependent’ are used interchangeably.

It’s not enough for us to decide from logical considerations that two variables are
associated or not. After all, what we’re interested in finding out is if two variables
are associated, not from our preconceived notions about what we think the state
of the world is, but directly from data concerning the variables. Also, we need to
8.8 Module 8. Association between Categorical Variables

remember that just because the values of one variable might tend to associate with
particular values of the other variable in our data, that doesn’t necessarily mean the
two variables are definitely associated. It could be that sampling variation has raised
its ugly head and is trying to mislead us. We need to be sensitive to this in coming
up with a method of establishing association.

An example will help focus on what’s important here.

Example 8.2
Suppose one hundred statistics students are randomly chosen. Of these, sup-
pose 40 are male and 60 are female, and 70 are dark haired and 30 are fair
haired. The contingency table representing these data therefore has row and
column totals as shown here.

Dark hair Fair hair


Female 60
Male 40
70 30 100

We haven’t seen any information yet about how many female students are dark-
haired, or how many male students are fair haired, so we can’t fill the cells of
the table in.
Let’s assume though that gender and hair colour are independent of each other,
as we believe they are. If this is the case then we should be able to more or
less ‘guess’ what the cell frequencies are.
Before reading on, see if you can do this on the basis of what you understand
by independence.
Independence means that the value of one variable doesn’t help determine the
value of the other. In other words, the distribution of one variable is the same
regardless of the value of the other variable. For this example, that means the
distribution of hair colour is the same for both males and females. Also, the
distribution of gender is the same for both dark haired and fair haired students.
How does this help determine the expected cell frequencies? Give it some more
thought if you haven’t determined them yet.
To answer the question, the distribution of gender for all 100 students is 60%
male, 40% female and the distribution of hair colour is 70%, 30% for dark and
fair hair. These distributions are expected to stay the same regardless of hair
colour and regardless of gender. That is, 70% of females are expected to have
dark hair and 30% fair. Similarly, 70% of males are expected to have dark
hair and 30% fair. But there are 60 females and 40 males. Therefore 70%
8.2. Associated or Not? 8.9

of 60 or 42 of the females and 70% of 40 or 28 of the males are expected to


have dark hair. The remaining females and males are expected to have fair
hair. We therefore get the following table of expected frequencies, assuming
independence of course.

Dark hair Fair hair


Female 42 18 60
Male 28 12 40
70 30 100

Notice, in case you’re wondering, the same expected counts are obtained by
thinking in terms of the gender distribution rather than the hair colour distri-
bution. In other words, 60% of the 70 of the dark-haired students are female
and 40% of the 30 fair-haired students are male giving the same frequencies as
above. This is no accident. It will always be the case.
These cell counts or frequencies that we worked out assuming independence
are called expected counts or expected frequencies. Even if the two variables
are not in fact independent we can still work out expected frequencies using
this logic. There’s a formula that allows us to do this directly from the totals
in the table without going through this logic each time. It is
row total × column total
Exp = (8.1)
table total
Applying this formula to the top left-hand cell (dark-haired female) gives
60 × 70
Exp = = 42
100
as found above. You should confirm this formula works for the other three
expected counts in the table.
Of course, the actual or observed counts may be quite different to the expected
counts. Because of sampling variation, we’d be surprised (or suspicious) if the
observed counts exactly match the expected counts even when the two variables
are independent. We’d be asking, whatever happened to sampling variation?
What is being demonstrated here is the meaning of independence in statistical
terms. If two categorical variables are independent then the respective marginal
and conditional distributions are similar. In this example we saw that this is
the same as noting that two variables are independent if the observed and
expected frequencies are similar.
If the observed and expected frequencies are a lot different however, we’d sus-
pect that the two variables are dependent or associated. How we decide what
we mean by a ‘lot’ different is coming up.
Can you feel a hypothesis test coming on?
8.10 Module 8. Association between Categorical Variables

We now turn our attention to deciding whether or not differences between the ob-
served and expected counts could be due to sampling variation.

8.3 The Chi-Square Test of Independence


A clever way of deciding whether the differences between the observed and expected
frequencies (or counts) in a contingency table are due to sampling variation is the
chi-square test of independence. ‘Chi’ is the Greek letter χ, pronounced ‘kye’ with a
hard ‘k’.

Incidentally, this test is more accurately called the chi-square test of association be-
cause we are trying to find out if two variables are associated, rather than not associ-
ated. Regardless it’s universally referred to as the chi-square test of independence or
Pearson’s chi-square test of independence in honour of the statistician whom we can
blame for its existence.

As in all hypothesis testing, we have a null and an alternative hypothesis. The general
form of these for a chi-square test of independence is always the same; namely
H0 : A and B are not associated
Ha : A and B are associated,
where A and B are categorical variables. Don’t get these around the wrong way. Just
remember that a null hypothesis indicates an absence of an effect—in this case, an
absence of association, i.e., no association between A and B.

It’s also fine to use the word independence in place of ‘not associated’ and ‘dependent’
in place of ‘associated’. So ‘A and B are not associated’ can be expressed as ‘A and
B are independent’ and ‘A and B are associated’ can be expressed as ‘A and B are
dependent’.

Of course we always put these statements into the context of the data being tested.

So, for the smoking versus job data from Example 8.1, we would write something like
H0 : Whether a student smokes is independent of whether or not the student has a job
Ha : Smoking status and whether or not a student has a job are associated

The first step in the analysis is to present the data as follows:


8.3. The Chi-Square Test of Independence 8.11

Observed counts

Student does not Student has job


have job
Student doesn’t 161 73 234
smoke
Student smokes 51 22 73
212 95 307

We now use the totals to compute the expected frequencies from formula (8.1). Check
these out for yourself.

Expected counts

Student does not Student has job


have job
Student doesn’t 161.59 72.41 234
smoke
Student smokes 50.41 22.59 73
212 95 307

Notice that two decimal places have been included. (Expected frequencies do not
have to be whole numbers since they can be thought of as long run averages, and it
is recommended they be expressed to two decimal places of precision.)

Also note that marginal totals have been included. They will be the same (allowing for
roundoff) as the totals in the previous table if the calculation of expected frequencies
is correct. You should check the row and column totals once the expected frequencies
have been worked out as a check on your calculations.

We now come to calculation of the test statistic, called chi-square.

The formula for the chi-square test statistic is


X (Obs − Exp)2
χ2 = (8.2)
Exp
all cells

where Obs stands for observed count and Exp for expected count. Notice how this
formula is based around the difference between the observed and expected counts in
each cell of the contingency table.
8.12 Module 8. Association between Categorical Variables

What does this mean for our data? There are four cells in the table. Therefore there
are four contributions of the sort (Obs − Exp)2 /Exp to calculate. The formula asks
us to add these to give
(161 − 161.59)2 (73 − 72.41)2 (51 − 50.41)2 (22 − 22.59)2
χ2 = + + +
161.59 72.41 50.41 22.59
= 0.029

The value of χ2 depends on the size of the differences between the observed and
expected counts. It doesn’t matter whether the difference is positive or negative.
Squaring the difference gets rid of the negatives. The bigger the differences, the
bigger the value of χ2 . Or putting it around the other way, if the observed and
expected frequencies are similar in each cell, the value of χ2 will be small.

What is the smallest value that χ2 could ever be?

That’s correct. Zero. If the corresponding observed and expected counts are exactly
the same, χ2 would be zero. Mind you, if that happened, we would wonder what
became of sampling variation, and we might be a bit suspicious of the data. But in
theory it could happen. Chi-square then must be positive, or zero at least.

For our data you will have noticed the observed counts and the expected counts are
quite similar. This suggests, even without working out the chi-square value, that
we are not going to be able to reject the null hypothesis of no association. The
small chi-square value of 0.029 confirms this observation. We’ll continue through the
procedure though to show how to convert the test statistic into a P -value as is usual
in hypothesis testing.

How do we find the P -value? Well, as with the t statistic, someone has kindly
worked out what sort of values are to be expected for χ2 when the null hypothesis is
true. That is, they’ve come up with a distribution of the values of χ2 when the two
categorical variables defining a contingency table are independent. This distribution
can be found as Table X at the back of De Veaux, Velleman & Bock. We can look up
the value we found for our test statistic and find a range in which the P -value lies.

Have a look at Table X. Notice the picture of the distribution. What shape has it
got? What values does it take on? Also notice we need to know what degrees of
freedom (df) are involved before we can use the table, just like we did for Table T.

Degrees of freedom
Chi-square measures the difference between the actual counts we observe and the
counts we expect to get if there is no association. If chi-square is ‘small’ (close to
8.3. The Chi-Square Test of Independence 8.13

zero) then the observed and expected counts are similar and we won’t be able to reject
H0 . If chi-square is ‘large’, then the observed and expected counts differ considerably
and we would be able to reject H0 .

The question is what do we mean by small or large?

Well that depends not only on the size of the differences between the observed and
expected counts but also on the number of cells in the table.

After all, the larger the table (i.e., the more cells), the larger potentially chi-square can
be, because every cell contributes to chi-square. We need therefore to take account
of the number of rows and columns in the table when using Table X. The degrees of
freedom does this because it measures the size of the table.

The formula for degrees of freedom is easy. If there are r rows and c columns in the
contingency table we have
df = (r − 1)(c − 1) (8.3)

So, for example, if the table has two rows and four columns (excluding labels and
totals of course), df = (2−1)(4−1) = 3. The smallest possible contingency table with
just two rows and two columns (as in the gender/job example) has just one degree of
freedom.

To use Table X then we need to go to the row with the appropriate number of degrees
of freedom and look across that row and locate the whereabouts of the chi-square
value. From this the figures at the top of the table indicate the range for the P -value.

If df = 3 and χ2 = 10.5, for example, then the chi-square value is between 9.348 and
11.345 according to Table X. The figures at the top of these columns, 0.025 and 0.01,
are P -values. The P -value then for χ2 = 10.5 with 3 degrees of freedom is between
0.01 and 0.025 or 1% and 2.5%.

As another example, if χ2 = 15.8 and df = 3, then because 15.8 is greater than 12.838,
the P -value is less than 0.005.

Notice that as the χ2 value gets larger, the P -value gets smaller. This is typical of
test statistics and P -values. The larger the magnitude (i.e., ignoring the sign) of a
test statistic such as t or z or χ2 , the smaller the P -value and vice versa. This is
to be expected. A large magnitude for the test statistic indicates the data does not
agree with the null hypothesis—in other words there is only a small probability (the
P -value) that sampling variation alone is the cause.

Back to our example with smoking and jobs. The degrees of freedom for our 2×2 table
is one and χ2 = 0.029. From the df = 1 row in Table X we see that 0.029 is smaller
than the smallest entry, 2.706. We conclude that the P -value is greater than 10%.
Consequently, as we thought earlier, there is insufficient evidence to conclude that
8.14 Module 8. Association between Categorical Variables

smoking and whether or not a student has a job are associated, i.e., the differences
between the observed counts in the sample and the expected counts based on marginal
distributions are most likely due to sampling variation.

Let’s now consider an example where there is a statistically significant association.

Example 8.3
See De Veaux, Velleman & Bock, (5th edition) R6.42 (Review Section VI after
Chapter 23). The data in this question has to first be converted into a table of
observed counts as follows.

< 3 years HS 3+ years HS Some college


Planned 200 271 137
Unplanned 391 337 102

H0 : Unplanned pregnancies and education level are independent


Ha : There is an association between unplanned pregnancies and education level
The expected counts are determined using formula (8.1). Check these yourself.

< 3 years HS 3+ years HS Some college


Planned 249.88 257.07 101.05
Unplanned 341.12 350.93 137.95

The chi-square test statistic using (8.2) then gives

(200 − 249.88)2 (271 − 257.07)2 (137 − 101.05)2


χ2 = + +
249.88 257.07 101.05
(391 − 341.12) 2 (337 − 350.93)2 (102 − 137.95)2
+ + +
341.12 350.03 137.95
= 40.71

Check this calculation yourself using your calculator.


There are r = 2 rows and c = 3 columns giving df = (2 − 1) × (3 − 1) = 2.
In Table X with df = 2 we find 40.71 is larger than the largest value given of
10.597. Therefore the P -value is less than 0.005.
We conclude there is strong evidence of an association between unplanned preg-
nancies and education level, i.e., the differences between the observed counts
in the sample and and the expected counts based on marginal distributions are
very unlikely due to sampling variation.
8.3. The Chi-Square Test of Independence 8.15

Follow-up analysis
It’s not altogether satisfactory to conclude the existence of an association without
investigating the nature of the association or the major source(s) of the association.

An appropriate way of doing this involves looking carefully at the differences between
the observed and expected counts. These differences are called residuals. Therefore
the residual for each cell is given by
residual = observed count − expected count

For the data in Example 8.3 the residuals are


< 3 years HS 3+ years HS Some college
Planned −49.88 13.93 35.95
Unplanned 49.88 −13.93 −35.95

Make sure you understand where these are coming from. We’ve simply subtracted the
table of expected counts from the table of observed counts. Notice that the sum of
the residuals in each column and in each row is zero. This is typical of what happens
with residuals. You might recall that the sum of the residuals in a regression analysis
is zero.

Now it turns out that although these residuals tell us directly which observed counts
are bigger than the corresponding expected count and which are less, the importance
of each residual in the overall significance of the association depends also on the
expected number in each cell itself. Putting this another way, a residual of 10 (or
−10) is more important if the expected cell count is small, like 40 say, than if the
expected count is large, like say 100. To take account of this we convert the residuals
into standardised residuals by dividing each by the square root of its expected count:
observed count − expected count
standardised residual = √ (8.4)
expected count

The standardised residuals for the data in Example 8.3 are then
< 3 years HS 3+ years HS Some college
Planned −3.16 0.87 3.58
Unplanned 2.70 −0.74 −3.06

Compare the formula for chi-square (8.2) with the one for standardised residuals (8.4)
and you’ll see that a standardised residual for a cell is just the square root of the chi-
square contribution for the cell. Or putting it another way, the value of the chi-square
test statistic is just the sum of all the standardised residuals squared.
8.16 Module 8. Association between Categorical Variables

For example, in Example 8.3, the contribution to chi-square of the (Planned, Some
college) cell is (137 − 101.05)2
√ /101.05 = 12.79 and the value of the standardised
residual is (137 − 101.05)/ 101.05 = 3.58, which is the square root of the chi-square
contribution (allowing for round-off).

The relative size of the standardised residuals (ignoring any negative sign) indicates
the relative importance of the cell to the chi-square value. If the chi-square value
is significant (i.e., it is ‘large’ with a small P -value) then the cell with the largest
standardised residual is the main contributor. By considering the direction of the
residual for this cell, the main source of the significant association can be described
in the context of the data.

Example 8.4
Following-up on the analysis in Example 8.3 we see that the largest standard-
ised residual ignoring the negative sign is 3.58 for the (Planned, Some college)
cell. (If there was a negative sign, it would indicate that the observed count
is less than what is expected if there were no association between unplanned
pregnancies and education level.) We conclude that the main reason a statis-
tically significant association is observed is that more educated women tend to
have fewer unplanned pregnancies.

Note that it only makes sense to do a follow-up analysis of the residuals if there is a
statistically significant association. Only the largest one or two standardised residuals
need to be looked at and interpreted.

This module is self-contained so the following reading only reinforces material already
covered.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 3.

Conditions for the test


Like all inferential procedures, there are a set of conditions that need to be satisfied for
the method to work properly. For a chi-square test of independence, these conditions
are:

1. The sample is a simple random sample from the population of interest;


8.3. The Chi-Square Test of Independence 8.17

2. The sample is not more than 10% of the population (the 10% condition);

3. The expected counts are all at least five (the expected count condition).

The last condition insists in effect that the sample is reasonably large so that each
cell is expected to contain at least 5 individuals. Note, it’s not the observed counts
we worry about—some of these could even be zero—but the expected counts.

What do we do if some of the expected counts are less than 5? One approach is
to reduce the number of rows or columns of the contingency table by combining
categories together. A row (or column) containing an expected count less than 5 is
combined with another row (or column) so that the problem cell is eliminated. This
may need to be repeated to eliminate all problem cells. It would also need to make
sense in the context of the situation to combine categories in this way. Obviously a
2 × 2 contingency table can’t be reduced any further and there are other methods of
dealing with this problem which are not covered in this course.

We only need to know about the technique of combining rows or columns. Even then,
none of the set problems will require you to do this. Also, as usual, it’s not necessary
to state and check these conditions unless specifically requested in a problem.

Of course, in the real-life practice of statistics or implementation of sta-


tistical techniques, knowing and checking conditions and assumptions is
essential before proceeding.

This module is self-contained so the following reading only reinforces material already
covered.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 4.

Equivalence with Homogeneity of Proportions Test


An interesting way of looking at the chi-square test of independence is that it is
really testing whether or not some proportions are equal. After all, when there is no
association, the row (or column) conditional distributions are all the same; in other
words the proportions across the categories for one variable are the same for each
value of the other variable.
8.18 Module 8. Association between Categorical Variables

For example, if smoking and having a job are independent, the conditional distri-
butions by rows will be the same. But that’s the same as saying the proportion of
smokers with a job is the same as the proportion of non-smokers with a job; and the
proportion of smokers without a job is the same as the proportion of non-smokers
without a job. And since the proportions without a job can only be equal if the pro-
portions with a job are equal, we could state the null hypothesis as H0 : p1 = p2 where
p1 is the proportion of smokers with a job and p2 is the proportion of non-smokers
with a job.

Of course, independence also means the conditional distributions by columns are the
same. Therefore, the null hypothesis could also be stated as H0 : p1 = p2 where p1
is the proportion of those with a job who smoke and p2 is the proportion of those
without a job who smoke.

This idea extends to larger tables. For Example 8.3, the null hypothesis could be
written either as
H0 : Unplanned pregnancies and education level are independent, or
H0 : The distribution of unplanned and planned pregnancies is the same at each
education level,
and this second way of stating the null hypothesis could be re-expressed in terms of
proportions. That’s a bit messy because there’s quite a few proportions involved and
it doesn’t add to the idea of this section.

Also the null hypothesis could be written


H0 : The distribution of education levels is the same for women with planned or
unplanned pregnancies
and once again this could be re-expressed in terms of proportions.

The point is, a test of independence is the same as a test of equality (i.e., homogeneity)
of (conditional) distributions which is the same as testing equality of one or more sets
of proportions.

De Veaux, Velleman & Bock, (5th edition) treat the test of independence and test
of homogeneity of proportions separately (see Chapter 22, Section 2). There are
occasions where it makes sense to do that, but we don’t make the distinction in this
module and you won’t be penalised for not making the distinction. Consequently, if,
for example, interest is in assessing whether there is a difference between males and
females in the proportions who smoke, the appropriate null and alternative hypotheses
are
H0 : Gender and smoking are not associated
Ha : Gender and smoking are associated

Or, if we’re interested in whether the grade distribution of on-campus and online
8.3. The Chi-Square Test of Independence 8.19

statistics students differ, it’s appropriate to state the hypotheses as


H0 : Study mode and grade in statistics are independent
Ha : Study mode and grade in statistics are dependent

Using SPSS

For any but a small contingency table, a chi-square test requires a lot of calculation.
You will not be expected to calculate chi-square ‘by hand’ for Tables with more than
six cells.

Exercise 8.2
Watch the following spss videos. After watching the videos try to
replicate the analysis in Example 8.5 below.

ˆ Chi-square Test of Independence

ˆ Weight Cases

When the data is given in a summarised form, such as a table of counts as in Exam-
ple 8.3, these spss videos and the description below indicate the process for analysing
this data using a chi-square test. If the data is given in raw form, i.e., data values are
given for all cases, you will not need to apply Weight Cases. The following example
provides a brief outline of the method for summarised data such as might be found
in textbook exercises.

Example 8.5
Examples 8.3 and 8.4 are repeated here using spss.
To enter the data, each of the observed counts is identified by a row and column
number (see Figure 8.1). So, for example, 391 is in Row 2, Column 1. To make
the spss output look good of course, include variable value and data value
labels for the ‘Row’ and ‘Col’ variables and their coded values in the Variable
View window.
8.20 Module 8. Association between Categorical Variables

Figure 8.1: SPSS data input for unplanned pregnancy example.

Weighting is now applied as shown in Figure 8.2 using the Data/Weight Cases
procedure and ‘count’ as the weighting variable (see the spss video Weight
Cases). This tells spss that there are 200 individuals with < 3 years HS
who had planned pregnancies, 271 with 3 or more years HS who had planned
pregnancies, etc. (If this is not done, spss thinks each row of the Data View
spreadsheet represents just one individual.)

Figure 8.2: Data weighting procedure for chi-square test of independence.

The chi-square test is called up using ‘Analyze/Descriptive Statistics/Crosstabs’.


Enter ‘Row’ and ‘Col’ as shown in Figure 8.3. Lots of bells and whistles are
available. Some of the options under ‘Cells’ and ‘Statistics’ are checked, as
shown in Figure 8.4. The important one of course is the chi-square test itself.
The output is shown in Figures 8.5. Make sure you can read this output.
8.3. The Chi-Square Test of Independence 8.21

Figure 8.3: spss dialog box for contingency table and chi-square test procedure.

Figure 8.4: spss options for chi-square test of independence.

Notice the Observed, Expected Counts and Standardised Residuals are dis-
played followed by the value of the test statistic (called ‘Pearson’s Chi-Square’),
40.715, degrees of freedom and P -value (‘Asymp. Sig. (2-sided)’) = .000 (i.e.
it is 0.000 when rounded off to three decimal places). Note that the chi-square
test is a two-sided test.
Notice how spss also does a check on the expected cell frequency condition.
Of course, it won’t stop reporting the results of the test even if the condition
fails. We have to be aware of what is valid and what isn’t.
8.22 Module 8. Association between Categorical Variables

Figure 8.5: spss Output from chi-square test of independence.

Now it is time to check your understanding by answering the following questions:

ˆ What is the aim of a chi-square test of independence?

ˆ In general terms, what does the null hypothesis state in a test of independence?
What about the alternative hypothesis?

ˆ What condition and assumptions are required for the test?

ˆ What is the formula for the chi-square statistic?

ˆ What is the formula for the degrees of freedom for the test of independence?

ˆ Is a large value of χ2 evidence for or against H0 ?

ˆ Is the chi-square distribution symmetric, skewed to the right or skewed to the


left?

ˆ What is the formula for a residual? For a standardised residual?

ˆ How do we ascertain the main reasons for a significant association?


8.4. The Chi-Square Goodness-of-Fit Test 8.23

ˆ What exactly is the chi-square test statistic called in spss?

Exercise 8.3
Do De Veaux, Velleman & Bock, (5th edition), exercises 22.9 (by hand),
22.39 (use spss), 22.41 (use spss).

De Veaux, Velleman & Bock, (5th edition) describe three chi-square tests: a goodness
of fit test, a test of independence and a test of homogeneity of proportions. The last
two have been discussed in this module. Now it is time to turn our attention to the
Goodness-of-Fit test.

8.4 The Chi-Square Goodness-of-Fit Test


We sometimes wish to test whether the distribution of a categorical variable follows
a particular probability model. This test is appropriately called a Goodness-of-Fit
Test. While the Goodness-of-Fit Test is not one of association between two categorical
variables (that we looked at in Section 8.3), it does however involve a chi-square test
statistic. For this test we are interested in how good the data fits the probability
model that has been proposed, i.e. we conduct a Goodness-of-fit test to see if the
sample comes from the population with the claimed distribution (does the sample
data ‘fit’ the population distribution).

Example 8.6
The manufacturer of Beanies (a fruit-flavoured gummy-textured sweet in the
shape of a broad bean seed) is concerned that their mixing machine is not
calibrated correctly. If the mixing machine is properly calibrated to reflect
advertised ratios of colours, the distribution of colours of Beanies should be
30% red, 30% green, 20% orange and 20% purple. A random sample of 180g
bags of Beanies were removed from the packaging machine and yielded the
following data:

Colour
Red Green Orange Purple Total
Count 902 1096 654 806 3458
Percent 26.1 31.7 18.9 23.3 100%
Expected % 30 30 20 20 100%
8.24 Module 8. Association between Categorical Variables

How ‘good’ does the data fit the probability model we would expect if the
mixing machine was correctly calibrated?
Step 1: Write hypotheses
H0 : The distribution of colours of Beanies follows the manufacturer’s require-
ment of 30% red, 30% green, 20% orange and 20% purple.
Ha : The distribution of colours of Beanies is different to the manufacturer’s
requirement of 30% red, 30% green, 20% orange and 20% purple.
Step 2: Assuming that the distribution is as required, calculate the χ2 test
statistic
X (Obs − Exp)2
χ2 =
Exp
all cells
However, before we can do this we need to calculate expected counts based on
the manufacturer’s required distribution of colours.

Colour
Red Green Orange Purple
Observed Count 902 1096 654 806
Expected Count 1037.4 1037.4 691.6 691.6

The chi-square test statistic then gives

(902 − 1037.4)2 (1096 − 1037.4)2 (654 − 691.6)2


χ2 = + +
1037.4 1037.4 691.6
(806 − 691.6)2
+
691.6
= 41.95

Step 3: Calculate the P -value.


The degrees of freedom for this test is the number of categories −1, thus df =
n − 1 = 4 − 1 = 3 giving a P -value of < 0.005 from Table X.
Note: the degrees of freedom for the goodness-of-fit test is calculated differ-
ently from that for a test of independence.
Step 4: Write a conclusion in context
The observed chi-square value of 41.95 represents how similar the observed data
are to the counts we would expect, if H0 is true. The P -value of less than 0.005
(0.5%) indicates that we could obtain this chi-square value less than 0.5% of
time if the distributon of colours of Beanies is as the manufacturer has been
advertising. As the P -value is less than 0.5% there is very strong evidence
to accept Ha , and state that the observed counts of colours of Beanies are
different enough for us to conclude that the mixing machine is not correctly
calibrated.
8.4. The Chi-Square Goodness-of-Fit Test 8.25

Note that the same assumptions/conditions apply to the Goodness-of-Fit


test as the Test of Independence and should be checked before proceeding to
the hypothesis test.

SPSS Output
The initial output for performing the Chi-square Goodness-of-Fit test in spss
gives only the Sig. or P -value of .000 and the decision to reject the null hypoth-
esis. Notice that it says that this is a ‘One-sample Chi-Square Test’ which indi-
cates that the steps to doing this test in spss require you to click on ‘Analyze >
Nonparametric > Legacy Dialogs > One Sample’ to access the Goodness-of-Fit
test.

Figure 8.6: spss initial output for Goodness-of-Fit Test.

If you click on this table, it will show more detailed output that includes the
sample size, test statistic, degrees of freedom and P -value.
8.26 Module 8. Association between Categorical Variables

Figure 8.7: spss detailed output for Goodness-of-Fit Test.

This module is self-contained so the following reading only reinforces material already
covered.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 22, Section 1.

Now for some practice.

Exercise 8.4
Do De Veaux, Velleman & Bock, (5th edition), exercises 22.19 (by hand
and using spss), 22.21.

8.5 Closing Comments


Because categorical data is commonly collected in all fields, the chi-square test is very
popular. Students doing biology, psychology and marketing in particular will come
across the test in future studies.
8.5. Closing Comments 8.27

This module covers the material on contingency tables and chi-square tests. First
we dealt with converting the table to percentages in three different ways and talked
about the idea of association. We’ve now got a method of assessing the statistical
significance of an association – Chi-square Test of Independence. In addition we can
also test whether a frequency distribution fits a specific pattern – Goodness-of-Fit.

One final point. Remember that the chi-square test is applied to frequency data. Even
though we often convert frequencies to proportions for descriptive purposes and to
improve our appreciation of the relationships amongst the groups under consideration,
the chi-square calculations must be performed on the original frequencies or counts,
not the percentages.
8.28 Module 8. Association between Categorical Variables

8.6 Tutorial Module 8


Time to put your knowledge to the test and practise your skills! Work through these
tutorial exercises before you look at the worked solutions which will be made available
on the StudyDesk in due course. You will find it useful to work through the examples
in this Module for guidance.

Question 1: Investigate the following research question


Is there a relationship between the amount of coffee consumed and the number of
hours of sleep students manage to have per night?

The following data was collected and organized into the Contingency Table below to
see if there is a relationship between amount of coffee consumed (stated as ‘3 or less
coffees per day (Light)’ and ‘more than 3 coffees per day (Heavy)’) and the number
of hours slept per night (stated as ‘7 or less hours per night’ and ‘more than 7 hours
per night’) for 151 statistics students.

7 or less More than Total


hours per night 7 hours per night
3 or less coffees per day (Light) 57 19 76
More than 3 coffees per day (Heavy) 36 39 75
Total 93 58 151

Analysis and questions

(a) Display the joint distribution of level of coffee consumed and level of sleep in
terms of percentages.

7 or less More than


hours per night 7 hours per night
3 or less coffees
per day (Light)
More than 3 coffees
per day (Heavy)

(b) Using the joint distribution in (a), write 4 sentences describing what each of these
4 percentages mean (i.e. one of the sentences could be something like ”90% of
statistics students drink 3 or less cups of coffee per day and have 7 or less hours
of sleep per night”).

(c) Calculate the two marginal distributions involved in terms of percentages.

ˆ Marginal Distribution of Level of Sleep


ˆ Marginal Distribution of Level of Coffee Consumed
8.6. Tutorial Module 8 8.29

7 or less More than Total


hours per night 7 hours per night
3 or less coffees
per day (Light)
More than 3 coffees
per day (Heavy)
Total

(d) Using percentages,

(i) calculate the distributions of level of coffee consumed conditional on level of


sleep (i.e. calculate the distribution of level of coffee consumed for those who
sleep 7 or less hours and then the distribution of level of coffee consumed
for those who sleep more than 7 hours).

7 or less More than


hours per night 7 hours per night
3 or less coffees
per day (Light)
More than 3 coffees
per day (Heavy)

(ii) calculate the distributions of level of sleep conditional on level of coffee


consumed.

7 or less More than


hours per night 7 hours per night
3 or less coffees
per day (Light)
More than 3 coffees
per day (Heavy)

(e) Sketch a suitable graph to display the relationship between the two variables,
using the information from (d) (there are two ways of doing this, either using
the information from your answer to d(i) or the information from your answer to
d(ii)).
8.30 Module 8. Association between Categorical Variables

(f) From the data in the study, what percentage of statistics students are heavy coffee
drinkers AND sleep 7 or less hours per night?

(g) What percentage of students who are light coffee drinkers also sleep 7 or less
hours per night?

(h) What percentage of students who sleep 7 or less hours per night are light coffee
drinkers?

(i) What percentage of students slept 7 or less hours per night?

(j) What percentage of students are light coffee drinkers?

(k) Are the two categorical variables associated?


We used two categorical variables in this exercise and have examined the joint,
marginal and conditional distributions of these. We now want to examine if there
is any relationship between level of sleep acquired and the level of coffee consumed
by statistics students.
Do you think there is an association between level of sleep and level of coffee
consumed? Why or why not? Use the conditional distributions in (d) and/or
the bar graphs in (e) to help you decide on this answer and write a sentence to
support your conclusion.

(l) Now follow up on this conclusion by conducting an appropriate hypothesis test


using spss to calculate the test statistic and P -value. Do you come to the same
conclusion? (You will need to use Weight Cases when entering the data into
spss.)

NOTE on wording:

ˆ Saying that two categorical variables are not associated (or not related) is
the same as saying they are independent.
ˆ Saying that two categorical variables are associated (or related) is the same
as saying they are NOT independent.

(m) In (k) we examined the association between two categorical variables (level of
sleep and level of coffee consumed). To further test your understanding of cate-
gorical variables and association, can you suggest:

(i) Two other categorical variables that you would expect to be independent?
(ii) Two other categorical variables that you would expect to be associated?
8.6. Tutorial Module 8 8.31

Question 2: Investigate the following research question

Is there an association between hand dominance and leg dominance?

Students in a statistics lecture were asked to clasp their hands together, fingers inter-
twined, and then note which thumb was on top (right thumb equates to Right Hand
Dominance). They were then asked to cross their legs, one ankle over the other, and
note which ankle was on top (right ankle equates to Right Leg Dominance). Using the
answers to these questions, the following two-way table was produced (Figure 8.8):

Figure 8.8: Cross-tabulation of Hand Dominance with Leg Dominance.

(a) Conduct an appropriate hypothesis test by hand to answer the research question:

Step 0: Compute the conditional distribution of leg dominance for right hand
dominance and then left hand dominance. Compare these two distributions and
give a tentative answer to the question.
Step 1: Write the hypotheses for answering this question.
Step 1a: What assumptions/conditions needed to be satisfied to perform this
test?
Step 1b: Compute a table of the expected frequencies/counts.
Step 2: Calculate the test statistic and give the degrees of freedom for the test.
Step 3: Calculate the P -value.
Step 4: Write a conclusion in context.
Step 5: Compute the standardised residuals and interpret the relationship be-
tween Hand Dominance and Leg Dominance.

(b) Now repeat the hypothesis test using spss to complete the Steps 1b, 2 and 3 as
in (a). Compare your answers with those in (a).
8.32 Module 8. Association between Categorical Variables

Question 3: Is there an association between education level and age group?

An investigation of the differences in education level attained among different age


groups in the United States produced the following data (see De Veaux, Velleman &
Bock, (5th edition), Chapter 22, Question 50).

Figure 8.9: Data for Education Level by Age.

Using the data above, spss produced the following output:

Figure 8.10: spss output for Education by Age.

Use this output and the data above to answer the following questions:

(a) Write the hypotheses for testing if Education Level Attained is associated with
Age Group.
(b) What assumptions/conditions are needed to be satisfied to perform this test?
8.6. Tutorial Module 8 8.33

(c) State the test statistic, df and P -value (from the spss output) and use these to
give a conclusion in the context of the question.

(d) Now reproduce these results in spss for yourself (the data file for this question
can be found in the Module 8 materials on the StudyDesk ). Note that you will
be required to perform some analyses in spss for your assignment.

Question 4: Is the die fair?

Bill, Clara and Zoe are playing a board game, but Zoe suspects that the die they are
using may have been damaged in the manufacturing process such that outcomes of
a roll, the numbers 1 to 6 are not equally likely (as one would expect they should be
with a fair die). She decides to test her suspicion by tossing the die 300 times and
obtains the following results after analysing her data in spss:

Figure 8.11: Output from the ‘Goodness-of-fit Test’ in spss.

(a) State the hypotheses to test Zoe’s suspicion.

(b) State any assumptions/conditions that need to be satisfied to conduct this test.

(c) What is the P -value for this test?


8.34 Module 8. Association between Categorical Variables

(d) Give a conclusion for this test in the context of the question.

Do the calculations for this hypothesis test by hand (show all working):

(e) Assuming the null hypothesis is true, calculate the test statistic.

(f) Calculate the degrees of freedom for this test.

(g) Determine the P -value for this test using the χ2 Distribution Table.

Question 5: Differences

How is the Chi-square Goodness-of-Fit Test different to the Chi-square Test of Inde-
pendence?

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

8.7 Module 8 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ identify independence and association between two categorical variables;

ˆ create a contingency table from data for two categorical variables;

ˆ determine and interpret joint, marginal and conditional distributions from a


contingency table;

ˆ carry out by hand and using spss a chi-square test of independence and interpret
the results;

ˆ carry out by hand and using spss a chi-square test of goodness-of-fit, and in-
terpret the results;

ˆ calculate the standardised cell residuals both by hand and by using spss after
a chi-square test and interpret them;

ˆ state and assess the assumptions and conditions necessary for the validity of a
chi-square test;

ˆ determine when it is appropriate to apply each chi-square test.


Module 9
Linear
Regression
Analysis
9.2 Module 9. Linear Regression Analysis

Contents
9.1 Inference for Simple Linear Regression . . . . . . . . . . . . . . 9.4
9.1.1 The Simple Linear Regression Model . . . . . . . . . . . . . . . . . 9.5
9.1.2 Assumptions & Conditions . . . . . . . . . . . . . . . . . . . . . . 9.5
9.1.3 Inference about Regression Parameters . . . . . . . . . . . . . . . . 9.7
9.2 Multiple Regression Model . . . . . . . . . . . . . . . . . . . . . 9.15
9.2.1 Parameter Estimation by OLS Method . . . . . . . . . . . . . . . . 9.16
9.2.2 Assumptions for a Multiple Regression Model . . . . . . . . . . . . 9.18
9.2.3 Inferences on the Regression Parameters βi . . . . . . . . . . . . . 9.20
9.2.4 Comparing Multiple Regression Models . . . . . . . . . . . . . . . 9.23
9.3 Closing Comments . . . . . . . . . . . . . . . . . . . . . . . . . . . 9.24
9.4 Tutorial Module 9 . . . . . . . . . . . . . . . . . . . . . . . . . . 9.25
9.5 Module 9 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 9.27
9.3

Module objectives
Upon completion of this module students should be able to:

ˆ define the difference between the sample regression line and the true regression
line;
ˆ state the assumptions necessary for inferential t procedures to be valid for simple
linear regression;
ˆ use appropriate graphs to check the assumptions where possible for the validity
of inferential regression procedures;
ˆ interpret the meaning of the intercept, slope and error standard deviation in
the simple linear regression model;
ˆ perform and interpret the results of a test about the slope using spss;
ˆ determine and interpret a confidence interval for the slope using spss;
ˆ determine, using spss, a confidence interval for the mean response at a given
value of x and a prediction interval for an individual response at a given value
of x;
ˆ use a simple regression model to predict a particular value of a future response;
ˆ determine the residuals and interpret the residual plot;
ˆ define a linear multiple regression model;
ˆ define the parameters of a multiple linear regression model;
ˆ state the assumptions of a multiple regression model;
ˆ determine estimates of the parameters of a multiple regression model;
ˆ find a best fitted regression line of a multiple regression model;
ˆ interpret the coefficients of a multiple regression model;
ˆ predict a particular value of a future response of a multiple regression model;
ˆ find the residuals and interpret the residual plot;
ˆ find a confidence interval for the parameters of a multiple regression model;
ˆ perform hypothesis tests on the parameters of a multiple regression model;
ˆ find the adjusted coefficient of determination and interpret it; and
ˆ check the assumptions of a multiple regression model.
9.4 Module 9. Linear Regression Analysis

Time Allocation
You should take no more than one week to complete this module. Make sure that
you keep up to date with the work as it is difficult to catch up should you get behind.

Introduction
The idea of simple linear regression and correlation were introduced in Module 3.
It may be worth reviewing that material before studying this module. We assume
familiarity with the idea of explanatory and response variables, the equation of a
regression line and of being able to determine, using spss, the correlation coefficient
and least-squares regression line and be aware of the potential impact of outliers
and influential observations on a regression line. We build on material presented in
Module 3 to discuss inferences involving simple regression and then extend this to
multiple regression.

De Veaux, Velleman & Bock, (5th edition) covers the material in this module in
Chapters 23 & 24.

9.1 Inference for Simple Linear Regression

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Section 1

The main idea here is that the points we see in a scatterplot of two quantitative
variables are just a sample from a population. The line we put through the scatterplot
may not be the same as the ‘true’ line – the one we would put through the population
of points if we knew what they were. The statistics b0 , b1 , yb and e are all to do with
the line through the sample, and we can calculate all of them from the data. The
parameters β0 , β1 , µy and  (notice they’re all represented by Greek letters) are all to
do with the population, and we don’t know the values of any of them! Inference about
regression is all about trying to figure out the values of these parameters. To do this,
we need to make some assumptions, including, as usual one concerning normality,
because we’re going to use the t distribution model again.
9.1. Inference for Simple Linear Regression 9.5

9.1.1 The Simple Linear Regression Model


Statistical models are probabilistic in nature and contain at least one random com-
ponent that is intended to explain the apparent random variability of the dependent
variable (say y) for a given value of the independent variable (say x). For example if y
is the weight of a person and x is the person’s height, we may postulate a probabilistic
model of the form

y = β0 + β1 x +  (9.1)

where  is the random error or variability of y about the relationship y = β0 + β1 X.

In this case, this random or unexplained variability would include effects on weight
not related to height, for example, bone density, body shape, etc. Thus we would
not expect that all people with the same height would have the same weight. So
although we have postulated a model it will not, in general, fit any particular set of
data perfectly.

To complete the statistical model we need to know the properties of  and relationship
between the response and explanatory variables.

9.1.2 Assumptions & Conditions

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Sections 2
and 3.

After this reading, write down answers to the following questions:

ˆ What assumptions about the response variable y are required to proceed with
inference calculations?

ˆ How can linearity be checked?

ˆ How can independence be checked?

ˆ How can equal variance be checked?

ˆ How can normality be checked?

ˆ How is R2 interpreted?
9.6 Module 9. Linear Regression Analysis

ˆ What symbol is used for the error standard deviation?

ˆ What is the formula for the error standard deviation?

ˆ Why is n − 2 used in this formula?

The simple linear regression model requires the following assumptions/conditions:

(i) Linear relationship: the response and explanatory variables must be linearly
related. This is often checked by visual observation of the scatter plot.

(ii) Independence of errors: the error terms are independent of each other. This
also implies that the responses are independent. This is checked by a residual
plot.

(iii) Equal error variance: the error variance is constant (or equal). This means
that the spread of the response (y) is the same for all values of the explanatory
variable (x). That is, V ar[] = σ 2 . This is checked by a residual plot.

(iv) Normal distribution of errors: the errors around the idealised regression line
follow a normal distribution for all values of x. This is checked by the histogram
or normal probability plot of the residuals.

The main point here is that whenever we do some inference there are assumptions or
conditions involved. We can anticipate the normality condition because we’ll be using
the t distribution again, as we did in Modules 5 and 6. And just as mentioned in
those modules, the more data we have the less important is the normality assumption.
Independence always appears somewhere in any assumptions. We are not saying that
the explanatory and response variables are independent – if they were we wouldn’t be
bothering with a regression at all! We’re saying that the n pairs of measurements of
the x and y values are independent. This often can’t be checked except by knowing
how the data was measured or collected. Often the best we can report is that we
assume the n pairs of measurements of x and y were made independently. The equal
variance assumption or equal spread assumption is a new type of condition. All it’s
saying is that the scatter of points in the vertical direction is approximately uniform
along the regression line. You may see this referred to as homoscedasticity.

The usual trick in checking assumptions for regression is to make good use of the
residuals. The routine suggested in Chapter 23, Section 2 of De Veaux, Velleman &
Bock, (5th edition) is fine. Of course, judgement is involved in deciding whether to
go ahead with inference. That won’t be a major issue for the problems asked in this
course – we’ll be pretty specific about what has to be done. In the bigger scheme of
things though, checking assumptions sets apart the user of statistics from the abuser
of statistics.
9.1. Inference for Simple Linear Regression 9.7

9.1.3 Inference about Regression Parameters


This section involves quite a few formulas. Many of them are irrelevant to us because
we don’t ever use them. The ones that you need to know will be pointed out. The
only sensible way to do regression inference is using technology – for us, that is by
using spss (although some of you may also have calculators that will do these things
too). It takes far too long to do all the calculations by hand and it’s impractical to
check assumptions without using a computer. In general, we solve problems by using
spss and/or by interpreting output from spss.

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Section 3.

After this reading, write down answers to the following questions:

ˆ What is the difference between se and sx ?

ˆ What is the difference between se and σ ?

ˆ What is meant by the standard error of the slope?

ˆ How many degree of freedom has the Student’s t-model when used in simple
regression?

ˆ What does β1 = 0 indicate about x and y in a regression problem?

ˆ What test statistic is used for H0 : β1 = 0?

ˆ What is the formula for a confidence interval about β1 ?

ˆ What are the two types of prediction intervals that can be produced? What is
the difference between them?

ˆ Why do we call one interval a confidence interval, and the other a prediction
interval?

Which of the formulas mentioned should we know about or are we going to be asked
to use? Firstly, from the previous section (and from Module 3), we should be aware
that se represents the standard deviation of the residuals and is given by
rP
(y − yb)2
se = (9.2)
n−2
9.8 Module 9. Linear Regression Analysis

This statistic describes the amount of scatter around the line (in the vertical direction)
and is a fundamental quantity. It’s given a few different names – the error standard
deviation, the residual standard deviation or the standard error of the estimate in
spss. If its value were zero then all the points would be on the regression line.
Otherwise, like any standard deviation or standard error it’s positive. The larger it
is, the more scattered the points around the line and the larger will be the margins
of error in our intervals. The n − 2 in the denominator of (9.2) is called the degrees
of freedom for se , and as such is the degrees of freedom for all simple regression
calculations. We need to know the formula for a confidence interval for the slope β1 :
b1 ± t∗n−2 × SE(b1 ) (9.3)

Also, we need to know how to calculate the test statistic for H0 : β1 = 0; that is
b1 − β1
t= (9.4)
SE(b1 )
There’s a formula for SE(b1 ) but we don’t use it. When we need to use (9.3) or (9.4)
we will be given SE(b1 ) somewhere (probably in some computer output). That’s it.
We don’t need to do any inference on β0 , and when asked for prediction intervals for
y or confidence intervals for µy , we should be aware that they have the form
yb ± t∗n−2 × SE (9.5)
but we won’t ever have to use these formulas. Usually we’ll just have to read the
interval from computer output.

An example will clarify things.

Example 9.1
Students in a Statistics tutorial measured forearm length and head circumfer-
ence in centimetres as shown:
Forearm Head Forearm Head
Subject length (cm) c’ference (cm) Subject length (cm) c’ference (cm)
1 43.5 60 13 48.0 58
2 40.0 57 14 44.0 58
3 44.0 57 15 47.0 59
4 52.0 59 16 47.0 59
5 46.0 58 17 38.0 56
6 46.5 62 18 41.0 58
7 43.5 56 19 43.0 56
8 44.5 57 20 47.5 57
9 40.0 57 21 45.0 59
10 43.5 56 22 46.5 56
11 46.0 57 23 49.5 58
12 46.0 63 24 48.0 59
9.1. Inference for Simple Linear Regression 9.9

Interest is in the relationship between forearm length and head circumference


and whether it makes sense to predict head circumference from forearm length.
Since we’re planning to predict head circumference from forearm length, the
explanatory variable must be forearm length (in cm) and the response must be
head circumference (in cm).
The first check is to determine the form of the scatterplot (see Figure 9.1). The
relationship is not strong although it is reasonable to assume it is linear. We
are therefore comfortable about fitting a linear regression model.

Figure 9.1: Scatterplot for regression of head circumference on forearm length.

Tick the ‘Save’ and ‘Plots’ options in the spss Regression procedure as indicated
in Figures 9.2 and 9.3. One of the results is that we get six more variables in
the Data View Window: ‘PRE 1’ (= yb), ‘RES 1’ (= e, the residuals), ‘LMCI 1’
and ‘UMCI 1’ (95% confidence interval bounds for µy ), and ‘LICI 1’ and ‘UICI
1’ (95% prediction interval bounds for y). Part of the new Data View is shown
in Figure 9.4.
Make a scatterplot of the residuals ‘RES 1’ (or ‘unstandardised residuals’ as
spss calls them) against the predicted values ‘PRE 1’ (‘unstandardised predic-
tions’ in spss). We get Figure 9.5. There is a lot of scatter but it’s plausible to
accept that there is no systematic thickening or thinning so the equal spread
assumption is reasonable. Also there are no obvious outliers. No information is
given as to the order in which the data was collected so the residuals cannot be
plotted against time. We can therefore only state there is no reason to assume
the observations are not independent.
9.10 Module 9. Linear Regression Analysis

Figure 9.2: ‘Save’ options in the SPSS Linear Regression procedure.

Figure 9.3: ‘Plots’ options in the SPSS Linear Regression procedure.


9.1. Inference for Simple Linear Regression 9.11

Figure 9.4: Data View after running the Regression procedure.

Figure 9.5: Residuals plotted against predicted values.

A check on the normality condition is given in Figures 9.6 and 9.7. Normality
is acceptable on the basis of these plots.
We feel justified now in interpreting the results of the regression procedure,
shown in Figure 9.8. Ignore the table titled ‘ANOVA’. Notice the correlation
coeffcient is r = 0.380 and R2 = 14.4% telling us that approximately 14% of the
variability in head circumference is explained by variability in forearm length
for these data. This confirms the description of the relationship as being weak
– 86% of the variability in head circumference is explained by other variables
apart from forearm length.
9.12 Module 9. Linear Regression Analysis

Figure 9.6: Histogram of residuals.

Figure 9.7: Normality plot of residuals.

The error standard deviation or ‘standard error of the estimate’, as spss calls
it, is se = 1.719 cm. This is really the estimated standard deviation of head
circumferences for individuals of a particular forearm length. A good way of
knowing whether se is ‘large’ or ‘small’ is to compare it with the standard devia-
tion of all 24 head circumferences that were measured. That value is sy = 1.818
9.1. Inference for Simple Linear Regression 9.13

cm which is only slightly larger than se . In other words, knowing somebody’s


forearm length only reduces the precision with which we can estimate their
head circumference from a standard deviation of about 1.82 cm to about 1.72
cm. Not much! The table headed ‘Coefficients’ tells us that b0 = 48.316 and
b1 = 0.215. Therefore the regression line is

yb = 48.316 + 0.215x (9.6)

Figure 9.8: Output from the Linear Regression procedure in SPSS.

where y = head circumference in cm and x = forearm length in cm. This is the


equation spss uses to give the column of predicted values in Figure 9.4. The
textbook sometimes writes this equation in the form

\
head circumference = 48.316 + 0.215 × forearm length

which is fine except it should be made clear also that head circumference and
forearm length are measured in centimetres. The standard error of b0 and b1 are
given in the ‘Coefficients’ table. The important one is SE(b1 ) = 0.112 because
this tells us how precisely the slope of the line has been measured which leads
9.14 Module 9. Linear Regression Analysis

to a confidence interval and hypothesis test. From 9.3 we can obtain a 95%
confidence interval for the true slope β1 as

b1 ± t∗n−2 × SE(b1 ) = 0.215 ± 2.074 × 0.112


= (−0.017, 0.447)

The most important thing about this interval is that it doesn’t exclude zero.
Therefore we are unable to reject H0 : β1 = 0 in favour of Ha : β1 6= 0 at the
5% level of significance. In other words, the P -value for this test will be bigger
than 5%.
It’s easy to confirm this from the ‘Coefficients’ table where we see t = 1.927 is
the value of the test statistic with (2-sided) P -value = 0.067. Putting this result
into words we can conclude from these data there is only weak or slight evidence
in favour of arguing that head circumference can be predicted from forearm
length. We should also be aware how the formula (9.4) applies. Substituting
b1 and SE(b1 ) into (9.4) gives

b1 − β1 0.215
t= = = 1.92
SE(b1 ) 0.112
as reported by spss (within round off error).
Predictions are produced by spss in the Data View as we see in Figure 9.4.
So, for example, a 95% prediction interval for head circumference of a student
with forearm length of 46 cm is (54.57, 61.86) cm. In other words, with 95%
confidence a student with forearm length of 46 cm can be expected to have a
head circumference somewhere between 54.6 cm and 61.9 cm.
In the same vein, with 95% confidence the mean head circumference of all
Statistics students with forearm length of 46 cm is expected to be somewhere
between 57.5 cm and 59.0 cm.
What say we want to make a prediction for an x value that is different to those
in the data? For example, a forearm length of 45.4 cm. All we need do is
include that value (or those values) of x at the end of the ‘x’ column in the
Data View and run the regression procedure as before. The analysis won’t
change but prediction intervals will be generated for the ‘new’ data. Check
yourself that a 95% confidence interval when x = 45.4 is (57.35, 58.82) and a
95% prediction interval is (54.45, 61.73).

Use and interpretation of regression parameters


1. Prediction
The estimated equation may be used to predict (or estimate) a value of y for a
given value of x. However it is not prudent to extrapolate outside the range of
the values of x which are given. In Example 9.1, with the x-values ranging from
9.2. Multiple Regression Model 9.15

38 to 52, it would be unwise to predict the head circumference if the student


had a forearm length of 34 cm. This is because the model may not be the same
outside the given range of the x-values.

2. Interpretation
b1 is the slope of the fitted regression line. It is the rate of change in ŷ for
one unit change in x. In Example 9.1, the increase in head circumference per
additional cm increase in forearm length is b1 = 0.215 cm.
b0 is the intercept of the fitted line. This is the estimated value of yb when x = 0.
Usually there is no or very little interest in it from a statistical point of view.
Here for example, it is nonsensical to say that forearm length of zero, relates to
a head circumference of 48.3 cm.

Exercise 9.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 23.1, 23.3, 23.5,
23.7, 23.9, 23.11, 23.23, 23.57 (use spss).

9.2 Multiple Regression Model

Reading
Read: De Veaux, Velleman & Bock, (5th edition), Chapter 23, Sections 5
& 6 and Chapter 24.

The multiple regression model and inferences on it are covered in this section.

In real life the response variable y depends on several explanatory variables. If there
are two or more explanatory variables in a linear regression model, it is called a multi-
ple regression model. Let us consider the linear regression of y on k(≥ 2) explanatory
variables, say X1 , X2 , . . . , Xk , then the multiple regression model is defined as

y = β0 + β1 x1 + β2 x2 + · · · + βk xk + , (9.7)

where

1. β0 , β1 , β2 , . . . , βk are unknown regression coefficients or parameters;

2. x1 , x2 , . . . , xk are known constants (explanatory variables measured without er-


ror);
9.16 Module 9. Linear Regression Analysis

3.  is the random error term.

Notes:
1. β0 , β1 , . . . , βk are estimated using n sets of observations
(yi , xi1 , xi2 , . . . , xik ), i = 1, 2, . . . , n where n ≥ k + 1. This will allow us to ‘fit’
the model to the data.

2. The values of the xi (i = 1, 2, . . . , k) may be continuous or discrete or categorical;


i.e., they may

(a) be measurements on a continuous variable such as weight; or


(b) be counts such as the number of rooms in a house; or
(c) arise from the use of categorical data such as presence or absence of a
particular trait, e.g., blue eyes, or indicate the extent or level of a particular
treatment, e.g., high pressure, medium pressure, low pressure.

However, in all of the above, the x-values are assumed to be measured without error.
If the x-variables are measurements or counts as in (a) and (b) they are said to
be quantitative variables, whilst those in (c) are called qualitative (or categorical)
variables. In this module, we will only consider quantitative variables.

9.2.1 Parameter Estimation by OLS Method


The regression parameters (β’s) of the multiple regression model in (9.7) are estimated
by Ordinary Least Squares.

The fitted model is written as

yb = b0 + b1 x1 + b2 x2 + · · · + bk xk , (9.8)

where b0i s are estimated coefficients/parameters for i = 0, 1, 2, · · · , k. This equation


is used to predict the value of y for a given set of values of x0i s.
9.2. Multiple Regression Model 9.17

Example 9.2

To predict the house price (response variable, y) one could use the floor area
(x1 ), number of bedrooms (x2 ) and number of bathrooms (x3 ) as explanatory
variables. The following data set is typical of multiple regression analysis.
Here all three explanatory variables are numerical/quantitative.

Figure 9.9: House Price Data (n = 10) for Multiple Regression Analysis

In the spss Regression procedure we now have 3 explanatory variables:

Figure 9.10: SPSS Multiple Regression Analysis


9.18 Module 9. Linear Regression Analysis

The spss Regression procedure produces the following output:

Figure 9.11: SPSS Multiple Regression Analysis

The fitted regression line is obtained by using the regression coefficients pro-
duced by spss as follows:
yb = 543.382 + 0.076x1 − 76.905x2 + 15.832x3 , (9.9)
where the intercept is b0 = 543.382, coefficient of x1 is b1 = 0.076, coefficient
of x2 is b2 = −76.905, and coefficient of x3 is b3 = 15.832.

Interpretation of regression coefficients/estimates


b0 = 543.382 is the intercept of the fitted regression line. So, yb = 543.382 if all
x0i s are zero.
b1 = 0.076 implies that the value of yb changes by 0.076 thousand dollars if the
value of living area changes by one sqft, given that the other two explanatory
variables are in the model.
b2 = −76.905 implies that the value of yb changes by −76.905 thousand dollars
if the value of number of bedrooms changes by one unit/room, given that the
other two explanatory variables are in the model.
b3 = 15.832 implies that the value of yb changes by 15.832 thousand dollars if
the value of number of bathrooms changes by one toilet/room, given that the
other two explanatory variables are in the model.

9.2.2 Assumptions for a Multiple Regression Model


The multiple linear regression model requires the following assumptions. These are
discussed in relation to Example 9.2.

(i) Linear relationship: The response and explanatory variables must be linearly
related. This is checked by visual observation of the scatter plot. From Fig-
ure 9.12, the house price has strong linear positive relationship with both living
9.2. Multiple Regression Model 9.19

area and number of bathroom, but negative very weak linear relationship with
number of bedrooms. Multicollinearity – linear relationship between pairs of
explanatory variables can be an issue. From Figure 9.12, we can see that mul-
ticollinearity is not an issue in this case as none of the explanatory variables is
linearly related to any other explanatory variable.

Figure 9.12: Scatter plots of House Price Data

(ii) Independence of errors: The error terms are independent of each other. This is
checked by a residual plot. From Figure 9.13, there is no trend or pattern in
the residual plot, and hence there is no concern of violation of this assumption.

(iii) Homoscedasticity: The error variance is constant (or equal), i.e., the spread of
the response (y) is the same for all values of the explanatory variable (x). This
is checked by the residual plot. From Figure 9.13, there is no funnel type shape
in the residual plot, and hence there is no concern of violation of the equal
variance assumption.

(iv) Normal error distribution: The errors around the idealised regression line follows
a normal distribution for all values of x. This is checked by the histogram
or normal probability plot of residuals. From Figure 9.14, the histogram of
residuals does not show any serious violation of the normality assumption. A
Q-Q plot also supports the same.
9.20 Module 9. Linear Regression Analysis

Figure 9.13: Residual plot of House Price Data

Figure 9.14: Histogram of Residuals of House Price Data

Since the assumptions are satisfied, we could proceed with the inference on the re-
gression parameters.

9.2.3 Inferences on the Regression Parameters βi


When a model has been fitted, point estimates of the regression coefficients βi0 s are
determined and an estimate of σ 2 will be found if (n > k + 1). Note that for the
house price example, n = 10 is larger than k + 1 = 3 + 1 = 4.
9.2. Multiple Regression Model 9.21

When the assumptions of the model hold good, one could find confidence intervals
and test hypotheses about the true values of the regression parameters βi .

To test H0 : βi = βi0 , any fixed value, against Ha : βi 6= βi0 , use the test statistic

(bi − βi0 )
t=
SE(bi )

with n − (k + 1) = 10 − 4 = 6 degrees of freedom.

A (1 − α) × 100% confidence interval for βi is given by

bi ± tn−k−1,α/2 × SE(bi ).

Example 9.3
Refer back to the house price data in Example 9.2.

1. Test that the living area is not significant, (i.e., β1 = 0) using a two-sided
alternative at α = 0.05; and
2. Find a 95% confidence interval for the coefficient of living area, β1 .

Solution

1. Here H0 : β1 = 0; H1 : β1 6= 0. Using the information in the spss


output in Figure 9.11 the observed value of the test statistic is

(b1 − β10 ) (0.076 − 0)


t= = = 3.2
SE(b1 ) 0.024

Again from the spss output, the two-sided P -value is 0.018 which is less
than α = 0.05, so H0 is rejected. That is, there is sufficient evidence to
claim that the β1 is significantly different from zero at the 5% level of
significance.
2. Note here df = n − k − 1 = 10 − 3 − 1 = 6, hence tα/2 = t6, 0.05/2 = 2.447.
A 95% confidence interval for β1 is

b1 ± tα/2 × SE(b1 ) or 0.076 ± 2.447 × 0.024 or (0.0173, 0.1347).

Test of significance of all coefficients – ANOVA


To test simultaneously that all the coefficients are zero, that is, test H0 : β1 = β2 =
· · · = βk = 0 against Ha : at least one coefficient is not zero.

This test is the same as the Analysis of Variance (ANOVA) F test.


9.22 Module 9. Linear Regression Analysis

Example 9.4
Refer back to the house price data in Example 9.2. spss Output from the
Regression procedure for this overall test is shown in Figure 9.15.

Figure 9.15: ANOVA of House Price Data

With a P -value of 0.006, the null hypothesis is rejected and so at least one co-
efficient is not zero (we have shown in Figure 9.11 that the non-zero coefficients
are β1 and β2 ).

Prediction using the fitted Regression Model


Again, refer back to the house price data in Example 9.2.

Example 9.5
Using the fitted regression model 9.9, predict the house price for a property with
living area x1 = 4000, number of bedrooms x2 = 3 and number of bathrooms
x3 = 2.5.
Solution
The predicted value of house price (b
y ) is given by

yb = 543.382 + 0.076 × 4000 − 76.905 × 3 + 15.832 × 2.5


= 543.382 + 304 − 230.715 + 39.58 = 656.247

So the estimated or predicted house price for a property with x1 = 4000, x2 = 3


and x3 = 2.5 is 656.247 thousand dollars. In the given data set the house price
for this property (third on the list) was 649 thousand dollars. So the residual
is e = y − yb = 649.000 − 656.247 = −7.247

Warning
Do not predict the response for values of the explanatory variables outside the range
of observed values as the linear relationship of the data may not be valid beyond the
observed data.
9.2. Multiple Regression Model 9.23

Comment The coefficient of number of bedrooms is significant at the 5% level since


the P -value associated with it is 0.026.
The coefficient of number of bathrooms is not significant at the 5% level since the
P -value associated with it is 0.632 (much higher than 5%. This indicates that number
of bathrooms may be dropped from the regression model.

9.2.4 Comparing Multiple Regression Models


One of the challenges in multiple regression analyses is to find the best model (that
explains maximum variation in the response variable) with a minimum number of
explanatory variables. One way to do this is to maximise the value of the coefficient
of determination R2 . The closer the value of R2 to 100% the better is the fitted
regression model.

Figure 9.16: Coefficient of determination of House Price Data

The value of R2 increases as the number of explanatory variables in the model in-
creases but models with too many explanatory variables are difficult to understand
and interpret. The adjusted
 R2 allows us to compare multiple regression models with
different numbers of explanatory variables.

If you have different models for the same response variable with different numbers
of explanatory variables, then the one with the largest value of Radj 2 is preferred.
Normally the explanatory variables with insignificant coefficients should be removed
from the model.

Again refer to the house price data in Example 9.2.

Example 9.6
Find the value of R2 and Radj
2 and comment on those values.

Solution
For the house price data, the Regression procedure in spss produces the model
9.24 Module 9. Linear Regression Analysis

summary as in Figure 9.16. Here the value of R2 = 0.861, that is, 86.1% of the
variation in house prices is accounted for by the three regressors or explanatory
variables in the model. As R2 does not adjust for the number of regressors in
the model, we would always use the adjusted R2 instead.
From the spss output, Radj2 = 0.792, so, we conclude that 79.2% of the variation

in house prices is accounted for by the explanatory variables living area, number
of bedrooms and number of bathrooms in the house.

Exercise 9.2
Do De Veaux, Velleman & Bock, (5th edition), exercises 23.13, 23.15,
23.17, 23.19, 24.3, 24.5, 24.7, 24.9.

9.3 Closing Comments


In this module we have introduced the concept of linear statistical models and dis-
cussed inferences about the parameters of those models.

You should be aware that linear statistical models play a major role in the study of
Statistics and that, if you proceed to the course in STA3301 Statistical Models, the
ideas introduced in this module will be expanded upon and models involving more
parameters and different assumptions about the error term will be considered.
9.4. Tutorial Module 9 9.25

9.4 Tutorial Module 9


The following section contains Tutorial 9 – a good summary of the work learnt in
this module and most importantly, testing your knowledge. You are expected to use
spss to do the calculations required in regression analysis.

Question 1: Investigate the following research question

Can the maintenance cost of vehicle be predicted by the distance driven?

The maintenance cost (in dollar) and distance driven (in hundred km) of 10 vehicles
are given in the following table.

1 2 3 4 5 6 7 8 9 10
Maintenance Cost 456 828 500 489 387 553 531 560 475 474
Distance Travel 112 173 160 127 137 124 153 143 112 123

Answer the following questions:

a) Write the equation for an appropriate simple regression model.

b) State all the parameters of the model.

c) State the assumptions of the simple regression model.

d) Enter the data into spss and run the Regression procedure. Use the spss output
to answer the following questions.

e) Write the equation of the best fitted model.

f) Interpret the estimated slope.

g) State the estimate of the error variance σ 2 .

h) State the standard error of b1 .

i) Predict the value of the maintenance cost if a vehicle drove 14300km distance.

j) Test the significance of the slope parameter at α = 5%.

k) Find the 95% confidence interval for the slope parameter.

l) Find the 95% prediction interval for a the maintenance cost of a single vehicle if
it drove 14300km.

m) Produce a residual plot.


9.26 Module 9. Linear Regression Analysis

n) Check the assumptions of the simple regression model with reference to the residual
plot.
o) Check the near normal distribution of the response variable.
p) State the value of the coefficient of determination.
q) Interpret the value of the coefficient of determination.
r) Comment on how ‘good’ the distance travelled is in predicting the maintenance
cost.

Question 2: Investigate the following research question

Can the GPA of students be predicted by SAT and HS Scores?

The data set [Link] contains information on the grade point average (GPA)
along with SATMath, SATVerbal, HSMath and HSEnglish scores of a random sample
of 30 students. In the context of a multiple regression analysis, run the Regression
procedure in spss and use the output to answer the following questions. Include
relevant spss output in your answers.

a) Write the equation of an appropriate multiple regression model for the data.
b) State all the parameters of the model.
c) State the estimated regression parameters.
d) Write the equation of the fitted model.
e) Test if the regression is significant.
f) Interpret the estimated coefficient of SATMath.
g) Find the estimate of the error variance σ 2 .
h) Test the significance of the regression parameter of SATMath at α = 5%;
i) Calculate and save the unstandardised residuals and report the first three values.
j) Check the assumptions of the multiple regression model using the residual plot.
k) Check the near normal distribution of the response variable.
l) Interpret the value of the adjusted coefficient of determination.
m) Comment on how ‘good’ the predictors are in predicting the GPA of students.

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.
9.5. Module 9 Checklist 9.27

9.5 Module 9 Checklist


Once you have completed the tutorial activity and textbook problems throughout the
module you should be able to:

ˆ identify the situation when a simple linear regression model is appropriate;

ˆ formulate a simple linear regression model and write the linear model equation;

ˆ state and verify the assumptions of a simple linear regression model;

ˆ find the best fitted simple linear regression model for a given data set;

ˆ calculate and interpret the residuals for a simple regression model;

ˆ identify and interpret the parameters/estimates of a simple regression model;

ˆ find and interpret the standard error of the estimators of simple regression
parameters;

ˆ find the coefficient of determination and interpret its value;

ˆ perform a test of significance of regression parameters of simple regression


model;

ˆ construct confidence interval about the parameters of the simple regression


model;

ˆ predict a single response for given values of explanatory variable(s);

ˆ find a confidence interval of the mean response for given values of explanatory
variable(s);

ˆ find prediction interval for a single response for given values of explanatory
variable(s);

ˆ formulate a multiple regression model identifying the response and explanatory


variables;

ˆ find and interpret the parameters/coefficients of a multiple regression model;

ˆ find the best fitted multiple regression model;

ˆ perform a test on the significance of the regression coefficients of a multiple


regression model;

ˆ find confidence interval of the regression coefficients of a multiple regression


model;
9.28 Module 9. Linear Regression Analysis

ˆ check the assumptions of a multiple regression model using appropriate graphs;

ˆ interpret the test of significance of all regression parameters of a multiple re-


gression model using ANOVA method;

ˆ find and interpret the coefficient determination of a multiple regression model;


and

ˆ check the assumptions of a multiple regression model using appropriate graphs.


Module 10
Synthesis and
Consolidation
10.2 Module 10. Synthesis and Consolidation

Contents
10.1 Sample size, P -values and Effect size . . . . . . . . . . . . . . . . 10.3
10.2 P-hacking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10.6
10.3 Type I and II errors and their relationship to Power . . . . . . 10.9
10.4 Parametric and Non-parametric tests . . . . . . . . . . . . . . . 10.10
10.5 Choosing the best test to use . . . . . . . . . . . . . . . . . . . . 10.11
10.6 Tutorial Module 10 . . . . . . . . . . . . . . . . . . . . . . . . . . 10.15
10.7 Module 10 Checklist . . . . . . . . . . . . . . . . . . . . . . . . . . 10.17
10.1. Sample size, P -values and Effect size 10.3

Introduction
In each of the weekly modules in this course we have worked through material related
to specific topics or statistical methods. Each Module built on the content covered
in previous modules, however the focus was always on the specific module material.
In this module we will expand on some important statistical concepts and will also
consider how to evaluate scenarios that may require deeper synthesis of materials
from across different modules. The majority of the content for this module is covered
in this study book.

10.1 Sample size, P -values and Effect size


In this course so far you have learnt that a small P -value is associated with a ’large’
test statistic. The test statistic (for each statistical test) represents how ‘far away’
the data is from the null hypothesis; so a large test statistic means that the sample
data represents a population that is quite different to the population described by
the null hypothesis. That is, a large test statistic is associated with a small P -value
because the probability that the null hypothesis describes our population is small.

In the formulae of each test statistic for each of the various statistical tests you should
recognise that the sample size (n) is included in the calculations. For example, n might
be used in the calculation of the mean and standard deviation or in the standard error.
Sample size is also a key component of determining the P -value through the use of
degrees of freedom. So, how would we expect a P -value to change as sample size
changes?

Generally, the larger a sample size, the more likely a study will find a relationship
or difference statistically significant if the relationship or difference actually
exists. We will come back to the bold part of the previous sentence in a moment,
but first, we will give an example to show the relationship between samples size and
P -values.

Example 10.1
If we have two samples from the same population with the same sample size
and standard deviation but different means such that:

ˆ Group1: Mean =12; Standard Deviation = 5; Sample size = 10, and


ˆ Group2: Mean =10; Standard Deviation = 5; Sample size = 10;

a two-independent-samples t-test can be used to determine if the means are


significantly different and would find the test statistic t = 0.894, df = 18 and
10.4 Module 10. Synthesis and Consolidation

P -value = 0.383 (38.3%). This indicates that there is no significant difference


(P -value > 0.05) between the means of the two groups (it doesn’t matter what
we are measuring for this example, but lets imagine we are measuring the
length of something so that we can give our means units of cm – 12cm and
10cm respectively with a difference of 2cm).
If we repeated the analysis with each group retaining the same mean and
standard deviation values but now with samples sizes of n = 50 in each group
then the test statistic, df and P -value all change: t = 2; df = 98; and P -value
= 0.048. This now indicates that there is a significant difference (P -value
< 0.05) between the group means.
What if we consider sample size of 100 per group? Then the two-independent-
samples t-test would now find t = 2.83, df = 198 and P -value = 0.005. The
test is now even more statistically significant (P -value< 0.01).
The size of the difference between the means has not changed (12 − 10 = 2cm),
however the P -value keeps decreasing just by increasing the sample size (notice
also that the test statistic increases as sample size increases – think about how
the calculation of the standard deviation may be influencing this). Should we
be concerned about this?
No. The test and the P -values are doing exactly what we would expect them
to do based on the underlying mathematics and what we need them to do in
order to help us come to a well supported decision.
As the sample size increases we are gathering more and more evidence about the
variation that naturally exists within our two groups; so we are more and more
confident that the 2cm difference between the group means is a true difference
between the groups and not the result of our sampling. Increasing sample size
reduces the impact of random error and measurements become more precise.

Does this relationship between sample size and small P -values mean that as long as
we have large samples that statistical tests will always give a significant result, even
when there is no real difference or relationship?

Earlier in this section we mentioned that generally larger sample sizes lead to signif-
icant results (small P -values) if the relationship or difference actually exists.
This bold section is important. If the null hypothesis (H0 ) is true and no difference
between the means actually exists then no increase to the sample size will produce a
significant result. The relationship only holds if H0 is false (and if the assumptions
of the test have been met). This is important because we do not want to incorrectly
reject a true H0 (Type I error).

However, this relationship between sample size and P -values does mean that as sample
size increases a statistical test can detect smaller and smaller differences between
means or subtler relationships between variables. Again, this is not a flaw in the
methods; we want the tests and P -values to work this way. Another example:
10.1. Sample size, P -values and Effect size 10.5

Example 10.2
In the first example, when the difference between the means was 2cm and the
sample size was 50 per group, the test was significant with P -value < 0.05. If
the difference between means was only 0.5cm such that:

ˆ Group1: Mean = 12, Standard Deviation = 5; and


ˆ Group2: Mean = 11.5 Standard Deviation = 5;

the sample size would need to be approximately 800 per group for the test to
detect a difference that small (ony 0.5) and be significant (P -value < 0.05).
This is a good and cautious approach; small differences should require more
supporting evidence (that is larger sample sizes) before we reject the H0 .
But what if a difference of 0.5 cm is not practically or contextually meaningful?
The difference between statistical and practical significance is why P -values
should always be treated as just one step in the final interpretation of any
analysis – they are not the only thing that should be considered and analyses
should always be interpreted in context of the initial question and hypotheses,
and the data.
Remember that statistical significance just means that the P -value was less
than some predefined alpha (α) level. The P -value is not an indication of the
importance of the result. Small unimportant differences, in a practical sense,
can be statistically significant. Practical significance refers to the magnitude
of the effect that has been measured, or the effect size. An understanding of
the smallest effect size that would have some practical meaning is important
when interpreting any statistical analysis and often we need to refer to previous
research or literature to help determine this.

It is important to remember that sample size is not the only thing that can influence
P -values in hypothesis tests. In either of the examples above, if we had also altered
the SD, the measure of variability around the means of each of our groups, this
also would have influenced the calculation of the test statistic and the subsequent
P -value. Think about how SD contributes to the calculation of the test statistic and
what would happen if it were larger or smaller.

We use inferential analysis (with statistical significance determined by P -values) to


help us make more objective rather than purely subjective conclusions; the same
analysis on the same data will produce the same result. Scientific or evidence based
decision making is generally not based on the results or P -values of one study alone,
but rather the cumulative evidence of repeatable and reproducible studies correctly
applying statistical methods.

Objectivity does not mean the removal of uncertainty. There is always some un-
certainty because we rely on samples to estimate statistics about our populations of
10.6 Module 10. Synthesis and Consolidation

interest. Uncertainty is why we always draw conclusions based on probability rather


than statements of fact.

10.2 P-hacking
P-hacking is a term used to describe attempts to find a significant result when pre-
ceding analyses were not significant. P-hacking can be in the form of increasing the
sample size, trying different methods of analysis, trying different subsets of the data
or adjusting the measurement scale of the data etc. As shown in Figure 10.1, each of
these approaches (as well as additional ones) are legitimate approaches to data prepa-
ration and analysis, but not if they are done only because prior approaches have not
provided a statistically significant result.

If we try varying the data and the analysis repeatedly only because we have not
achieved a statistically significant result, then we are inserting bias into the process.
If these are legitimate variations for a particular data set, then the variations should
be part of the initial analysis plan. If analysis is stopped once a significant result is
found then we never explore the effect of these variations, which results in selective
reporting of only significant results. This violates one of the ethical tenets described
in Module 2.

How to avoid or correct for the potential influence of P-hacking:

ˆ Perform careful data cleaning and descriptive analysis before starting any in-
ferential analysis to ensure that you have objectively considered the potential
influence of outliers or the scale of the variables etc.
ˆ Perform all checks to ensure test assumptions are met.
ˆ Where analysis of data begins without clear research questions or hypotheses,
or new questions arise during the analysis process, recognise that this form of
analysis is Hypothesis Generating rather than Hypothesis Testing. We should
ensure that this is clear in any reporting and label results as preliminary (need-
ing follow-up studies to add weight to the conclusions). The same sample data
set should not be used to both generate and test a hypothesis as the data is
biased towards the hypothesis.
ˆ Apply a correction where multiple testing or multiple comparisons of sub-groups
of the original data set is needed. Correction for testing multiple comparisons
may even be needed when comparison of a series of subgroups within the data
was the purpose of the original analysis plan. The xkcd comic in Figure 10.2 is a
good example of using multiple tests in search of significant results. We will not
cover the different approaches to correcting for multiple testing in this course.
However if you are curious google ‘Bonferroni correction’ for one example.
10.2. P-hacking 10.7

Figure 10.1: Forms of P-hacking.


10.8 Module 10. Synthesis and Consolidation

Figure 10.2: Multiple testing, Hypothesizing after the result is known, statistical
and practical significance, claiming causality. (Reprinted from [Link]
under the CC BY-NC 2.5 license.)
10.3. Type I and II errors and their relationship to Power 10.9

Figure 10.2 above is a good example of several analysis ‘problems’. Multiple testing
has been mentioned in the fourth point above. It is also an example of point 3: hy-
pothesizing after the result is known – claiming or giving the impression in reporting
that the significant result was the original hypothesis when in fact the process of find-
ing the significant outcome was part of a hypothesis generating process. In addition,
the final claim of the research seems to be drawing a causal link between green jelly
beans and acne. Think back to Module 2. What did you learn about our ability to
claim causal relationships? What type of study design would have been needed to
make this claim? We will come back to these questions in the Tutorial 10 questions
below.

10.3 Type I and II errors and their relation-


ship to Power
In Module 5 you learnt about Type I and II errors. A Type I error is the probability
of incorrectly rejecting H0 when it is true and a Type II error is the probability of
failing to reject H0 when it is actually false. Another related concept is Power. The
Power of a test is the probability that it will correctly reject a false H0 . In any
analysis we would hope to minimise Type I and II errors and maximise Power.

De Veaux, Velleman & Bock, (5th edition) gives a good review of this topic in Chapter
19 – read from the section titled ‘Errors’ to the end of the chapter.

The relationship between Power and Type I and II errors means that minimising one
can lead to a trade off with the others. In reality the only parameter we can overtly
control is the significance level, α which is the probability of incorrectly rejecting a
true H0 (Type I error). The standard accepted α is 0.05 (5%); that is, a 5% chance
of incorrectly rejecting H0 . We discussed earlier in this Module that increasing the
sample size will reduce the P -value which effectively also reduces α. However, as we
reduce the probability of a Type I error we increase the probability of a Type II error.
In words, as we reduce the chance of incorrectly rejecting H0 we make it harder to
reject H0 in general. So, if H0 is in fact false we have increased the chance that we
will fail to reject it.

Referring to Figure 19.2 in De Veaux, Velleman & Bock, (5th edition) we can see that
if we move the vertical line to the right to reduce α (probability of Type I error), β
(probability of Type II error) increases. As Power is defined as 1 − β, as α decreases
so does Power.

The relationship between Power and Type I and II errors can be a mind-bender and
when first learning these concepts it is easy to go around in circles a bit. Don’t worry,
we have all been there.
10.10 Module 10. Synthesis and Consolidation

10.4 Parametric and Non-parametric tests


Within the class of inferential quantitative methods, a distinction can be made be-
tween parametric and non-parametric methods. Parametric methods have an under-
lying assumption about the distribution of the variables, i.e. the assumption that the
variables are normally distributed. Remember that this assumption relates only to
scale variables where we can describe the distribution by the mean and variance.

Non-parametric methods are sometimes called distribution-free methods as they do


not require that the assumption of normality has been met. This means that they
can be used on ordinal and categorical/nominal variables. You have learnt about two
non-parametric tests in this course when you analysed categorical (nominal) data
using the Chi-square Test of Independence and the Chi-square Goodness of Fit test.

Additionally, when scale (quantitative) variables do not meet assumptions of nor-


mality then an equivalent non-parametric test could be used in situations where one
exists. This then requires that scale/quantitative variables be treated as ordinal or
nominal. When the data is ordinal the appropriate test considers the median as the
measure of central tendency rather than the mean.

Where the assumptions are met the parametric method is a more powerful test and
the preferred approach. However, when using inferential methods we prefer to always
err on the side of caution. In practice this means that we prefer that a non-significant
conclusion be drawn rather than incorrectly concluding a significant result. From a
statistical perspective this means that it is often ‘worse’ to make a Type I error than
a Type II error and so where assumptions are not met or are uncertain, the use of
a non-parametric test is preferable. As non-parametric tests have fewer assumptions
they place a larger ‘burden of proof’ on the data to demonstrate a clear pattern before
the test will indicate a significant result, so the use of these methods still minimises
the chance of a Type I error.

Figure 10.3 gives some examples of non-parametric tests that could be used when
a parametric test is not the best option. Other than the Chi-square tests already
mentioned, we will not cover any of these other non-parametric tests any further in
this course. Note that not all parametric tests have a non-parametric equivalent and
there is no non-parametric equivalent of confidence intervals.
10.5. Choosing the best test to use 10.11

Figure 10.3: Non-parametric tests equivalent to parametric tests (most of these


non-parametric tests are not covered in this course)

10.5 Choosing the best test to use


Often the hardest part of the statistical analysis process ends up being correctly
understanding the language rather than difficulties with maths!

In a teaching context we use different scenarios to present students with different


problems specific to the individual test or method we are teaching. But how do you
approach a problem when you are only given the context and no instructions on what
method should be used? Or how do you critically assess whether analysis produced
by someone else has been properly evaluated to be the best option for their data?

To determine the appropriate statistical methods that you should use, you will need
to consider the following:

(i) What types of data are the variables you wish to analyse? Are they categorical
(nominal), ordinal or quantitatve (ratio/scale)?
(ii) What are your dependent and independent variables?
10.12 Module 10. Synthesis and Consolidation

(iii) Should you use descriptive or inferential methods or both?

(iv) Should you use parametric or non-parametric methods?

Figure 10.4: How to analyse this data?

Figure 10.4 gives an idea of the broad level decisions that need to be made. In
practice, inferential methods should always be preceded by some descriptive methods
and generally both numeric and graphical descriptive methods are needed. Let’s look
at an example.

Example 10.3
Scenario: What is the difference (if any) in the average annual income of
accountants and real estate agents?
Although we don’t have a lot of information, we know enough about types of
studies, types of variables and different statistical methods to map out a quite
detailed approach to the analysis.

(i) The interest seems to be in comparing one group to another without any
treatment or intervention being considered. Therefore this data is proba-
bly collected from an observational study rather than an experiment.
(ii) Annual income would probably be measured on a continuous scale in
dollars or thousands of dollars. The sample data would need to include
values for one group (of accountants) and values for the other group (of real
estate agents). These would be different groups of people so each person
would belong to only one occupation and would only be ‘measured’ once.
Therefore, the data would be independent and not paired.
10.5. Choosing the best test to use 10.13

(iii) The scenario indicates that the interest is in the difference between av-
erage annual incomes. From the tests we have covered in this course we
could use a t-test or a confidence interval to consider differences between
means. We have already determined that the occupation groups are not
paired so we could consider two- independent-samples t-test with annual
income as the dependent variable and occupation type (accountant or real
estate agent) as the independent variable.
(iv) Before performing the t-test analysis, descriptive numerical summaries of
the mean and standard deviation of income for accountants and real estate
agents separately should be calculated. Histograms of the distribution of
yearly income for each group should also be considered to help identify
any outliers and evaluate the assumptions of the relevant t-test.
(v) If there are concerns about the assumption of normality not being met
then a non-parametric equivalent test could be considered.

We could also approach this exercise from another perspective.


Scenario: In order to determine if there is a difference in the average annual
income of accountants and real estate agents, histograms were used to explore
the data graphically and the correlation analysis and regression analysis were
performed.
Can you identify any flaws in this approach?

(i) Correlation should be performed on paired data not independent groups.


Two variables should be measured for each case and correlation analysis
used to consider the direction and strength of the linear relationship be-
tween the variables. Correlation would not help determine if there is a
difference in mean annual incomes between the groups.
(ii) Regression analysis would also not be appropriate to address the desired
question. Like correlation, regression is useful when the goal is to un-
derstand the linear relationship between the variables and to predict the
values of the dependant variables based on values of the independent vari-
able. In this scenario it is not clear which variable would be dependent on
the other – predicting accountant incomes from real estate agent incomes
does not make much sense and vice versa.
(iii) While histograms would be appropriate for this data they would not be
the best graphical summary to use for correlation or regression analysis
which would be better explored using a scatterplot.

An additional helpful resource is this YouTube video - it is less than 10min long.
Choosing which statistical test to use:
[Link]
10.14 Module 10. Synthesis and Consolidation

Now it is time to check your understanding by answering the following questions:

1. What effect can increasing sample size have on P -values?

2. What is the difference between statistical and practical significance?

3. What is P-hacking?

4. Is the Power of a test independent from the significance level?

5. Why are non-parametric tests sometimes described as distribution-free tests?

Exercise 10.1
Do De Veaux, Velleman & Bock, (5th edition), exercises 19.9, 19.25,
19.28, 19.30, 19.31.

If you require further explanation related to these questions, please ask in the Forum
on the StudyDesk.
10.6. Tutorial Module 10 10.15

10.6 Tutorial Module 10


The following section contains Tutorial 10 – a good summary of the work learnt in
this module and most importantly, testing your knowledge.

Question 1: Sample size – short answer

Discuss this statement: In order to show that there is a statistically significant differ-
ence between my control group and my treatment group (to prove that the treatment
is effective), I just need to make sure my sample size is big enough.

Question 2: P-hacking – short answer

Describe your understanding of the ethical concerns related to P-hacking and the
reporting of statistical results. You may need to refer back to the Supplementary
Material in Module 2.

Question 3: Power – short answer

Explain how Power and the significance level (α) are related.

Question 4: Which test?

For each of the following scenarios describe what descriptive and inferential analysis
(from those covered in this course) you think would be most appropriate and why?
Make sure you mention what type you think the mentioned variables are, what you
think are the dependent and independent variables and whether you think a non-
parametric method should also be considered.

(a) What is the difference (if any) in the average annual income between two business
types, dairy farmers (n=15) and banana growers (n=10)?

(b) What linear relationship (if any) exists between the angle of inclination of a beach
retaining wall and the amount of sand deposited along the beach, measured at
20 predetermined positions along the beach.

(c) Predict the concentration of nitrous oxides in smoke plumes from the height above
the smoke stack.

(d) What is the strength and direction of the linear relationship between the weight
of children under 6 years of age and their height?
10.16 Module 10. Synthesis and Consolidation

(e) What is the difference (if any) between the average scores of students on two
pieces of assessment within a statistics course?

(f) What is the relationship between voters party affiliation (Labour, LNP, Greens)
and their opinion on whether Australia should be a republic (Yes, No)?

(g) A sample of 20 cards has been drawn from the deck of cards. Was each suit
equally likely to be drawn?

Question 5: You never really understand something until you have to


explain it quickly and correctly

You need to describe to a colleague how a particular statistical test works but you
must do it within 100 words. Give your best description of each of the following with
as much important information included as possible.

(a) Two-independent-samples t-test

(b) Linear correlation

(c) Chi-square Test of Independence

Question 6: Critically evaluate

Scenario: Data was collected to determine if revenue for small businesses can be
predicted from the advertising spending (both variables measured in dollars).

Describe any problems you can identify in each of the following approaches to address
this scenario. Give detailed answers.

(a) A bar graph was used to compare the total advertising spending for each business
and then a paired samples t-test was used to see if there was a difference between
average revenue and average advertising spending.

(b) A histogram for advertising spending and then another histogram for revenue were
used to explore the data and remove outliers. A correlation was then performed
to determine the strength of the relationship between advertising and revenue.

(c) A histogram for advertising spending and then another histogram for revenue were
used to explore the data and then a two-independent-samples t-test was used to
see if there was a difference between average revenue and average spending on
advertising.

(d) A scatterplot was used to explore the relationship between the spending and
revenue variables. A confidence interval using two-independent-samples method
was then performed.
10.7. Module 10 Checklist 10.17

(e) A scatterplot was used to explore the relationship between the spending and
revenue variables and then regression analysis was used to model changes in
advertising spending (dependent variable) from changes in revenue (independent
variable).

The answers to the above tutorial questions will be made available on the StudyDesk
during the week in which they are taught.

10.7 Module 10 Checklist


Once you have watched the lecture recording, completed the tutorial activity and
textbook problems you should be able to:

ˆ state what a P -value actually represents and what factors influence its value;

ˆ describe the implications on the P -value when increasing or decreasing sample


size;

ˆ describe the difference between statistical and practical significance and why
this is important in research;

ˆ describe the importance of reporting effect size;

ˆ describe what P-hacking is and the consequences of this in research, including


the ethical considerations around good statistical practices;

ˆ describe the concept of the power of a test;

ˆ state the implications of Type I and Type II errors on the Power of a test;

ˆ define the difference between a non-parametric and parametric test and when
you should use each;

ˆ choose the correct statistical procedure for a given scenario;

ˆ describe and critically evaluate different scenarios.


10.18 Module 10. Synthesis and Consolidation

You might also like