Statistics
Statistics
Introduction to Statistics
Statistics is a branch of mathematics that is concerned with the planning, collection, organisation, analysis,
reporting of data and the interpretation of results. The aim of this module is to give you an introduction to
statistics. This module gives you an overview of different methods of displaying and organising data,
calculating measures of central tendency, calculating measures of spread and comparing like data.
Correlation and Regression using Excel will inform the understanding of linear relationships.
Collecting Data
When data are collected from every member of the group, a census is held. The group in this instance is
called the population.
When only selected members of the population contribute to the data, this is referred to as a sample of the
population and a survey takes place. Mostly, it is impractical and too expensive to obtain data from a
population so a sample is selected. Information from a sample is often used to predict information about the
population. This is the process used to predict the outcomes of elections.
The selection of a sampling method is very important. For example it may be important to collect data from
different geographical locations, different ages or different socioeconomic backgrounds. For quality control
in industry, systematic sampling, taking (say) every tenth item to check quality could be used. It takes
careful planning to conduct investigations with little or no bias. Bias is an unfair preference towards one
group which may lead to a distortion of the statistical results.
The researcher must also decide on sample size. If the sample size is too small, then the validity of the
finding could be in doubt. If the sample size is too large, the cost of obtaining the data may be prohibitively
high. It is possible to obtain an appropriate sample size using statistical processes that considers the accuracy
of results required with the number to survey to give that accuracy.
Statistical data can be of different types, the type of the data may determine the statistical processes that can
be undertaken. Data can be classified as either Categorical or Numerical. Categorical data are usually in word
form. Numerical data are usually in number form, however, some data in number form such as postcodes
could be considered Categorical data as no statistical analysis would be performed; for example; the average
postcode.
There are two types of numerical data: discrete and continuous. Most of this module is about numerical
data.
Discrete data are numerical data where the values are in set amounts, for example, the number of people in a
classroom. There could be 25, 26, 27 people in a classroom but not any amounts in between such as 25.72
people.
Page 1
[last edited July 2024]
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone
Continuous data are numerical data where the values can be any value, for example, the height of people. The
height of people can be measured, in theory, to any precision. The only limitation is the limitation of the
measuring instrument, but in theory the measurement can be to any precision.
(b) Number of students at a lecture – this is discrete because only a whole number of students could
attend.
(c) The number of children in a family – this is discrete because only a whole number of children is
possible.
(d) Age – this is continuous as age can be measured to any precision. For convenience, age in years may
be recorded – this is still continuous because age can be measured to any precision. The decision of
discrete vs continuous should be based on what is possible rather than what is chosen for convenience.
(f) Shoe size – this is discrete because shoe size has set values of half sizes and none in between.
Page 2
Numeracy
Module contents
Introduction Note: The tables in this section have not been
• Organising Data - Tables formatted for insertion in an assessment or
other publication. For table formatting requirements
• Graphs please refer to the style guide for your discipline
• Measures of Central Tendancy (e.g. APA7 or Harvard styles), and the guidance of
your Unit Assessor.
• Methods of Spread
• Comparing Like Data
• Correlation and Regression
Answers to activity questions
Outcomes
• Display data appropriately using charts and graphs.
• Organise data using tables.
• Calculate descriptive statistics for sets of data.
• Calculate correlation coefficient and equations of regression lines using Excel.
Data
1. For the data (right) about the Average Daily Hours of Location Latitude Av. Daily
Sunshine, calculate the mean, mode and median. °S Hours of
Sunshine
2. For the same data, calculate the range, interquartile Darwin 12.42 6.9
range and standard deviation. Brisbane 27.48 7.4
Perth 31.93 8.8
3. For the same data, use a 5 number summary to draw a Sydney 33.86 6.8
box and whisker plot. Adelaide 34.93 7.0
Canberra 35.3 7.2
4. Use Excel to do a scatterplot of Hours of Sunshine vs Melbourne 37.81 6
Latitude. Latitude is the independent variable Hobart 42.89 5.9
5. Use Excel to determine the correlation and the regression Macquarie 54.5 2.3
equation. Island
Page 3
[last edited July 2024]
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone
Numeracy
If the heights of 25 students are collected, then no table will be required. These 25 individual values can be
analysed without the need to organise the data.
If the grade point average of all the students at this university were to be analysed, then the data would need
to the organised into an appropriate table before any statistical analysis could take place. The structure of the
table will vary depending on the type of data (discrete or continuous) and the variation between lowest and
highest values. With spreadsheet computer programs such as Excel, the need to organise data using the
traditional table structures is not as important as in the past.
The first table to be considered is a Frequency Distribution Table. To assist in the development of ideas in
this module, some key examples will be used.
49 50 50 49 50 48 50 53 48 53
51 47 54 48 52 50 47 55 49 55
52 50 51 51 50 50 53 49 52 48
54 49 46 50 50 49 50 50 50 54
46 51 51 47 48 50 52 50 51 50
When the data are initially entered into a table, tally marks are used to record each piece of data. Tally marks
are small vertical lines. After four tally marks, the fifth is put through the four to make grouping of 5.
When recording the occurrence of each value, data should be systematically entered. It is easy to make mistakes
when tallying, so be systematic and check your totals twice. Watch the video below to see how to do this.
Page 4
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited July 2024]
After the data are entered into the table they will have the appearance below. The Frequency is the number
of occurrences of that value or how frequently it occurs.
49 ΙΙΙΙ Ι 6
51 ΙΙΙΙ Ι 6
52 ΙΙΙΙ 4
53 ΙΙΙ 3
54 ΙΙΙ 3
55 ΙΙ 2
Total 50
After this initial table is constructed, extra columns are added to the table to help with the summary data
required.
Commonly added columns are Relative Frequency, %Relative Frequency and Cumulative Frequency.
The Relative Frequency is the proportion of the data that has that value. This can be expressed as a decimal
or fraction. It is calculated by taking the frequency for each score and dividing by the total number of
scores. Find a total for this column.
The %Relative Frequency is the Relative Frequency made into a percentage. This is achieved by
multiplying the Relative Frequency by 100. Find a total for this column.
The Cumulative Frequency is a running total of the frequency. The cumulative frequency is found on any
line of the table by adding the frequency for that line to the total of frequencies for the previous lines. Do not
obtain a total for the cumulative frequencies as it is already a (running) total. The cumulative frequency
column is used extensively for finding the median and the quartiles (covered later).
Page 5
Number of Relative % Relative Cumulative
Frequency
Matches Frequency Frequency Frequency
2
46 2 = 0.04 0.04 x 100=4% 2
50
47 3 0.06 6% 2+3=5
48 5 0.1 10% 5+5=10
49 6 0.12 12% 10+6=16
50 16 0.32 32% 16+16=32
51 6 0.12 12% 32+6=38
52 4 0.08 8% 38+4=42
53 3 0.06 6% 42+3=45
54 3 0.06 6% 48
55 2 0.04 4% 50
Total 50 1.00 100%
From this table, answer to questions expressed in certain ways can be obtained.
For example:
1. How many boxes contained 51 matches? The frequency for this was 6, so there were 6 boxes that
contained 51 matches.
2. For what percentage of boxes contained 50 matches? 32% of the values were 50.
Page 6
Grouped Frequency Distribution Tables
When the difference between the lowest and highest scores is larger, groups may have to be used. Consider
the following table for data for amount spent by seventy children at a recent show.
Forming groups for continuous data is similar to those below, that is, in the form of:
or
The Grouped Frequency Distribution Table below was derived from a list of 70 values less than $100. The
groups were formed and the frequency of each group recorded in the frequency column. The next four
columns were calculated from the groups or their frequencies.
Page 7
Amount Group Relative % Relative Cumulative
Frequency
Spent$ Midpoint Frequency Frequency Frequency
$0 but less 2
2 5 = 0.029 0.04 x 100=4% 2
than $10 70
$10 but less
3 15 0.043 4.3% 2+3=5
than $20
$20 but less
5 25 0.071 7.1% 5+5=10
than $30
$30 but less
4 35 0.057 5.7% 14
than $40
$40 but less
2 45 0.029 2.9% 16
than $50
$50 but less
8 55 0.114 11.4% 24
than $60
$60 but less
14 65 0.2 20% 38
than $70
$70 but less
18 75 0.257 25.7% 56
than $80
$80 but less
12 85 0.171 17.1% 68
than $90
$90 but less
2 95 0.029 2.9% 70
than $100
Total 70 1.00 100
When a table like this is given, there are no original data! The original values are lost into the groups. This
is the big disadvantage of Grouped Frequency Distribution Tables. The only assumption that can be made
about the original values is that they are evenly spread throughout the group.
For example; the 18 values in the group ‘$70 but less than $80’ are assumed to be equally spaced out
between $70 and $80. The higher the number of children surveyed, the more likely this is to be true. This
idea is used when calculating some statistical measures in future topics.
If the data are discrete, the construction of the table is a little different.
The addition of a column labelled Group Boundaries is required for the construction of a frequency
histogram and polygon (a graph).
Page 8
The data below are the numbers of mp3 players sold per day during a 60 day sale.
5. How many players sold per day are represented by the 27th score (when put in number sold order)?
The frequency distribution table has ordered the data. For the group 21 - 30, the cumulative frequency
is 31. The previous group’s cumulative frequency is 16. This means that the 17th (the one after the
16th) box through to the 31st box between 21 – 30 days. As the 27th score is between the 16th and the
31st score, the 27th score will be between 21 – 30 players sold. A more detailed look at this will occur
later in this module.
18 22 41 19 30 31 27 20
32 27 31 25 35 24 19 40
35 32 44 37 17 20 45 32
23 27 34 19 47 33 24 41
39 26 29 30 44 24 28 32
Page 9
If the numbers are entered from the table above starting with the first row, working from left to right, the
Stem and Leaf will be:
Stem Leaf
1 8 9 9 7 9
2 2 7 0 7 5 4 0 3 7 4 6 9 4 8
3 0 1 2 1 5 5 2 7 2 4 3 9 0 2
4 1 0 4 5 7 1 4
3|2 means 32 years old
When constructing a Stem and Leaf Plot, make sure that the numbers are equally spaced. The length of each
leaf can give some general information about the data. Each table should have a statement that gives place
value to the data. Once obtaining this Stem and Leaf plot, it is now useful to order each of the leaves
(leafs!).
Stem Leaf
1 7 8 9 9 9
2 0 0 2 3 4 4 4 5 6 7 7 7 8 9
3 0 0 1 1 2 2 2 2 3 4 5 5 7 9
4 0 1 1 4 4 5 7
3|2 means 32 years old
Note: Stem And Leaf Plots are presented with leaves ordered – this is convention.
The advantages of a Stem and Leaf plot over a frequency distribution table are:
(i) The data are grouped and the original values are retained.
(ii) The data are in numerical order from the lowest value (top row on left) to the highest value
(bottom row on right). This will be very useful later on when calculating 5 number summaries.
If there is concern about the number of groups, that is, not enough groups, then each group can be
subdivided into two groups. There are various ways this is done, but the system used here is to replace the
existing 2 stem (values from 20 - 29) with the stems: 2 (values 20 - 24) and 2* (values 25 - 29).
Page 10
Now the data are spread over 7 groups instead of 4.
Stem Leaf
1* 7 8 9 9 9
2 0 0 2 3 4 4 4
2* 5 6 7 7 7 8 9
3 0 0 1 1 2 2 2 2 3 4
3* 5 5 7 9
4 0 1 1 4 4
4* 5 7
3|2 means 32 years old
Page 11
Activity
1. The table of data below represents the speed of 40 cars passing a school at 9am on a school day.
Car Speed Data (in km/hr)
12 41 44 45 28 40 32 62 46 25
31 35 31 20 59 27 49 19 58 38
22 50 46 14 33 48 25 32 52 69
40 52 57 27 61 42 39 64 52 27
(a) Enter the data into a grouped frequency distribution table. Include a cumulative frequency
column. Make the first ‘group 10 to but less than 20’.
(b) Enter the data into a Stem and Leaf Plot.
(c) What percentage of cars were doing 40km/hr or more?
2. The number of students attending a class (maximum 25) for 30 lessons is given in the table below:
Students Attending Class
25 24 24 25 24 23
25 24 23 24 25 25
24 25 20 23 25 24
22 24 25 24 23 21
25 23 24 25 24 22
3. The systolic blood pressures in mmHg (this is the higher value of the two blood pressure figures) of 30
patients are given in the table below.
Systolic Blood Pressures of 35 patients at a Cardiac Clinic
122 175 114 92 128 155
138 115 88 134 141 146
112 124 107 121 118 145
126 188 134 110 122 139
133 149 120 102 95 109
144 127 143 161 137
(a) Enter the data into a Stem and Leaf Plot. Use the key: 11|4 means 114.
(b) If hypertension (high blood pressure) is defined by a systolic blood pressure 140 or above, what
percentage of this group are suffering hypertension?
Page 12
4. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m).
(a) Enter the data into a Frequency Distribution Table. Your FDT must have at least 5 groups.
(b) How many have thrown less than 22m? (Use cumulative frequency to answer this)
(c) What percentage threw 21 to but less than 22m? (Use a % Relative Frequency Column).
Page 13
Numeracy
Please Note: The graphs in this section illustrate concepts only, not the
correct formatting for assessments. For formatting requirements
refer to the style guide for your discipline (e.g. APA7, Harvard) and
your Unit Learning Materials and/or consult your Unit Assessor
Topic 2: Graphs
The graphs you use to present information depend upon the nature and the type of data.
Below are two graphs; a bar graph for the gender breakdown in a Lecture Group and a column graph for the
number of matches in 50 boxes (covered previously). These were drawn using Excel. Notice the equal
spacing and thickness of the bars. The spacing of bars reflects the nature of discrete data.
15
Gender
Frequency
10
Male
5
0
0 5 10 15 20 25 30 46 47 48 49 50 51 52 53 54 55
Frequency Number of matches
Histograms
Numerical data can be used to construct a histogram. A histogram is basically a column graph with the bars
joined together. This reflects the nature of continuous data, however, discrete data can also be graphed as a
histogram
The graph below was drawn using Excel. Frequency is shown on the vertical axis. The horizontal axis
shows the amount spent at the show and is labelled with the class mark of each group. This is not ideal. The
preferred way to label the axis is to show the boundaries for each group on the bar, this is also shown below.
Page 14
[last edited July 2024]
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone
Spending Money at a Show
20
18
16
14
12
Frequency
10
8
6
4
2
Preferred labelling
0
Excel labelling 0 10 20 30 40 50 60 70 80 90 100
5 15 25 35 45 55 65 75 85 95
$Amount
A line called the frequency polygon can be drawn on the frequency histogram.
The line is drawn from the centre of each column to the centre of the next. It should start from the
horizontal axis from the centre of an imaginary group below and extend back to the horizontal axis to the
centre of an imaginary group above.
The polygon can be drawn with or without the histogram present. The polygon is drawn below.
10
8
6
4
2
0
0 10 20 30 40 50 60 70 80 90 100 110
$Amount
70
60
50
Cumulative Frequency
40
30
20
10
0
0 10 20 30 40 50 60 70 80 90 100
5 15 25 35 45 55 65 75 85 95
$Amount
Page 16
Line Graphs
Another important type of graph is the Line Graph.
Because none of the data collected for our scenario is suitable for a line graph, a simple example has been
made up.
% Pass Rate
2005 79 60
2006 77 40
2007 82 20
2008 85 0
2004 2005 2006 2007 2008 2009
2009 87 Year
bike
other
8%
12%
Car
50%
bus
18%
walk
12%
Visually, it is easy to say that most students travel to uni by car. With the percentages given, it is also
possible to quantify information from the graph.
For example: If there are 2000 students on campus today, approximately how many travelled by bike?
Number travelling by bike =12% of 2000 = 0.12 x 2000 = 240 students.
A variation of the pie graph is the percentage bar graph. Ideally a bar of length 100mm (or 200mm, 300mm
etc.) is divided by in the percentages given.
Page 17
Car 50% Bus 18% Walk Bike Other
12% 8% 12%
The graph below compares average weekly earnings in May 2007 for males and females working full time
in different states and territories.
1200
1000
800
Male
600
Female
400
200
0
NSW Vic Qld SA WA Tas NT Act
Location
Page 18
Activity
1. The height of a plant is measured every Friday morning for twelve weeks. Which type of graph would
be best to show the growth of the plant?
2. A ‘sound and vision’ shop sells CDs, DVDs, console games and computer games. Which type of
graph would best display the relative sales of the different items sold?
3. A student wishes to compare the number of motorcycle fatalities between the states. To get a more
accurate picture, data are collected for the years 2007, 2008 and 2009. Which type of graph allows
all data to be presented in one graph?
4. The number of students attending a class (maximum 25) for 30 lessons is given in the table below:
Number attending Tally Frequency
20 Ι 1
21 Ι 1
22 ΙΙ 2
23 ΙΙΙΙ 5
24 ΙΙΙΙ ΙΙΙΙ Ι 11
25 ΙΙΙΙ ΙΙΙΙ 10
Total 30
Construct a column graph to represent this data.
5. The Shot Put distances thrown by 27 world champion shot putters are given in the Frequency
Distribution Table below. The unit is metres (m).
Distance Tally Frequency Cumulative % Relative
Frequency Frequency
5
20 to but less than 20.5 ΙΙΙΙ 5 5 × 100 =
18.5%
27
2
20.5 to but less than 21 ΙΙ 2 7
27
× 100 =
7.4%
6
21 to but less than 21.5 ΙΙΙΙ Ι 6 13 × 100 =
22.2%
27
6
21.5 to but less than 22 ΙΙΙΙ Ι 6 19
27
× 100 =
22.2%
4
22 to but less than 22.5 ΙΙΙΙ 4 23
27
× 100 =
14.8%
2
22.5 to but less than 23 ΙΙ 2 25
27
× 100 =
7.4%
2
23 to but less than 23.5 ΙΙ 2 27
27
× 100 =
7.4%
Total 27 100%
(a) Construct a frequency histogram for this information. As a second step, put a frequency polygon
on the histogram.
(b) Construct a cumulative frequency histogram and then add an Ogive to the histogram.
Page 19
Page 20
Numeracy
Measures of Central Tendency are summary statistics that attempt to represent the data by summarising them
as a single, 'typical' value. There are three commonly used measures to describe the 'central' or 'typical' value,
the Mean, Median and Mode.
Mean
The mean is commonly referred to as the average.
The mean is the most common measure used. It is found by adding up all the values and dividing by the
number of values. Because of this, every value contributes to the mean.
For a set of values written as X1, X2, X3, X4, X5,………. Xn, the sample mean is calculated using the equation:
=X =
sum of the values ∑X
number of values in the sample n
The equation for the population mean uses some different pronumerals.
=µ =
sum of the values ∑X
number of values in the population N
Because the calculation of the sample and population means is the same (except for the pronumerals used in
the equation), calculators have just one key to cover both.
Page 21
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited July 2024]
Finding the mean of Individual Values
These data are the maths test results of a sample of 15 students in ascending order.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
X=
∑X
n
23 + 45 + 50 + ............. + 75 + 81 + 90
X=
15
X = 60.4
The mean can easily be calculated on a calculator, especially a scientific calculator with STATS mode
(covered later).
Now let’s look at the example about the money spent by children at the show. This was organised using a
Grouped FDT.
Page 22
Finding the mean is performed in a similar manner to above except the Group Midpoint is used to represent
the group.
Group f ×X
Frequency
Amount Spent$ Midpoint
(f)
(X)
$0 but less than $10 2 5 2 x 5 = 10
$10 but less than $20 3 15 3 x 15 = 45
$20 but less than $30 5 25 125
$30 but less than $40 4 35 140
$40 but less than $50 2 45 90
$50 but less than $60 8 55 440
$60 but less than $70 14 65 910
$70 but less than $80 18 75 1350
$80 but less than $90 12 85 1020
$90 but less than 190
2 95
$100
Total Σf =70 ΣfX =
4320
Remember this method assumes that the values in each group are evenly spread. This assumption is not
always true so the figure obtained for the mean using this method can slightly different to the mean if the
original values where used.
A scientific calculator is usually capable of performing this operation when STATS mode is used.
The instructions below apply to the Casio fx-82AU and the Sharp EL531 only. If you have a different
scientific calculator to this, you will need to read the user's guide for your calculator.
Page 23
3. Obtaining the statistical measure required.
Now your calculator is in STATs mode, the next stage is to enter the data. There are slight differences
depending on how the data are organised.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
Page 24
If you make a mistake, use the big round blue ‘Replay’
button to navigate back to the position of the incorrect
90 then press M+
value and enter the correct value. the calculator display will show
DATA SET=
15
DATA SET=
The display 15 informs you that 15
numbers have been entered into the STATs memory.
Now the data are entered, the calculated value of the sample and population standard deviations can be
obtained.
Mean = 60.4
Page 25
Finding the Mean from a Frequency Distribution Table
Let’s recall the data about the number of matches in 50 boxes which was organised in a FDT with no
grouping.
Number of
Frequency (f)
Matches (X)
46 2
47 3
48 5
49 6
50 16
51 6
52 4
53 3
54 3
55 2
Total Σf = 50
It is worth recalling that this table is informing us that there are 2 occurrences of the value 46, 3 occurrences
of 47, 5 occurrences of 48, ….
Instead of entering the value of 50, 16 times a modified entry process is used. The process for finding the
standard deviation is exactly the same.
With a Frequency Distribution Table, the data are entered into the Sharp EL531 calculator in the format:
score, frequency .
On the Casio fx-82AU, the table the values are entered into needs to be changed into a frequency like the
example above. To do this:
Page 26
On a Casio fx-82AU: On a Sharp EL531
Get the calculator into STATS mode: Each row from the table is entered as:
Press Mode . The display will be: 46 , 2 M +
DATA SET =
the calculator display will show 1
47 , 3 M +
DATA SET=
the calculator display will show 2
DATA SET=
The display 2 informs you that 2
rows have been entered into the STATs memory.
Press 2 for STAT
48 , 5 M +
DATA SET=
the calculator display will show 3
.
.
.
.
.
Press 1 for 1-VAR .
.
This time the calculator will display an X column and
a FREQ column. Enter a value followed by =.
55 , 2 M + the calculator display will show
Navigate to the frequency column using the big blue DATA SET=
REPLAY key. 15
Obtaining values for the mean is exactly the same process as in the previous section.
Mean = 50.24
Page 27
If the data are grouped, the process is almost identical except that the group midpoint is used.
Group
Frequency
Amount Spent$ Midpoint
(f)
(X)
$0 but less than $10 5 2
$10 but less than $20 15 3
$20 but less than $30 25 5
$30 but less than $40 35 4
$40 but less than $50 45 2
$50 but less than $60 55 8
$60 but less than $70 65 14
$70 but less than $80 75 18
$80 but less than $90 85 12
$90 but less than
95 2
$100
Total Σf =70
The data are entered into the calculator as Group Midpoint, Frequency or Group Midpoint; Frequency
depending on the calculator used.
Median
The median is the middle value of the data if the values are arranged in order.
As a measure of centre, it is a value based on its position; it is not influenced by the size of the values that
are above or below it.
Therefore it is quite different to the mean because it is based on position and for the mean, every value
contributes.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
When finding the median for a set of individual values, the first step is to order the values. Usually this is
done in ascending order, but descending order gives the same result.
The position of the median can always be found by using a simple expression:
n +1
The position of the median is .
2
Page 28
The median is the 8th value, which is 59. The median is the 8th value, so there are 7 values before it and 7
values after it.
23, 45, 50, 54, 55, 57, 59 59 59, 61, 63, 75, 75, 81, 90
7 values median 7 values = 15 values
Number of Cumulative
Frequency (f)
Matches (X) Frequency
46 2 2
47 3 5
The 17th through to the 32nd
48 5 10 values are in this group.
52 4 42
53 3 45
54 3 48
55 2 50
Total Σf =50
n + 1 50 + 1 51
As there are 50 values, the position of the median is: = = = 25.5
2 2 2
In this case the 25th and 26th values are required. The median is the average of the 25th and 26th values. As
the 25th and 26th values are both 50, the median is 50.
Now let’s look at the example using a Grouped FDT. Finding the median can be performed in one of two
ways: interpolation method or graphical method.
Page 29
(a) Interpolation Method (Amount Spent at a Show Data)
Group Cumulative
Frequency
Amount Spent$ Midpoint Frequency
(f)
(X)
$0 but less than $10 2 5 2
$10 but less than $20 3 15 5
$20 but less than $30 5 25 10
$30 but less than $40 4 35 14
This means that the 24th
$40 but less than $50 2 45 16 score is $60
$50 but less than $60 8 55 24
$60 but less than $70 14 65 38 This means that the 38th
$70 but less than $80 18 score is $70
75
There are 14 values in
56
The group width
$80 but =70-60
less than $90 12 this group.
85 68
=10
$90 but less than
2 95 70
$100
Total Σf =70
n + 1 70 + 1 71
As there are 70 values, the position of the median is: = = = 35.5
2 2 2
In this case the 35th and 36th values are required. Both of these values are located in the group ‘$60 but less
than $70’. This means that the median is greater than 60 but less than 70. This means that the median is
(35.5 – 24 =) 11.5/14 of the way through the ‘$60 but less than $70’ group. The median is:
Median = 60 +
( 35.5 − 24 ) ×10 (Group width)
14
Median = 68.2
n +1
-cf m −1
Lm +
2
Median = × Group width
fm
where Lm is the lower limit of the median group
f m is the frequency of the median group
cf m −1 is the cumulative frequency of the previous group
There are 70 values in the table. This is shown on the vertical axis (Cumulative Frequency).
This is to 35.
Page 30
The second step; from this position, move horizontally across the graph until the cumulative frequency
Ogive is found.
The third step is to move down (vertically) to the horizontal axis and read the value.
70
60
Cumulative Frequency
50
40
30
20
10
0
0 20 40 60 80 100 120
Amount Spent $
Remember this method assumes that the values in each group are evenly spread. This assumption is not
always true, so the figure obtained for the median using this method can slightly different to the actual
median if the original values where used.
Page 31
Mode
The Mode is the value that occurs the most often.
The mode may not exist because values only occur once.
The mode could even be quite different to the mean and median, it may be much higher or much lower
depending on the meaning of the data.
There may also be more than one mode. If there are two modes, the data are said to be ‘bimodal’.
8
6
4
2
0
0 1 2 3 4 5
Number of Children
From the graph, it can be observed that '1 child per household' has the highest frequency. This is the mode of
these values.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
When inspecting this data it can be seen that the 59 occurs 3 times.
This is the value that occurs the most frequently, that is, it has the highest frequency.
To find the mode, the frequency column is inspected. The highest frequency of 16 is related to a box
containing 50 matches. The mode of this data is 50 matches. Take care to give the mode as the score (50
matches) not the frequency (16).
Page 32
Number of
Frequency (f)
Matches (X)
46 2
The mode for this
47 3 data is 50 matches as
it has the highest
48 5 frequency (occurs
the most often)
49 6
50 16
51 6
52 4
53 3
54 3
55 2
Total Σf =50
When data are grouped, it is more relevant to refer to the modal group.
From the scenario, ‘Amount Spent at the Show’, the group with the highest frequency is '$70 but less than
$80' which has a frequency of 18.
It can be stated that the modal group is '$70 but less than $80'.
Group
Frequency
Amount Spent$ Midpoint
(f)
(X)
$0 but less than $10 2 5
$10 but less than $20 3 15
$20 but less than $30 5 25
$30 but less than $40 4 35
The modal group for
$40 but less than $50 2 45 this data is ‘$70 but
$50 but less than $60 8 55 less than $80’ as it
has the highest
$60 but less than $70 14 65 frequency (occurs
the most often)
$70 but less than $80 18 75
$80 but less than $90 12 85
$90 but less than
2 95
$100
Total Σf =70
Either the group ‘$70 but less than $80’ or the Group Midpoint ‘$75’ can be given as the mode.
Page 33
Which measure of centre should be used and when?
Every value from the data is used to calculate the mean. As long as the data are spread evenly, without excessive
variation, the mean is usually used. A lecturer would use the mean to find a measure of centre because the
marks would be usually spread out from (say) 30 to 100%. In this data the low or high values are not
excessively different to a measure of centre.
In real estate, the median is often used as a measure of centre because there are some excessively high
values that are quite different to the bulk of the market. Real estate sales generally consist of many
properties at the cheaper part of the market and fewer properties at the expensive part of the market. If the
mean was to be used, the fewer, more expensive properties would have a huge effect on the mean. Because
the median is a measure of centre based on location, variation in the price of high value properties and the
number of high value properties will have no effect on the median but would have a significant effect on the
mean.
In the clothing industry, the mode is often used as a measure of centre. If a shop sells mostly size 12
clothing, then size 12 is the mode of the sizes sold. The mean is of little use because it could be a value that
is not even a possible size. The median is a better measure but fails to reflect the nature of the sales
environment.
Page 34
Activity
1. The lifetime, in hours, of a sample of 15 light bulbs is: 351, 429, 885, 509, 317, 753, 827, 737, 487,
726, 395, 773, 926, 688, 485.
Calculate the mean, mode and median of the values. Do not organise into a table.
2. The number of children in a 10 families is: 1, 5, 2, 2, 2, 3, 1, 4, 3, 2. Calculate the mean, mode and
median of the values. Do not organise into a table.
3. For the Car Speed Data, calculate the mean, mode and median of the values after organising the data
in Frequency Distribution Table (The FDT can be found in the answers to the Topic ‘Graphs’).
4. The number of students attending a class (maximum 25) for 30 lessons is given in the table below:
(a) Calculate the mean, mode and median using a Frequency Distribution Table. (The FDT can be
found in the answers to the Topic ‘Graphs’).
(b) Discuss the appropriateness of each measure of central tendency as a typical value.
5. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m).
Calculate the mean, mode and median using a grouped Frequency Distribution Table. (The FDT can
be found in the answers to the Topic ‘Graphs’)
Page 35
Numeracy
Set A: 20, 35, 50, 65, 80 Set B: 40, 45, 50, 55, 60
The reason the sets are different is because of the spread of the values; the values in Set A are more spread
out than those in Set B.
Range
In statistics, not only is a measure of centre important but a measure of spread is also important. The most
basic measure of spread is the range.
Set A Set B
Range = Highest Value - Lowest Value Range = Highest Value - Lowest Value
Range = 80 – 20 Range = 60 – 40
Range = 60 Range = 20
From this it is possible to say that Set A is has a higher spread than Set B. The problem with the range is that
it uses the lowest and highest values. In any set of data, the lowest or highest score could be an odd or
unusual value.
Using odd or unusual values to measure spread will not produce a result that properly reflects the data. If
you consider students sitting an examination, a student not feeling well sits an exam and records an
unusually low score. This score is not typical of the group and so the range is much larger than it should be
for the group.
Scores that are atypical of the group are called outliers. There are ways of identifying outliers, but these will
not be covered in this unit.
Page 36
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited June 2023]
Finding the range of Individual Values
The following data are the maths test results of a sample of 15 students in ascending order.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
Once the data are ordered, the lowest and highest values are easy to locate.
47 3
48 5
49 6
50 16
51 6
52 4
53 3
54 3
55 2
Total Σf =50
When data are grouped it is more difficult to determine the highest and lowest values so an assumption must
be made.
Group
Frequency
Amount Spent$ Midpoint
(f)
(X)
$0 but less than $10 2 5
$10 but less than $20 3 15
$20 but less than $30 5 25
$30 but less than $40 4 35
$40 but less than $50 2 45
$50 but less than $60 8 55
$60 but less than $70 14 65
$70 but less than $80 18 75
$80 but less than $90 12 85
Page 37
$90 but less than
2 95
$100
Total Σf =70
The only assumption (because the original values are not present) that can be made here is that the lowest
value is $0 and the highest value is $100.
The range is 100 – 0 = 100
Interquartile Range
The next measure of spread is the Interquartile Range (IQR). It is still a range but it is between the first and
third quartiles. The first and third quartiles are values that are ¼ and ¾ of the way through the ordered data.
Another way to think about the IQR is to consider it the range of the middle 50% of the data.
The five values shown in this diagram are commonly referred to as a 'five number summary'. A five number
summary is very useful for calculating the range, IQR and for drawing a box and whisker diagram (covered
later).
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
Once again the data must be ordered. Previously a simple equation was used to locate the median. This is
similar.
n +1 3(n + 1)
th th
The first quartile is the score. The third quartile is the score.
4 4
Page 38
For the data above:
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
= Q3 − Q1
IQR
= 75 − 54
IQR
IQR = 21
The advantage of using the IQR is that the first and third quartiles are stable values; they are not outliers or
'odd' values.
Number of Cumulative
Frequency (f)
Matches (X) Frequency The 11th through to the 16th values are
49 matches, this includes the 12th and
46 2 2 13th values.
47 3 5
The first quartile is 49.
48 5 10
49 6 16
The 33rd through to the 38th values are
50 16 32 51 matches.
51 6 38 The 38th value is 51 matches.
52 4 42
53 3 45 The 39th through to the 42nd values are
52 matches.
54 3 48
The 39th value is 52 matches.
55 2 50
Total Σf =50
Page 39
n +1 3 ( n + 1)
4 4
50 + 1 3 × 51
= =
4 4
51 153
= =
4 4
= 12.75th value = 38.25th value
To find the 12.75th value, means finding the 12th and 13th values. In this case the 12th and 13th values are both
49, so the First Quartile is 49.
To find the 38.25th value, we have to calculate the value that is 0.25th (i.e. a quarter) of the way between the
38th value and the 39th value. In this case the 38th value is 51 and the 39th value is 52, so the Third Quartile is
51+ 0.25 x Difference between the 39th and 38th values, i.e. 52-51.
When data are grouped, the cumulative frequency is used to obtain the quartiles. There are two methods to
find the quartiles; the interpolation method and the graphical method.
Group Cumulative
Frequency
Amount Spent$ Midpoint Frequency
(f)
(X)
$0 but less than $10 2 5 2
$10 but less than $20 3 15 5 The 17th through to the 24th
values are found in the group
$20 but less than $30 5 25 10 ‘$50 but less than $60’
Page 40
First Quartile (Q1) Third Quartile (Q3)
n +1 3 ( n + 1)
4 4
70 + 1 3 × 71
= =
4 4
71 213
= =
4 4
= 17.75th value = 53.25th value
The first quartile is found in the group ‘$50 but less than $60’.
In this group there are 8 values (the 17th to the 24th) and the class width is 10.
The third quartile is found in the group ‘$70 but less than $80’.
In this group there are 18 values (the 39th to the 56th) and the class width is 10.
1.75 15.24
The first quartile (Q1 ) =50 + ×10 The third quartile (Q3 ) =
70 + ×10
8 18
= 52.19 = 78.47
= Q3 − Q1
IQR
=
IQR 78.47 − 52.19
IQR = 26.28
There are 70 values in the table. This is shown on the vertical axis.
To find the first quartile, come one quarter of the way up the vertical axis. This is to 17.5 (One quarter of
70). From this position, move horizontally across the graph until the cumulative frequency line is found,
then move down (vertically) to the horizontal axis and read the value.
To find the third quartile, come three quarters of the way up the vertical axis. This is to 52.5 (three quarters
of 70). From this position, move horizontally across the graph until the cumulative frequency line is found,
then move down (vertically) to the horizontal axis and read the value.
Page 41
Ogive of Money Spent
80
70
Cumulative Frequency 60
50
40
30
20
10
0
0 20 40 60 80 100 120
Amount Spent $
The graph value for the first quartile is (about) 52 and the third quartile is (about) 78, giving an IQR of 78 –
52 = 26.
Standard Deviation
The third measure of spread is the Standard Deviation. The key word in this measure is deviation. The word
deviation in this context is how much each value deviates (differs) from the mean.
(X − X )
2
2
X=
∑X 2 - 5 = -3 9
n
4 4 - 5 = -1 1
2+ 4+6+8
=
6 4 6-5=1 1
20
=
8 4 8-5=3 9
=5
∑( X − X )
2
=
20
The standard deviation of a sample is calculated using
∑ ( X −=
X)
2
20
=S = = 2.58
6.66666
n −1 4 −1
Page 42
If the set of data is a population, then the standard deviation has a slightly different equation. In this equation X
represents the raw scores, the population mean is represented by mu, while N is the population size.:
∑ ( X − µ )=
2
20
σ
= = =
5 2.24
N 4
A scientific calculator is usually capable of performing this operation when STATS mode is used.
The instructions below apply to the Casio fx-82AU and the Sharp EL531 only. If you have a different
scientific calculator to this, you will need to read the user's guide for your calculator.
23, 45, 50, 54, 55, 57, 59, 59, 59, 61, 63, 75, 75, 81, 90.
Page 43
On a Casio fx-82AU: On a Sharp EL531
Each number is entered into the list, 23 =, 45 =: Each number from the list is entered as:
23 then press M+
DATA SET =
the calculator display will show 1
45 then press M+
DATA SET=
the calculator display will show 2
.
.
.
If you make a mistake, use the big round blue ‘Replay’ .
button to navigate back to the position of the incorrect .
value and enter the correct value.
90 then press M+
DATA SET=
the calculator display will show 15
DATA SET=
The display 15 informs you that 15
numbers have been entered into the STATs memory.
Now the data have been entered, the calculated value of the sample and population standard deviations can be
obtained.
Page 44
Now press 4 to get more options. RCL Σx the symbol Σx means the sum of the
numbers entered.
It is worth recalling that this table is informing us that there are 2 occurrences of the value 46, 3 occurrences
of 47, 5 occurrences of 48, ….
Instead of entering the value of 50, 16 times a modified entry process is used. The process for finding the
standard deviation is exactly the same.
With a Frequency Distribution Table, the data are entered into the Sharp EL531 calculator in the format:
score, frequency .
On the Casio fx-82AU, the table the values are entered into needs to be changed into a frequency like the
example above. To do this:
Page 45
Press SHIFT MODE . Eight options will be displayed.
DATA SET=
The display 2 informs you that 2
rows have been entered into the STATs memory.
Press 2 for STAT
48 , 5 M +
DATA SET=
the calculator display will show 3
.
.
.
.
.
Press 1 for 1-VAR .
.
This time the calculator will display an X column and
a FREQ column. Enter a value followed by =. 55 , 2 M + the calculator display will show
Navigate to the frequency column using the big blue DATA SET=
REPLAY key. 15
Page 46
Obtaining values for the population and sample standard deviation is exactly the same process as in the
previous section.
If the data are grouped, the process is almost identical except that the group midpoint is used.
Group
Frequency
Amount Spent$ Midpoint
(f)
(X)
$0 but less than $10 5 2
$10 but less than $20 15 3
$20 but less than $30 25 5
$30 but less than $40 35 4
$40 but less than $50 45 2
$50 but less than $60 55 8
$60 but less than $70 65 14
$70 but less than $80 75 18
$80 but less than $90 85 12
$90 but less than
95 2
$100
Total Σf =70
The data are entered into the calculator as Group Midpoint, Frequency or Group Midpoint; Frequency
depending on the calculator used.
Page 47
Activity
1. The lifetime, in hours, of a sample of 15 light bulbs is: 351, 429, 885, 509, 317, 753, 827, 737, 487,
726, 395, 773, 926, 688, 485.
Calculate the range, inter-quartile range and standard deviation of the values. Do not organise into a
table.
3. For the Car Speed Data, calculate the range, inter-quartile range and standard deviation of the values
after organising the data in Frequency Distribution Table (This can be found in the answers to the
Topic ‘Graphs’).
4. The numbers of students attending a class (maximum 25) for 30 lessons are given in the table below:
Students Attending Class
25 24 24 25 24 23
25 24 23 24 25 25
24 25 20 23 25 24
22 24 25 24 23 21
25 23 24 25 24 22
Calculate the range, inter-quartile range and standard deviation using a Frequency Distribution Table.
(This can be found in the answers to the Topic ‘Graphs’).
5. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m).
Calculate the range, inter-quartile range and standard deviation using a grouped Frequency
Distribution Table. (This can be found in the answers to the Topic ‘Graphs’)
Page 48
Numeracy
Before continuing, it is important to make a point about the limitations of this process. When samples of
data are collected there is always random variation between samples. For example; an environmental
scientist collected tadpoles from two different ponds each day for a week. The number of tadpoles in the
first pond was lower than the second pond. It was suspected that the first pond receives water runoff from an
industrial area which is polluted. The lower number of tadpoles in the first pond could be explained by one
of two explanations: (a) chance variation or (b) the first pond was polluted. Perhaps if the difference is
small, the variation is chance variation. If the difference is large, the variation is due to pollution in the first
pond. What is a small or large variation?
In this module it is not possible to separate chance variation from variation due to some effect. In most first
year university statistics courses the topic ‘Hypothesis Testing’ is covered. Hypothesis testing is a process
that gives some certainty to situation outlined above.
Brand A Brand B
53 40 37 41 40 56 56 60 50 52 27 18 60 97 35 79 55 73 44 68
58 55 53 52 54 59 59 64 45 42 77 84 93 84 61 78 69 74 55 78
For each brand a 5 number summary is determined. This was mentioned in the section on ‘Interquartile
Range’. Diagrammatically, it is represented as below:
A five number summary is necessary for drawing a box and whisker diagram.
One way to organise this data is to use a back to back stem and leaf plot.
This method of organising data works well here because the data are easily put into a stem and leaf plot
(see section 2 Tables) and some trends about the data may be seen from the plot.
Page 49
Student Learning Zone | +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited July 2024]
Brand A 5th value Stem Brand B
th
6 value 1 8 5th value
6th value
2 7
17th value
7 3 5
52100 4 4 10th value
18th value 998665543322 5 55
40 6 0189 11th value
11th value
7 347889
18th value
10th value 8 44
9 37 17th value
2|4 means 24 hours
Lowest First Quartile is the Median is the Third Quartile is the Highest
Value = 37 n +1 n +1 3 ( n + 1) Value = 64
4 2 4
20 + 1 20 + 1 3 × 21
= = =
4 2 4
=
21
=
21 = 17.75th value
4 2 This means that the third
= 5.25th value = 10.5th value quartile is three quarters of
This means that the first Median = the way between the 17th and
quartile is a quarter of the way (53+54)÷2=53.5 18th values.
between the 5th and 6th values.
First Quartile Third Quartile
= 42 + 0.25 x (45- = 59+0.75x(59-59)
42) =59
= 42.75
Lowest First Quartile (Q1) = Median = Third Quartile (Q3) Highest
Value = 37 42.75 53.5 = 59 Value = 64
1) Draw short vertical lines at the 5 values from the 5 number summary.
3) Connect the first quartile to the lowest value with a whisker. Likewise the third quartile to the
highest value. These should be located in the middle between the top and bottom of the
box. You will notice that the median lies somewhere within the box encompassing the
interquartile range.
Page 50
Five value summary of brand A data used in constructing box and whisker plot:
37 42.75 53.5 59 64
15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100
Lowest First Quartile is the Median is the Third Quartile is the Highest
Value = 18 n +1 n +1 3 ( n + 1) Value = 97
4 2 4
20 + 1 20 + 1 3 × 21
= = =
4 2 4
=
21
=
21 = 17.75th value
4 2 This means that the third
= 5.25th value = 10.5th value quartile is three quarters of
This means that the first Median = the way between the 17th and
quartile is a quarter of the way (69+73)÷2=71 18th values.
between the 5th and 6th values.
First Quartile Third Quartile
= 55+ 0.25 x (55- = 84+0.75x(84-84)
55) =84
= 55
Lowest First Quartile (Q1) = Median = Third Quartile (Q3) Highest
Value = 18 55 71 = 84 Value = 97
18 55 71 84 97
15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100
Page 51
Now to do the comparing the Box and Whisker Plots should be arranged side by side.
18 55 71 84 97
Brand B
Brand A
37 42.75 53.5 59 64
15 20 25 30 35 40 45 50 55 60 65 70 75 80 85 90 95 100
Comparing:
1. Based on the lengths of the overall plots and the lengths of the boxes, Brand B batteries are more
variable in lifetime than Brand A.
2. Based on the first quartiles, median and third quartiles, it is possible to state that Brand B will mostly
last longer than Brand A. The lowest value for Brand B is lower than the lowest for Brand A. This is
not a major consideration as lowest or highest values may be ‘odd’ values. Because the lower quartile,
median and upper quartiles are free of ‘odd’ values, they are the most reliable values to base
generalisations on.
Page 52
Activity
1. The lifetime, in hours, of a sample of 15 light bulbs is: 351, 429, 885, 509, 317, 753, 827, 737, 487,
726, 395, 773, 926, 688, 485.
2. The Car Speed Data were collected from a school zone just prior to an advertising campaign. The data
collected were:
After an advertising campaign to make motorists more aware of speeding in school zones, the
following data were collected.
22 45 42 27 27 40 45 26 75 25
31 48 30 16 28 42 19 55 46 38
34 50 14 59 33 42 30 32 36 18
38 41 50 14 31 39 25 46 39 27
Draw side by side box and whisker plot to determine if the campaign was successful.
3. The heart rate of a group of athletes was compared to the heart rate of a group of office workers after
climbing a set of stairs. The data are presented below:
Athletes
96 111 88 79 101 91
104 106 121 93 103 96
85 117 126 97 83 93
112 106 110 116 91 88
Office Workers
114 107 94 113 113 97
118 103 121 127 117 131
145 132 118 108 145 126
100 138 120
Draw side by side box and whisker plots and make generalisations about the two groups.
Page 53
Numeracy
Once the correlation coefficient suggests a reasonable linear relationship exists, the linear relationship can
be described by an equation. The process of finding the “theoretical” or “ideal” linear equation is called
Regression.
The scenario used as an example is relevant to university study. Many studies have shown that time on task
is an important indicator to the overall success of students at university. In this example the success of
students on an examination (as a percentage) will be one variable and the hours spent studying will be the
other variable. The data will be entered directly into Excel. The spreadsheet is shown below:
Page 54
Student Learning Zone | +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited July 2024]
The next step is to draw the graph of the data. For the scenario, the dependent variable is Exam Result%
because the underlying relationship could be that Exam Result% depends upon the Study Time. Study Time
is the independent variable. When two variables are graphed together, the graph is called a scatter plot or
scatter diagram.
To graph the data from above
>Highlight the data in the table
>Click on the Insert tab
> In the Charts Area click on Scatter
> Choose the option that plots points (not lines)
Strong correlation is indicated by the points making a straight line. The straight line may have a positive
slope, indicating positive correlation or a negative slope, indicating negative correlation. If the points form a
line with slope close to zero or seem to be randomly arranged, then there is no correlation.
For positive correlation, it is possible to say 'when the value of one variable increases, the value of the other
variable will also increase'.
For negative correlation, it is possible to say 'when the value of one variable increases, the value of the other
variable will decrease'.
100
90
80
Exam Result%
70
60
50
40
30
0 5 10 15 20 25 30 35
Study Time (Hours)
The appearance of the graph suggests positive correlation. Generally speaking this means that as the Study
Time increases there is an increase in the Exam Result.
Correlation has a number associated with it called the correlation coefficient. The correlation coefficient
ranges from -1 to 1. A correlation coefficient of -1 represents strong negative correlation, 0 represents no
correlation, 1 represents strong positive correlation. The following table shows the link between the
appearance of graphs, the correlation, possible correlation values and the link between the two variables.
Page 55
Graph Correlation Correlation Link between variables
description coefficient
10 The greater the value of
8 the x variable, the
6 greater the value of the y
4
Strong variable.
0.8 to 1
2
Positive
0
0 5 10
Page 56
10 The greater the value of
8 the x variable, the lower
6 the value of the y
4
Strong variable.
-1 to -0.8
2
Negative
0
0 5 10
Looking at the table above and comparing to the graph for our scenario, it is clear that there is some
correlation. The wording ‘Quite Strong Correlation’ is almost suitable, so a term to suggest slightly lower
correlation could be ‘Reasonable Correlation’. The correlation coefficient could be estimated to be about
0.7; overall the correlation could be described as Reasonable Positive Correlation.
The data must be present in a table. Excel can give a graph for the visual assessment of the correlation and
the CORREL function will give the correlation coefficient.
Enter array 1 by going to the sheet and highlighting the Exam Result values with no heading.
Enter array 2 by going to the sheet and highlighting the Study Time value with no heading. The
spreadsheet should look like the spreadsheet below.
Page 57
The correlation coefficient was 0.892; there is quite strong positive correlation meaning that there is
evidence to suggest that the greater the amount of Study Time, the higher the amount of Exam Result will
be.
The correlation coefficient confirms the relationship between the variables is stronger than the graph
suggests.
In the graph above, Excel has plotted the regression line. The spreadsheet can be modified to include the
slope and y-intercept of the regression line. From this, the equation can be obtained. (You may need to brush
up on Equations of Straight Lines from the Linear Relationships module.)
In cell B20, write the words “Slope =” and in cell B21 write the words “Y-intercept =” The slope is found
by doing:
Click on the cell C20, click on the Formulas tab
> click on More Functions
> click on Statistical
Come down the list to SLOPE and click.
Enter y values (Exam Result%) by going to the sheet and highlighting the exam result values with no
heading.
Enter x values (Study Time) by going to the sheet and highlighting the study time values with no heading.
Enter y values (Exam Result%) by going to the sheet and highlighting the exam result values with no
heading.
Page 58
Enter x values (Study Time) by going to the sheet and highlighting the study time values with no
heading.
Note1: In Business text books, the equation may be written as R = 31.6 T+ 1.86
Note2: In some university units, students are required to calculate correlation coefficients and the
regression equation using formulas. Scientific calculators and computer spreadsheets use these formulas
in these calculations.
The purpose of describing the relationship between the variables as an equation is to predict values.
For example: If a student studies for 24 hours, what exam result would be expected?
=R 1.86T + 31.6
R = 1.86 × 24 + 31.6
R = 76.24
Because of the least-squares method used to calculate regression lines, the value of the y variable (exam
result) can be calculated from the value of the x variable (study time) but not the other way round. This
means that calculating the hours of study required to get a result of (say) 90% should not be attempted. For
Page 59
this to take place, a new regression equation would have to be recalculated giving study time as the
dependent (y) variable and exam result as the independent (x) variable.
In summary,
(a) The strength of the relationship between two variables is called correlation.
(b) If the graph indicates strong positive or negative correlation supported by the correlation coefficient
then determining the regression is worthwhile.
(c) The regression equation can be used to predict the dependent variable from a given independent
variable.
(d) Prediction should only be within the range of independent values that were used to calculate the
regression equation. In this example the regression equation was based on independent variable values
from 5 to 30 hours. Predicting within the range of independent values is called ‘Interpolation’.
(e) Prediction outside the range of independent values is called ‘Extrapolation’. Dependent values found
by extrapolation should be considered unreliable and this practice should be avoided.
(f) The regression equation only applies to this exam with a cohort of student similar to those used to
form the equation.
Page 60
Activity
1. The table below contains information about different types of milks. The information given is the
grams of fat and the energy of the food in kilojoules.
(a) Enter the data into Excel and produce the scatterplot. The hypothesis is ‘the Energy in the milk
depends on the amount of Fat’. Energy is the dependent variable. Comment on the correlation by
viewing your scatterplot.
(b) Modify your spreadsheet to calculate the correlation coefficient. Comment on this.
(c) Modify your spreadsheet to calculate the slope and y-intercept of the regression line.
2. A large company is making note of the amount spent on advertising each month and then comparing
this with the amount of monthly sales. The data are given in the table below:
(a) Enter the data into Excel and produce the scatterplot. Comment on the correlation by viewing
your scatterplot.
(b) Modify your spreadsheet to calculate the correlation coefficient. Comment on this.
(c) Modify your spreadsheet to calculate the slope and y-intercept of the regression line.
3. Information about climate is given below for towns on roughly the same latitude but varying
longitudes. The longitude is not given but distance inland (east) of the coastal location of Yeppoon is
Page 61
given. The purpose of the question is to determine if there is a correlation between the climate of a
location and its distance from the sea. (Distance Inland is the independent variable)
Using Excel, find the correlation coefficient of the Distance Inland vs the three climate data given.
Calculate the regression line as appropriate and make comments about what other information could be
applicable in this situation
Page 62
Numeracy
3. For the Latitude Data, use a 5 number summary to draw a box and whisker plot.
L=2.3 Q1=5.95 M=6.9 Q3=7.3 H=8.8
2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 7.5 8.0 8.5 9.0
4. Use Excel to do a scatterplot of Hours of Sunshine vs Latitude. Latitude is the independent variable.
Daily Hours of Sunshine for 9 Locations
10
9
Av Daily Hours of Sunshine
8
7
6
5
4
3
2
1
0
0 10 20 30 40 50 60
Latitude (°S)
Page 63
Student Learning Zone +61 2 6626 9262 | learningzone@[Link] | [Link]/learning-zone [last edited July 2024]
Correlation = - 0.69 Moderate negative correlation
Regression line:
H= −0.11L + 10.2
where L is the latitude
H is the av. daily hours of sunshine
Cumulative
Speed Tally Frequency
Frequency
10 to but less than 20 km/hr ΙΙΙ 3 3
20 to but less than 30 km/hr ΙΙΙΙ ΙΙΙ 8 11
30 to but less than 40 km/hr ΙΙΙΙ ΙΙΙ 8 19
40 to but less than 50 km/hr ΙΙΙΙ ΙΙΙΙ 10 29
50 to but less than 60 km/hr ΙΙΙΙ ΙΙ 7 36
60 to but less than 70 km/hr ΙΙΙΙ 4 40
40
Page 64
(b) Enter the data into Stem and Leaf Plot.
Stem Leaf
1 2 4 9
2 0 2 5 5 7 7 7 8
3 1 1 2 2 3 5 8 9
4 0 0 1 2 4 5 6 6 8 9
5 0 0 2 2 2 7 8
6 1 2 4 9
4|5 means 45 km/hr
2. The number of students attending a class (maximum 25) for 30 lessons is given in the table below:
(a) Are the data discrete or continuous? The data are discrete because the number of student can only
be a whole number.
(b) Enter the data into a frequency distribution table (groups not required). Include a relative
frequency and % relative frequency column.
1
21 Ι 1
30
= 0.0333 3.3%
2
22 ΙΙ 2
30
= 0.0666 6.7%
5
23 ΙΙΙΙ 5
30
= 0.166 16.7%
11
24 ΙΙΙΙ ΙΙΙΙ Ι 11
30
= 0.366 36.7%
10
25 ΙΙΙΙ ΙΙΙΙ 10 = 0.3333 33.3%
30
Total 30 1.00 100%
Page 65
(c) What proportion of lessons contained 22 students? The relative frequency for 22 students is
0.067 (to 3 decimal places)
(d) What percentage of lessons were fully attended? From the table - 33.3%
(a) Enter the data into a Stem and Leaf Plot. Use the key: 11|4 means 114.
Stem Leaf
8 8
9 2 5
10 2 7 9
11 0 2 4 5 8
12 0 1 2 2 4 6 7 8
13 1 3 4 4 7 8
14 1 3 4 5 6 9
15 5
16 1
17 5
18 8
11|4 means 114
(b) If hypertension (high blood pressure) is defined by a systolic blood pressure 140 or above, what
percentage of this group are suffering hypertension? With a Stem and Leaf, the original values
are retained, so counting is required. There are 10 out of 35 which is 28.6% (to 1 d.p.)
4. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m).
(a) Enter the data into a Frequency Distribution Table. Your FDT must have at least 5 groups.
2
22.5 to but less than 23 ΙΙ 2 25
27
× 100 =7.4%
Page 66
2
23 to but less than 23.5 ΙΙ 2 27
27
× 100 =
7.4%
Total 27 100%
(b) How many have thrown less than 22m? (Use cumulative frequency to answer this) Using the
cumulative frequency column, find the row that contains " but less than 22" - the answer is
given by the cumulative frequency of this row, 19.
(c) What percentage threw 21 to but less than 22m? (Use a % Relative Frequency Column) This is
two groups of the table. Each group has 22.2%, so 44.4% threw 21 to but less than 22m.
Graphs
1. The height of a plant is measured every Friday morning for twelve weeks. Which type of graph would
be best to show the growth of the plant? Line
2. A ‘sound and vision’ shop sells CDs, DVDs, console games and computer games. Which type of
graph would best display the relative sales of the different items sold? Pie
3. A student wishes to compare the number of motorcycle fatalities between the states. To get a more
accurate picture, data are collected for the years 2007, 2008 and 2009. Which type of graph allows
the data to be presented in one graph? Composite Column
4. The number of students attending a class (maximum 25) for 30 lessons is given in the table below:
8
6
4
2
0
20 21 22 23 24 25
Number Attending Class
Page 67
5. The Shot Put distances thrown by 27 world champion shot putters are given in the Frequency
Distribution Table below. The unit is metres (m).
(a) Construct a frequency histogram for this information. As a second step, put a frequency polygon on
the histogram.
4
Frequency
0
20 20.5 21 21.5 22 22.5 23 23.5
Distance Thrown (m)
(b) Construct a cumulative frequency histogram and then add an Ogive to the histogram.
25
Cumulative Frequency
20
15
10
0
20 20.5 21 21.5 22 22.5 23 23.5
Distance Thrown (m)
X=
∑X No value is repeated – there is no
mode.
As there are 15 values, the position
of the median is:
n The mode would not significant in n + 1 15 + 1 16
317 + 351 + .... + 885 + 926 this example. = = = 8
X= 2 2 2
15 The 8th value is the median.
X = 619.2 Counting through the ordered list,
The mean is approx. 619 hours. The the 8th value is 688. The median is
mean is a good indicator of a central also a good indicator or typical
or typical value in this question. value in this question.
2. The number of children in a 10 families is: 1, 5, 2, 2, 2, 3, 1, 4, 3, 2. Calculate the mean, mode and
median of the values. Do not organise into a table.
X=
∑X The mode is 2, it occurs 4 times, or
has a frequency of 4.
As there are 15 values, the position
of the median is:
n The mode is a good indicator of a n + 1 10 + 1 11
1 + 1 + .... + 4 + 5 typical value in this example. = = = 5.5
X= 2 2 2
10 The 5.5th value is the median.
X = 2.5 The 5th value is 2 and the 6th value is
The mean is 2.5 children. The 2 also, so the median is 2. The
mean is a good indicator of a median is also a good indicator or
central or typical value in this typical value in this question.
question. However it is not
possible to have 2.5 children
Page 69
3. For the Car Speed Data, calculate the mean, mode and median of the values after organising the data
in Frequency Distribution Table (This can be found in the answers to the Topic ‘Graphs’).
Page 70
There is also a graphical method for working out the median.
20
15
Median is approx. 42
10
5
0
10 20 30 40 50 60 70
Speed (km/hr)
Page 71
(b) Discuss the appropriateness of each measure of central tendency as a typical value.
5. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m).
Page 72
There is also a graphical method for working out the median.
25
Cumulative Frequency
20
15
10
Median is approx 21.6
0
20 20.5 21 21.5 22 22.5 23 23.5
Distance Thrown (m)
Measures of Spread
1. The lifetime, in hours, of a sample of 15 light bulbs is: 351, 429, 885, 509, 317, 753, 827, 737, 487,
726, 395, 773, 926, 688, 485.
Range SD
Range = Highest Value - Lowest Value Using a Scientific Calculator
Range = 926 - 317 s = 203 (rounded to nearest whole number)
Range = 609
IQR
The first quartile (Q1) is the The third quartile (Q3) is the
n +1 3 ( n + 1)
4 4
15 + 1 3 × 16
= =
4 4
16 48
= =
4 4
= 4th value = 12th value
Q1 is 429 Q3 is 773
The IQR is 773 – 429 = 344
2. The number of children in a 10 families is: 1, 5, 2, 2, 2, 3, 1, 4, 3, 2. Calculate the range, inter-quartile
range and standard deviation of the values. Do not organise into a table.
Page 73
1, 1, 2, 2, 2, 2, 3, 3, 4, 5
Range SD
Range = Highest Value - Lowest Value Using a Scientific Calculator
Range = 5 - 1 s = 1.27 (rounded to 2dp)
Range = 4 (assuming sample)
IQR
The first quartile (Q1) is the The third quartile (Q3) is the
n +1 3 ( n + 1)
4 4
10 + 1 3 × 11
= =
4 4
11 33
= =
4 4
= 2.75th value = 8.25th value
Q1 is 1.75 Q3 is 3.25
The IQR is 3.25 – 1.75 = 1.5
3. For the Car Speed Data, calculate the range, inter-quartile range and standard deviation of the values
after organising the data in Frequency Distribution Table. The assumption is that the original values
are lost and it is a sample.
Page 74
Range SD
Range = Highest Value - Lowest Value Using a Scientific Calculator
Range = 70 - 10 s = 14.5 (rounded to 1dp)
Range = 60 (assuming sample)
IQR
The first quartile (Q1) is the The third quartile (Q3) is the
n +1 3 ( n + 1)
4 4
40 + 1 3 × 41
= =
4 4
41 123
= =
4 4
= 10.25th value = 30.75th value
using interpolation using interpolation
10.25 − 3 30.75 − 29
The first quartile (Q1 ) = 20 + ×10 The third quartile (Q3 ) = 50 + ×10
8 7
7.25 1.75
= 20 + ×10 = 50 + ×10
8 7
= 29.1 = 52.5
The IQR is 52.5 – 29.1 = 23.4
The Ogive can also be used to calculate the quartiles and IQR
20
15 IQR = 52 – 29 = 23
10
5
0
10 20 30 40 50 60 70
Speed (km/hr)
Page 75
4. The number of students attending a class (maximum 25) for 30 lessons is given in the table below. The
assumption is that the original values are lost and it is a sample.
Range SD
Range = Highest Value - Lowest Value Using a Scientific Calculator
Range = 25 - 20 s = 1.27 (rounded to 2dp)
Range = 5 (assuming sample)
IQR
The first quartile (Q1) is the The third quartile (Q3) is the
n +1 3 ( n + 1)
4 4
30 + 1 3 × 31
= =
4 4
31 93
= =
4 4
= 7.75th value = 23.25th value
The 7th and 8th values are both 23. The 23rd and 24th values are both 25.
Q1 = 23 Q3 = 25
The IQR is 25 – 23 = 2
5. The Shot Put distances thrown by 27 world champion shot putters are given in the table below. The
unit is metres (m). The assumption is that the original values are lost and it is a sample.
Range SD
Page 76
Range = Highest Value - Lowest Value Using a Scientific Calculator
Range = 23.5 - 20 s = 0.902 (rounded to 3dp)
Range = 3.5 (assuming sample)
IQR
The first quartile (Q1) is the The third quartile (Q3) is the
n +1 3 ( n + 1)
4 4
27 + 1 3 × 28
= =
4 4
28 84
= =
4 4
= 7th value = 21st value
Interpolation is not required here as the By interpolation
7th value is 21. 21 − 19
The third quartile (Q3 ) =
22 + ×10
4
2
= 22 + × 0.5
4
= 22.25
The IQR is 22.25 – 21 = 1.25
The Ogive can also be used to calculate the quartiles and IQR
25
Cumulative Frequency
20
15
IQR = 22.2 – 21 = 1.2
10
0
20 20.5 21 21.5 22 22.5 23 23.5
Distance Thrown (m)
Page 77
Comparing Data
1. The lifetime, in hours, of a sample of 15 light bulbs is: 351, 429, 885, 509, 317, 753, 827, 737, 487,
726, 395, 773, 926, 688, 485.
From the Measures of Spread section, the quartiles are 429 and 773.
300 350 400 450 500 550 600 650 700 750 800 850 900 950
2. The Car Speed Data were collected from a school zone just prior to an advertising campaign. The data
collected were:
After an advertising campaign to make motorists more aware of speeding in school zones, the
following data were collected.
22 45 42 27 27 40 45 26 75 25
31 48 30 16 28 42 19 55 46 38
34 50 14 59 33 42 30 32 36 18
38 41 50 14 31 39 25 46 39 27
To organise the data, a back to back stem and leaf plot will be used.
After advertising Stem Prior to Advertising
9 8 6 4 4 1 2 4 9
8 7 7 7 6 5 5 2 2 0 2 5 5 7 7 7 8
9 9 8 8 6 4 3 2 1 10 0 3 1 1 2 2 3 5 8 9
8 6 6 5 5 2 2 2 1 0 4 0 0 1 2 4 5 6 6 8 9
9 50 0 5 0 2 2 2 7 8 9
6 1 2 4 9
5 7
Page 78
Lowest First Quartile is the Median is the Third Quartile is the Highest
Value = 12 n +1 n +1 3 ( n + 1) Value = 69
4 2 4
40 + 1 40 + 1 3 × 41
= = =
4 2 4
=
41
=
41 = 30.75th value
4 2 This means that the third
= 10.25th value = 20.5th value quartile is three quarters of
This means that the first Median = 40 the way between the 30th and
quartile is a quarter of the way 31th values.
between the 10th and 11th
values. Third Quartile
First Quartile = 50 + 0.75 x (52-
= 27 + 0.25 x (28- 50)
27) =51.5
= 27.25
Lowest First Quartile (Q1) = Median = Third Quartile (Q3) Highest
Value = 12 27.25 40 = 51.5 Value = 69
Page 79
12 27.25 40 51.5 69
Prior
After
14 27 35 44.25 75
10 15 20 25 30 35 40 45 50 55 60 65 70 75 80
The quartiles and the median are lower after the campaign, based on this there is some evidence
of a drop in speed. The lowest and highest value have increased, these can be outliers and so
unreliable. The Stem and Leaf Plot suggests that the highest is an ‘odd’ value.
3. The heart rate of a group of athletes was compared to the heart rate of a group of office workers after
climbing a set of stairs. The data are presented below:
Athletes
96 111 88 79 101 91
104 106 121 93 103 96
85 117 126 97 83 93
112 106 110 116 91 88
Office Workers
114 107 94 113 113 97
118 103 121 127 117 131
145 132 118 108 145 126
100 138 120
Page 80
For the Athletes
Lowest First Quartile is the Median is the Third Quartile is the Highest
Value = 79 n +1 n +1 3 ( n + 1) Value = 126
4 2 4
24 + 1 24 + 1 3 × 25
= = =
4 2 4
=
25
=
25 = 18.75th value
4 2 This means that the third
= 6.25th value = 12.5th value quartile is three quarters of
This means that the first Median = the way between the 18th and
quartile is a quarter of the way (97+101)/2=99 19th values.
between the 6th and 7th values.
First Quartile Third Quartile
= 91 = 110 + 0.75 x
(111-110)
=110.75
Lowest First Quartile (Q1) = Median = Third Quartile (Q3) Highest
Value = 79 91 99 = 110.75 Value = 126
79 91 99 110.75 126
Athletes
Office
Workers
94 107.5 114 129 145
70 75 80 85 90 95 100 105 110 115 120 125 130 135 140 145 150
Page 81
Energy Content of various Milks
1200
1000
800
Energy (kJ)
600
400
200
0
0 5 10 15 20
Percentage Level of Fat (%)
(b) The correlation coefficient = 0.989; This also suggests very strong correlation.
2. A large company is making note of the amount spent on advertising each month and then comparing
this with the amount of monthly sales.
Chart Title
180
160
140
Sales ($x1000)
120
100
80
60
40
20
0
1 2 3 4 5 6 7
Advertising Cost ($x1000)
(b) The correlation coefficient = 0.735 suggesting Quite Strong Positive Correlation
Page 82
3. (a) Distance Inland vs January Average Max Temp. The scatterplot suggests strong positive
correlation.
January Max Temp for Towns
40
38
Temperature (°C) 36
34
32
30
28
26
0 200 400 600 800 1000 1200
Inland Distance (km)
=y 0.0071x + 31.5 or
=
TJan max 0.0071d + 31.5
where TJan max - Jan Average Max Temp ( °C)
d - Inland Distance (km)
(b) Distance Inland vs July Average Minimum Temp. There is some (negative?) correlation here. It
appears that once you get 200km inland that the minimum temperature stops decreasing and then
increases at a lower rate.
July Min Temp for Towns
13
12
11
Temperature (°C)
10
9
8
7
6
5
4
0 200 400 600 800 1000 1200
Distance Inland (km)
Correlation coefficient = -0.37. The bounds for "no correlation" are between -0.4 and 0.4, as
the coefficient falls within these bounds, the regression equation is not calculated.
Page 83
(c) Distance Inland vs Annual Rainfall. Scatterplot suggests strong negative correlation. This
means the greater the distance inland, the lower the annual rainfall.
Regression line:
y= −0.505 x + 796 or
RAv = −0.505d + 796
where RAv - Average Annual Rainfall (mm)
d - Inland Distance (km)
Page 84
Comments:
Graph (a) is not surprising. Maximum temperature can be influenced by other factors such as altitude but
distance from the sea is a significant one.
Graph (b) is a very interesting graph. Perhaps altitude has an effect over 200 km inland. This could lead to
further investigations.
Graph (c) is expected. However rainfall can be affected by topography such as mountain ranges. This
usually creates a rainy side or a rain shadow depending on the direction of the prevailing wind.
There is one location (Blackwater) that is lower than the regression line suggests, perhaps this
location is in a rain shadow.
Page 85