Statistics
Definition: The study of the collection, analysis, interpretation, presentation, and
organisation of data.
Collecting Data Terminology
Population: the entire group which we are interested in but usually
can’t access directly
Sample: a representative portion of the population that we can
examine more easily
Survey: is an investigation or tool to collect information about a given sample or population.
A survey of the entire population is called a _________________.
Bias is a tendency which causes a difference in the results of a survey and the true facts.
Types of Sampling Methods
____________________________________ - Each individual in the population has an equal chance of
being chosen
e.g. each individual is assigned a number and a random number generator selects a number.
____________________________________ - Individuals are selected at regular intervals.
e.g. every 20th car is stopped on a highway, and the driver is surveyed regarding alcohol
consumption.
____________________________________ - The population is divided into subgroups and
individuals are surveyed
within that group. e.g. a population is surveyed regarding their heart health but are first
divided into age groups.
Errors may occur in sampling. Examples include:
- Sample size is too ___________.
- The sample does not represent the entire ____________________.
- Poor or _________________________ measurements are recorded.
- An inappropriate sampling method is chosen.
Types of Data
Example 1
Classify the following data as categorical or numerical:
a. brand of car
b. weekly wage
c. no. of siblings
d. petrol prices
e. size of pizzas
Now further classify the types of data above as either: discrete, continuous, ordinal or
nominal.
Example 2
a. Identify the sampling method used in each of these surveys.
i. Athletes are divided into sport type then tested for heart fitness.
ii. Every fifth competition paper ordered by surnames is selected for analysis regarding
the working shown.
iii. Every family in an apartment block is assigned a number and then five random
numbers are chosen so that families can be surveyed regarding their thoughts on
the standard of the cleaning in the block.
b. Wasram is to seek opinions about what people in the community think of the most
recent Soccer World Cup. He surveys his team-mates at the soccer club using a
questionnaire. Do you think the sample will contain bias? If so, explain why.
Statistics – Mean, Median, Mode
and Outliers
Cut and paste this in your books! Then
put the examples after!
Example 1
For the given data sets, find the following:
i. The mean
ii. The median
iii. The mode
a) 4 , 8 ,10 , 2 , 3 ,6 ,2 b) 12 ,15 , 21 , 24 , 18 ,16 ,17 , 13
Example 2
The amount of money a teenager earned in her last six babysitting appointments is:
$15, $25, $35, $18, $20, $25 .
a) Calculate the mean for this set of data.
b) How much would the teenager need to earn on the next babysitting appointment for the
mean to equal $25 ?
9K – Measures of spread, range, and interquartile range.
Five figure summary statistics
To find the 5 figure statistics you must first put your data in order from ____________ to
____________. Here is how you find each of the 5 figure statistics:
Minimum value Lower quartile Median (Q2) Upper quartile Maximum value
(Q1) (Q3)
The diagram on the
right, shows how to find
the quartiles if you have
an even or odd number
of data points.
Measures of spread Outliers
These tell you how spread your data is. There Outliers are data elements that sit outside
are two you need to know: the vicinity of the rest of the data. To
determine if there are outliers you must
Range=¿ look at the upper and lower fence:
Lower fence =
Upper Fence =
Interquartile range ( IQR )=¿
Example 1
Here is a set of measurements collected by measuring the lengths, in metres, of 10 long-jump attempts:
6.7, 9.2, 8.3, 3.0, 5.1, 7.9, 8.4, 9.0, 8.2, 8.8, 7.1
a) List the data in order from smallest to largest.
b) Find the range.
c) Find:
i) the median (Q2)
ii) the lower quartile (Q1)
iii) the upper quartile (Q3)
iv) the internal quartile range (IQR)
d) Interpret the IQR.
e) Check if there are any outliers, if so what is it?
9I Stem and Leaf Plots
A stem-and-leaf plot uses a stem number and
a leaf number to represent
_____________________ data.
The ‘key’ tells you how the plot is to be read.
_____________________________ stem-and-leaf
plots can be used to compare two sets of
data.
The stem is drawn in the middle with the
leaves on either side.
Types of data distributions
Numerical data distributions are characterised by shape
(___________________________________________) and spread.
SYMMETRICAL OR NORMALLY DISTRIBUTED
NEGATIVELY SKEWED
POSITIVELY SKEWED
Outliers
Extremely
high or low
value
Example 1 compared to
the majority
Example 2
A shop owner has two shops. The daily sales in each shop over a 16-day period are monitored
and recorded as follows.
Shop A: 3, 12, 12, 13, 14, 14, 15, 15, 21, 22, 24, 24, 24, 26, 27, 28
Shop B: 4, 6, 6, 7, 7, 8, 9, 9, 10, 12, 13, 14, 14, 16, 17, 27
a. Draw a back-to-back stem-and-leaf plot with an interval of 10.
b. Find the median and mode for both shops.
Compare and comment on differences between the sales made by the two shops
Data Displays: Numerical vs Categorical
Data
Which display to use?
Frequency Tables and Histograms
A ____________________ table shows the number of values within a set of categories or
class intervals.
Grouped numerical data can be illustrated using a ____________________.
- The height of a column corresponds to the frequency of values in that
_____________________.
- A percentage frequency histogram shows the frequencies as a
____________________ of the total. Example
Boxplots
A boxplot is a graphical representation of a _____________________________ summary.
- The median is shown by a vertical line drawn within the box
- The box is used to represent the middle 50% of scores (IQR)
- The lines (whiskers) extend out of the box to the smallest and largest data values of
the data set.
Example 1 Consider the given data set: 12, 8, 19, 13, 22, 15, 1, 17, 24, 19
a. Determine whether any outliers exist.
b. Draw a box plot to summarise the data, marking outliers if they exist.
______________________________________ are two or more box plots drawn on the same scale.
They are used to compare data sets within the same context.
Example 2
Example 3
Creating boxplots on the CAS
Example 4
7A and 7C (Scatterplots and correlation coefficient)
Bivariate Data
Data with two variables
Displayed on a Cartesian plan with data as pairs written
( x 1 , y 1 ) (x 2 , y 2)
Usually represented on a scatterplot in order to show the
relationship between two variables
When constructing a scatterplot we need to include:
o A title
o Axis markers – evenly spaces and consistent interval
o Axis labels with units
Explanatory and Response Variables
The explanatory variable is always on the x-axis
The response variable is always on the y-axis
“The explanatory variable AFFECTS or EXPLAINS the response
variable”
Example: The number of hours you spend studying for a maths
test (explanatory variable), should have an affect/explain your test
score (response variable)
Label each variable in the following scenarios:
1. Ice creams sold and the temperature of the day
2. Month of the year and amount of money spent on puffer jackets
3. The amount of time spent sleeping and the age of a person
Constructing a scatterplot
Example 2: The operators of a casino keep records of the number of people playing a ‘jackpot’ type game.
The table below shows the number of players for different prize amounts.
Number of players 260 280 285 340 390 428 490
Prize ($) 1000 1500 2000 2500 3000 3500 4000
a) Construct a scatterplot by hand ensuring you include all necessary features
Interpreting Scatterplots – Form, Direction and Strength
Correlation
Reporting on scatterplots:
There is a [strength] [direction] [form] association between the [response variable y] and the
[explanatory variable x].
For example: There is a moderately positive linear relationship between the hours studied and the
score on the maths test.
Example
1 (2007
Exam 2
Q3)
The
scatterplot
below
displays the
mean
surface
temperature
and mean duration of warm spells in Australia for 13 years
selected at random from 1960 to 2005.
Describe the association in terms of form, direction, and
strength.
Pearson’s Correlation Coefficient (r)
• Pearson’s correlation coefficient, r,
measures the strength of the linear
relationship between two variables.
• For this value to be accurate the
scatterplot must be linear, numeric and
have no obvious outliers
Reading the chart:
• If I calculated ‘r’ to be 0.61 then the
scatterplot would be showing a moderately
positive linear relationship between the
two variables
• If ‘r’ was -0.12 it would have no linear
association
Example 3:
INTERPRETING PEARSONS
CORRELATION COEFFICIENT (r)
There exists a strong positive linear
correlation between the chocolate
consumption (kg/yr/capita) and the
Nobel Laureates (per 10 million
people) as shown by the r value of
0.7910.
There exists a ____________________
(e.g. strong negative linear
correlation) between the
_____________________(explanatory
and response variable) as shown by
the r value of __________ (to 2dp)
7D and 7E Fitting and interpreting linear models to bivariate data.
Bivariate Data
When bivariate data have a strong linear correlation, we can model the data with a
straight line. This is called a line of best fit.
A line of best fit is positioned by eye by balancing the number of points above the
line with the number of points below the line.
We can use the line of best fit to see the relationship between our two vairables, or to
predict what might happen next!
Example 1: For the following scatterplot, answer the following
questions:
🧮 Finding the equation of a line of
best fit
1. Draw your scatter plot
Plot all your data points on a
a) Determine the equation of the line shown on the scatterplot
2. Sketch the line of best fit
Sketch a line that is equal
distance from the most points,
some above and some below.
3. Pick two clear points on
your line
These should be two points on
your line, not your data
points!
Example: (x₁, y₁) and (x₂, y₂)
4. Find the slope (b)
Use the formula:
y 2− y 1
b=
b) Which is the EV and which is the RV?
c) The slope tells us on average, that construction cost increases by ________ ($’000)
for each ______ m2 in house size.