Understanding Scatter Plots and Correlation
Understanding Scatter Plots and Correlation
Quantitative
(measurements/counts)
Types of Variables
Qualitative
(groups)
28
ge 26
A
24
22
20
1930 1940 1950 1960 1970 1980 1990
Year
Trend
Do you see
Do you see
Scatter
Do you see
Do you see
Anything unusual
Do you see
any outliers?
Do you see
any groupings?
r=0 r = 0.9
160
150
22 23 24 25 26 27 28 29
Foot size (cm)
What will happen to the correlation coefficient if the tallest Year 10 student is
removed? Tick your answer:
(Remember the correlation coefficient answers the question: “For a linear relationship, how
well do the data fall on a straight line?”)
s)
ar Life Expectancies and Gestation Period
e for a sample of non-human Mammals
(Y 40
y
nc
ta 30 Elephant
ec
p
20
x
E
e
Lif
10
50
0 10000 20000 30000 40000
People per Doctor
80
y
nc
ta
ec
p 70
x
E
e
Lif 60
50
0 100 200 300 400 500 600
People per Television
make another suggestion as to what needs to
be done in a country to increase life
expectancy?
Two variables may be strongly associated (as measured by the correlation coefficient for linear
associations) but may not have a cause and effect relationship existing between them. The explanation
maybe that both the variables are related to a third variable not being measured – a “lurking” or
“confounding” variable.
Only talk about causation if you have well designed and carefully carried out experiments.
Data Sources
[Link]
[Link]
[Link]
[Link]
[Link]
Can knowing the percentage total fat content of a cracker help us to predict the energy
content?
If I switch to a different brand of cracker with 100mg per 100g less salt content, what
change in percentage total fat content can I expect?
The energy content of 100g of cracker for 18 common cracker brands are shown in the dot plot
with summary statistics below.
Based on the information above, my prediction for the energy content of a cracker is
____________ Calories per 100g.
Another quantitative variable which could be useful in predicting (the explanatory variable)
the energy content (the response variable) of 100g of cracker is ____________________.
The Consumer magazine gives some nutritional information from an analysis of these 18 brands of
cracker. Some of this information is shown in the table below:
500
(calories/100g)
Energy
450
400
0 10 20 30
Total Fat (%)
500
(Calories per 100g)
Energy
450
400
500
(calories/100g)
Energy
450
400
0 10 20 30 40 50 60
Number of Crackers/100g
From these plots, the best explanatory variable to use to predict energy content is
________________________________________________________ because
_________________________________________________________________________
Roughly, my line predicts the energy content for a cracker with a 10% total fat content is about
550
Which Line?
500
(calories/100g)
Energy
450
400
0 10 20 30
Total Fat (%)
data point
(8, 25)
Observed y-value 25 y
Fitted line
Minimise
There is one and only one least squares regression line for every linear regression
for the least squares line but it is also true for many other lines
530
quantitative variables. Check for a
510
430
410
390
370
350
0 5 10 15 20 25 30
If the assumptions (straightness of line) The data suggests a linear trend. The association is
appear to be satisfied then fit a linear positive and very strong. The data suggests constant
regression. scatter about the trend line. It is sensible to do a linear
regression.
450
430
410
390
370
350
0 5 10 15 20 25 30
Name the variables, the units of We have two quantitative variables, total fat content (%)
measure, and who/what is measured and salt content (mg per 100g) measured on 18 common
(units of interest). Specify the cracker brands. We are investigating the relationship
question/problem of interest. between these two variables for the purpose of describing
how the total fat content changes as the salt content
changes.
The scatter plot is the basic tool for
investigating the relationship between
2 quantitative variables. Check for a
Common Cracker Brands
35
linear trend – never do a linear
30
regression without first looking at the Total Fat (%) 25
scatter plot 20
15
10
5
0
0 500 1000
Salt (mg per 100g)
25
20
15
10
5
0
0 500 1000
Salt (mg per 100g)
Use the equation to answer the Under this regression, in a 100g of cracker, a decrease of
original question. about 2.4% of total fat content is associated, on average,
with each 100mg decrease in salt content.
Calorie, fat, carbohydrate, protein content for various foods including fast foods by chain:
[Link]
Four scatter plots with fitted lines are shown below. The equation of the fitted line and the
value of R2 are given for each plot.
Energy (calories/100g)
500 500
450 450
400 400
350 350
100 300 500 700 900 1100 1300 0 10 20 30 40
Salt (mg/100g) Total Fat (%)
30
500
Total Fat (%)
25
20
450
15
400 10
5
350 0
0 20 40 60 0 500 1000 1500
Number of crackers per 100g Salt (mg/100g)
Comment on any relationship between the scatter plot and the value of R2.
What do you think R2 is measuring?
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
y
35
6.5 29.0
30
6.9 30.6
25
7.8 34.2 20
4 5 6 7 8 9 10 11 12
8.1 35.4
x
8.4 36.6
9.1 39.4
10.3 44.2
12.0 51.0
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
None
The variability in the fitted values is exactly the same as the variability in the observed values.
The fitted line explains ______________ of the variability in the observed values.
y
4
6.5 4.3
3
6.9 7.9 2
1
7.8 4.8 0
4 6 8 10 12 14
8.1 5.8
x
8.4 2.2
9.1 1.4
10.3 6.8
12.0 6.0
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
No variability
The variability in the residuals is exactly the same as the variability in the observed values.
The fitted line explains ______________ of the variability in the observed values.
45
y = 3.4883x - 1.8364
40
35
30
y
25
20
15
10
4 6 8 10 12 14
x
R2 = 0.866
R2 gives the fraction of the variability of the y-values accounted for by the linear regression
(considering the variability in the x-values).
R2 is often expressed as a percentage.
If the assumptions (straightness of line) appear to be satisfied then R 2 gives an overall
measure of how successful the regression is in linearly relating y to x.
R2 lies from 0 to 1 (0% to 100%).
The smaller the scatter about the regression line the larger the value of R2.
Therefore the larger the value of R2 the greater the faith we have in any estimates using the
equation of the regression line.
R2 is the square of the sample correlation coefficient, r.
For the above example, the linear regression accounts for 86.6% of the variability in the y-
values from the variability in the x-values.
Exercise: List the plots from greatest R2 to least R2.
50 0.1
0.09
40 0.08
Optical absorbance
0.07
Hardness (units)
30 0.06
0.05
20 0.04
0.03
10 0.02
0.01
0 0
150 200 250 300 350 400 0 0.5 1 1.5 2 2.5 3
Cem ent (gram s) Dissolved organic carbon (m g/L)
12 80
11.5
Pavement condition index
70
Time (seconds)
11
60
10.5
50
10
9.5 40
0.1 0.12 0.14 0.16 0.18 0.2 0.22 0.24 10 11 12 13 14 15 16 17 18 19
500
_____________________________________
450
400 _____________________________________
350
100 300 500 700 900 1100 1300
Salt (mg/100g)
Common Cracker Brands
y = 4.9844x + 380.82 _____________________________________
550 R2 = 0.982
Energy (calories/100g)
500
_____________________________________
450 _____________________________________
400
_____________________________________
350
0 10 20 30 40
Total Fat (%)
Common Cracker Brands
y = 0.3717x + 440.06 _____________________________________
550 R2 = 0.0166
Energy (calories/100g)
_____________________________________
500
450 _____________________________________
400
_____________________________________
350
0 20 40 60
Number of crackers per 100g
Common Cracker Brands
y = 0.0237x - 2.6556
_____________________________________
35 R2 = 0.4892
30
_____________________________________
Total Fat (%)
25
20
15
_____________________________________
10
5
0 _____________________________________
0 500 1000 1500
Salt (mg/100g)
The following table shows the winning distances in the men’s long jump in the Olympic Games
for years after the Second World War.
9
Distance (metres)
8.5
7.5
7
0 10 20 30 40 50 60
Years since 1944
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
8.5
7.5
7
0 10 20 30 40 50 60
Years since 1944
We estimate that for every 4-year increase in years (from one Olympic Games to the next) the
Using this linear regression we predict that the winning distance in 2004 will be ___________.
Excel output for a linear regression on 14 observations (with the 1968 observation removed)
8.5
7.5
7
0 10 20 30 40 50 60
Years since 1944
We estimate that for every 4-year increase in years (from one Olympic Games to the next) the
Using this linear regression we predict that the winning distance in 2004 will be ___________.
fitted line?
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
predicted winning distance in 2004?
____________________________________________________________________________
value of R2?
____________________________________________________________________________
An outlier, in a regression context, is a point that is unusually far from the trend.
If it is an actual unusual observation then try to understand why it is so different from the other
observations.
Observatio 1 2 3 4 5 6 7 8 9 10
n
1st reading 116 122 136 132 128 124 110 110 128 126
2nd reading 114 120 134 126 128 118 112 102 126 124
Observatio 11 12 13 14 15 16 17 18 19 20
n
1st reading 130 122 134 132 136 142 134 140 134 160
2nd reading 128 124 122 130 126 130 128 136 134 160
Blood Pressure
160
150
Second reading
140
130
120
110
100
100 110 120 130 140 150 160
First reading
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
Blood Pressure
y = 0.9337x + 4.9068
160 R2 = 0.8673
150
Second reading
140
130
120
110
100
100 110 120 130 140 150 160
First reading
We estimate that for every 10-unit increase in the first blood pressure reading the second
For a person with a first reading of 140 units we predict that the second reading will be
________________
150
Second reading
140
130
120
110
100
100 110 120 130 140 150 160
First reading
We estimate that for every 10-unit increase in the first blood pressure reading the second
For a person with a first reading of 140 units we predict that the second reading will be
________________
fitted line?
___________________________________________________________________________________
value of R2?
___________________________________________________________________________________
An x-outlier can alter the position of the fitted line substantially, i.e. it can influence the position
of the fitted line.
The fitted line may say more about the x-outlier than about the overall relationship between the
two variables.
If a data set has an x-outlier then carry out two linear regressions; one with the
x-outlier included and one with the x-outlier excluded. Investigate the amount of influence the
x-outlier has on the fitted line and discuss the differences.
30
25
Petal width (mm)
20
15
10
0
0 10 20 30 40 50 60 70 80
Petal length (mm)
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
20
15
10
0
0 10 20 30 40 50 60 70 80
Petal length (mm)
Fisher's Iris Data (Iris setosa) Fisher's Iris Data (Iris versicolor)
y = 0.2012x - 0.4822 y = 0.2273x + 3.437
7 20
R2 = 0.11 R2 = 0.3797
6
18
Petal width (mm)
5
Petal width (mm)
16
4
14
3
2 12
1 10
0 8
8 10 12 14 16 18 20 25 30 35 40 45 50 55 60
Petal length (mm) Petal length (mm)
20
15
10
40 45 50 55 60 65 70 75
Petal length (mm)
Comment.
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
If there are groupings in your data that behave differently then consider fitting a different linear
regression line for each grouping.
An x-outlier or data that has groupings can make the value of R 2 seem large when the linear
regression is just not appropriate.
On the other hand, a low value of R 2 may be caused by the presence of a single outlier and all
other points have a reasonably strong linear relationship.
1200
Concentration (units/litre)
1000
800
600
400
200
0
0 2 4 6 8 10 12 14 16 18
Time (hours)
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
Suppose that a patient had a heart attack 17 hours ago. Predict the creatine kinase
concentration in the blood for this patient.
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
1200
Concentration (units/litre)
1000
800
600
400
200
0
0 10 20 30 40 50 60
Time (hours)
A fitted line will often do a good job of summarising a relationship for the range of the observed
x-values.
Predicting y-values for x-values that lie beyond the observed x-values is dangerous. The linear
relationship may not be valid for those x-values.
The removal of an x-outlier will mean that the range of observed x-values is reduced. This
should be discussed in the comparison between the two linear regressions
(x-outlier included and x-outlier excluded).
150
145
Time (minutes)
140
135
130
125
120
0 10 20 30 40 50 60
Years since 1 Jan 1946
Concerns:
____________________________________________________________________________
____________________________________________________________________________
Possible solutions:
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
145 145
Time (minutes)
Time (minutes)
140 140
135 135
130 130
125 125
120
120
0 10 20 30 40 50 60 0 10 20 30 40 50 60
Years since 1 Jan 1946 Years since 1 Jan 1946
145
Time (minutes)
140
135
130
125
120
0 10 20 30 40 50 60
Time (minutes)
140 128
138
136 127
134
126
132
130 125
128
126 124
0 5 10 15 20 25 20 30 40 50 60
Years since 1 Jan 1946 Years since 1 Jan 1946
Comments:
____________________________________________________________________________
____________________________________________________________________________
2000
1500
Weight (kg)
1000
500
1000 2000 3000 4000 5000 6000
Engine size (cc)
Concerns:
____________________________________________________________________________
____________________________________________________________________________
Possible solutions:
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________
____________________________________________________________________________