Advanced Quantitative Methods in Geography
Advanced Quantitative Methods in Geography
1
Introduction
clusively, to statistics
• Descriptive and inferential (univariate and multi-
variate to a lesser extent) statistical methods, BUT
• Background and theory that you need to make
use of statistics properly in the context of geogra-
phy
3
Goals
• The goals of this course are:
– To help students understand the purpose,
meaning, and use of statistics in geographic
research.
– To introduce basic statistical methods used by
geographers in research.
– To learn how to use computer spreadsheets to
simplify geographic problem solving.
– To learn how to apply statistical techniques
to real geographical problems.
Terminology
• Population
– A collection of items of interest in research
– A complete set of things
– A group that you wish to generalize your re-
search to
– An example – All the trees in a Park
• Sample
– A subset of a population
– The size smaller than the size of a population
– An example – 100 trees randomly selected
from a Park
Terminology
• Representative – An accurate reflection of the
population (A primary problem in statistics)
• Variables – The properties of a population that
are to be measured (i.e., how do parts of the popu-
lation differ?
• Constant – Something that does not vary
• Parameter – A constant measure which describes
the characteristics of a population
• Statistic – The corresponding measure for a
sample
Two Sorts of Statistics
• Descriptive statistics
– To describe and summarize the characteristics
of the sample
– Fall within the class of exploratory techniques
• Inferential statistics
– To infer something about the population from
the sample
– Lie within the class of confirmatory methods
(i) Descriptive Statistics
• Measures of dispersion
– Describe how the observations are distributed
– Variance, standard deviation, range, etc
Measures of Central Tendency – Mean
x
i 1
i
x1 x2 xn
Summation Notation: Components
i n
x
i 1
i
indicates what we
are summing up
indicates we are
taking a sum
i n n
x x x
i 1
i
i 1
i i
Summation Notation: Examples
x
i 1
i x1 x2 x3 x4 x5 x6 x7 x8 x9 x10
1 2 3 4 5 6 7 8 9 10
Example II: Only observations 3 through 5 are included
in the sum: 5
x
i 3
i x3 x4 x5 3 4 5 12
Summation Notation: Rules
a a a a na
i 1
4
i 1
4 4 4 4 4 4 5 20
Measures of Central Tendency – Mean
n N
xi x i
x i 1 i 1
n N
Measures of Central Tendency – Mean
• Example I
– Data: 8, 4, 2, 6, 10
5
x i
(8 4 2 6 10)
x i 1
6
5 5
• Example II
– Sample: 10 trees randomly selected from Battle Park
– Diameter (inches):
9.8, 10.2, 10.1, 14.5, 17.5, 13.9, 20.0, 15.5, 7.8, 24.5
10
x i
(9.8 10.2 24.5)
x
i 1
14.38
10 10
Measures of Central Tendency – Mean
• Example III
Annual mean temperature (°F)
x 59.70
Mean
1198.10 (mm)
Mean
58.51 (°F)
Chapel Hill, NC
(1972-2001)
Measures of Central Tendency – Mean
• Advantage
– Sensitive to any change in the value of any observation
• Disadvantage
– Very sensitive to outliers
xi y i
x i 1 y i 1
n n
Map Coordinates
• Geographic coordinates – The geographic coordinate
system is a system used to locate points on the surface
of the globe (degrees of latitude and longitude)
Geographic coordi-
nates of Chapel Hill,
NC
Source: [Link]
Weighted Mean
• We can also calculate a weighted mean using some
weighting factor:
e.g. What is the average income of all
n people in cities A, B, and C:
wi xi City
A
Avg. Income
$23,000 100,000
Population
x i 1
n
B $20,000 50,000
w
C $25,000 150,000
i
i 1 Here, population is the weighting fac-
tor and the average income is the variable
of interest
Weighted Mean Center
• We can also calculate a weighted mean center in
much the same way, by using weights:
n n
2, 4, 6, 8, 10 median: 6
• Example II
– Sample: 10 trees randomly selected from a Park
– Diameter (inches):
9.8, 10.2, 10.1, 14.5, 17.5, 13.9, 20.0, 15.5, 7.8, 24.5
(mean: 14.38)
7.8, 9.8, 10.1, 10.2, 13.9, 14.5, 15.5, 17.5, 20.0, 24.5
Mean = 8.10 m
Source: [Link]
median: (6.0 + 7.1) = 6.55
# Tree Height # Tree Height
(m) (m)
1 4.5 6 7.1 mode: 7.5
2 4.8 7 7.5
3 5.0 8 7.5
4 5.3 9 8.0
5 6.0 10 25.4
30 40 25 50 45
50 55 45 48 61
60 75 70 45 72
24 45 200 205 65
65 39 58 45 65
24, 25, 30, 39, 40, 45, 45, 45, 45, 45, 48, 50, 50,
55, 58, 60, 61, 65, 65, 65, 70, 72, 75, 200, 205
Source: [Link]
Which one is better: mean, median,
or mode?
• It also depends on your goals
• Consider a company that has nine employees with
salaries of 35,000 a year, and their supervisor
makes 150,000 a year
• What if you are a recruiting officer for the com-
pany that wants to make a good impression on a
prospective employee?
• The mean is (35,000*9 + 150,000)/10 = 46,500 I
would probably say: "The average salary in our
company is 46,500" using mean
Source: [Link]
Which one is better: mean, median,
or mode?
• The mean is valid only for interval data or ratio
data
• The median can be determined for ordinal data as
well as interval and ratio data
• The mode can be used with nominal, ordinal, in-
terval, and ratio data
• Mode is the only measure of central tendency that
can be used with nominal data
• It also depends on the nature of the distribution
and your goals
Some Characteristics of Data
• Not all data is the same. There are some limitations
as to what can and cannot be done with a data set,
depending on the characteristics of the data
• Some key characteristics that must be considered
are:
• A. Continuous vs. Discrete
• B. Grouped vs. Individual
• C. Scale of Measurement
A. Continuous vs. Discrete Data
• Continuous data can include any value (i.e., real
numbers)
– e.g., 1, 1.43, and 3.1415926 are all acceptable values.
– Geographic examples: distance, tree height, amount of
precipitation, etc
• Discrete data only consists of discrete values,
and the numbers in between those values are not
defined (i.e., whole or integer numbers)
– e.g., 1, 2, 3.
– Geographic examples: # of vegetation types,
B. Grouped vs. Individual Data
• The distinction between individual and
grouped data is somewhat self-explanatory, but
the issue pertains to the effects of grouping data
• While a family income value is collected for each
household (individual data), for the purpose of
analysis it is transformed into a set of classes
(e.g., <$10K, $10K-20K, > $20K)
• e.g., elevation (1000m vs. < 500m, 500-1000m,
1000-2000m, > 2000m)
B. Grouped vs. Individual Data
Source: Earickson, RJ, and Harlin, JM. 1994. Geographic Measurement and Quantitative Analysis.
USA: Macmillan College Publishing Co., p. 91.
Source: [Link]
Source: [Link]
We Need Both Measures!
Source: [Link]
MODIS 16-day NDVI composite (Terra, 08/12/2004-08/27/2004)
(Original data obtained from Global Land Cover Facility, [Link]
Landsat ETM+ image at Chapel Hill, NC (2002-05-24)
(7-4-1)
Monthly mean temperature (°F) at Chapel Hill, NC (2001).
Measures of Dispersion
• In addition to measures of central tendency, we
can also summarize data by characterizing its
variability
• Measures of dispersion are concerned with the
distribution of values around the mean in data:
– Range
– Interquartile range
– Variance
– Standard deviation
– z-scores
Measures of Dispersion - Range
• Range – this is the most simply formulated of all
measures of dispersion
• Given a set of measurements x1, x2, x3, … ,xn-1, xn ,
the range is defined as the difference between the
largest and smallest values:
Range = xmax – xmin
• This is another descriptive measure that is vulner-
able to the influence of outliers in a data set,
which result in a range that is not really descriptive
of most of the data
Measures of Dispersion - Range
Mean = 8.10 m
Source: [Link]
50 55 45 48 61
60 75 70 45 72
24 45 200 205 65
65 39 58 45 65
24, 25, 30, 39, 40, 45, 45, 45, 45, 45, 48, 50, 50,
55, 58, 60, 61, 65, 65, 65, 70, 72, 75, 200, 205
5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47
Mean deviation =
S |
i=1
x i – x |
n
Measures of Dispersion –
Mean Deviation
• Mean deviation -- Why is it necessary to take absolute
values of statistical distances before summing them to
get the mean deviation?
• Because the statistical distances would be both positive
and negative, and when summed using the mean devia-
tion formula, they would sum to zero
1
( xi x) xi x xi n x xi n n xi 0
Measures of Dispersion –
Variance, Standard Deviation
• As an alternative to taking the absolute values
of the statistical distances, we can square
each deviation before taking their sum, which
yields the sum of squares
n
sum of squares =
S (x
i=1
i – x) 2
2
( xi x ) 2
(x x)
i
2
s i 1
i 1
2
n 1 N
Sample variance Population variance
• Note the differences in the two formulae, both in
the notation but also the denominators!
Measures of Dispersion – Variance
• To ensure that the sample variance gives an un-
biased estimate of the true, unknown variance of
the population from which the sample was drawn,
sample variance is computed by taking the sum
of the squared deviations, and then dividing by n –
1, instead of by n
• Unbiased implies that if we were to repeat this
sampling many times, we would find that the aver-
age or mean of our many sample variances would
be equal to the true variance
Chapel Hill Bend
Mean 1198.10 298.07
Variance 36786.34 6737.095
(x x)i
2
s i 1
n 1
• Compared with variance, standard deviation has a
scale closer to that used for the mean and the original
data
Chapel Hill Bend
Mean 1198.10 298.07
Variance 36786.34 6737.095
Standard Deviation 191.80 82.08
Mean = 59.70
Standard deviation = 12.97 Monthly temperature at Chapel Hill, NC (2001)
Other measures of dispersion
(x x) 2 n
2
i
(x x)
i
2
s i 1
s i 1
n 1 n 1
s
CV 100%
x
Measures of Skewness and Kurtosis
• A fundamental task in many statistical analyses is
to characterize the location and variability of a
data set (Measures of central tendency vs. mea-
sures of dispersion)
• Both measures tell us nothing about the shape of
the distribution
• A further characterization of the data includes
skewness and kurtosis
• The histogram is an effective graphical technique
for showing both the skewness and kurtosis of a
data set
Histograms
Fig. 3. Histogram of crown width (m) measured in situ for a random sample
of Quercus robur trees in Frame Wood (n = 63; mean = 9.3 m; SD = 4.64 m).
Source: Koukoulas & Blackburn, 2005. Journal of Vegetation Science: Vol. 16, No. 5, pp. 587–596
Frequency & Distribution
• A histogram is one way to depict a frequency
distribution
• Frequency is the number of times a variable takes
on a particular value
• Note that any variable has a frequency distribution
• e.g. roll a pair of dice several times and record the
resulting values (constrained to being between and
2 and 12), counting the number of times any given
value occurs (the frequency of that value occur-
ring), and take these all together to form a fre-
quency distribution
Frequency & Distribution
• Frequencies can be absolute (when the frequency
provided is the actual count of the occurrences) or
relative (when they are normalized by dividing the
absolute frequency by the total number of observa-
tions [0, 1])
• Relative frequencies are particularly useful if you
want to compare distributions drawn from two dif-
ferent sources (i.e. while the numbers of observa-
tions of each source may be different)
Histograms
• We may summarize our data by constructing his-
tograms, which are vertical bar graphs
• A histogram is used to graphically summarize
the distribution of a data set
• A histogram divides the range of values in a data
set into intervals
• Over each interval is placed a bar whose height
represents the frequency of data values in the in-
terval.
Building a Histogram
• To construct a histogram, the data are first
grouped into categories
• The histogram contains one vertical bar for each
category
• The height of the bar represents the number of
observations in the category (i.e., frequency)
• It is common to note the midpoint of the category
on the horizontal axis
Building a Histogram – Example
• 1. Develop an ungrouped frequency table
– That is, we build a table that counts the number of oc-
currences of each variable value from lowest to high-
est:
TMI Value Ungrouped Freq.
4.16 2
4.17 4
4.18 0
… …
13.71 1
• We could attempt to construct a bar chart from this table,
but it would have too many bars to really be useful
Building a Histogram – Example
• 2. Construct a grouped frequency table
– Select an appropriate number of classes
48
Percent of cells in catchment
44
40
36
32
28
24
20
16 A proxy for
12
8 Soil Moisture
4
0
4 5 6 7 8 9 10 11 12 13 14 15 16
Source: Earickson, RJ, and Harlin, JM. 1994. Geographic Measurement and Quantitative Analysis. USA: Macmil-
lan College Publishing Co., p. 91.
Further Moments of the Distribution
(x x)
i
3
skewness i 1
3
ns
• If skewness equals zero, the histogram is symmetric
about the mean
• Positive skewness vs negative skewness
Further Moments – Skewness
Source: [Link]
Further Moments – Skewness
• Positive skewness
– There are more observations below the mean
than above it
– When the mean is greater than the median
• Negative skewness
– There are a small number of low observations
and a large number of high ones
– When the median is greater than the mean
Further Moments – Kurtosis
• Kurtosis measures how peaked the histogram is
n
(x x)
i
4
kurtosis i
4
3
ns
• The kurtosis of a normal distribution is 0
• Kurtosis characterizes the relative peakedness or flat-
ness of a distribution compared to the normal distribu-
tion
Further Moments – Kurtosis
• Platykurtic– When the kurtosis < 0, the frequen-
cies throughout the curve are closer to be equal
(i.e., the curve is more flat and wide)
• Thus, negative kurtosis indicates a relatively flat
distribution
• Leptokurtic– When the kurtosis > 0, there are
high frequencies in only a small part of the curve
(i.e, the curve is more peaked)
• Thus, positive kurtosis indicates a relatively
peaked distribution
Further Moments – Kurtosis
platykurtic leptokurtic
Source: [Link]
• Histograms
• Box plots
Functions of a Histogram
• The function of a histogram is to graphically
summarize the distribution of a data set
• The histogram graphically shows the following:
1. Center (i.e., the location) of the data
2. Spread (i.e., the scale) of the data
3. Skewness of the data
4. Kurtosis of the data
4. Presence of outliers
5. Presence of multiple modes in the data.
Functions of a Histogram
Source: [Link]
Further Moments – Skewness
• Skewness measures the degree of asymmetry exhibited
by the data n
(x x)
i
3
skewness i 1
3
ns
• If skewness equals zero, the histogram is symmetric
about the mean
• Positive skewness vs negative skewness
• Skewness measured in this way is sometimes referred to
as “Fisher’s skewness”
Further Moments – Skewness
Source: [Link]
n
Mode
Median
i
( x x ) 3
skewness i 1
Mean
ns 3
A B
Median
Mean
i
( x x ) 3
skewness i 1 3
ns
skewness i 1 3
Mean
ns
i
( x x ) 3
Mode
skewness i 1
ns 3 Median
Mean
A B
(x x) i
3
skewness i 1
ns 3
Source: [Link]
(x x)
i
4
kurtosis i
4
3
ns
• The kurtosis of a normal distribution is 0
• Kurtosis characterizes the relative peakedness or flat-
ness of a distribution compared to the normal distribu-
tion
Kurtosis – Review
• Platykurtic– When the kurtosis < 0, the frequen-
cies throughout the curve are closer to be equal
(i.e., the curve is more flat and wide)
• Thus, negative kurtosis indicates a relatively flat
distribution
• Leptokurtic– When the kurtosis > 0, there are
high frequencies in only a small part of the curve
(i.e, the curve is more peaked)
• Thus, positive kurtosis indicates a relatively
peaked distribution
n
(x x)
i
4
kurtosis i
4
3
ns
Source: [Link]
Source: [Link] (First three)
[Link] (Last)
Box Plots
• We can also use a box plot to graphically sum-
marize a data set
• A box plot represents a graphical summary of
what is sometimes called a “five-number sum-
mary” of the distribution
– Minimum
– Maximum
– 25th percentile
max. 75th
– 75th percentile %-ile
– Median median
25th
• Interquartile Range (IQR) min. %-ile
Rogerson, p. 8.
Box Plots
• Example – Consider first 9 Commodore prices ( in
$,000)
6.0, 6.7, 3.8, 7.0, 5.8, 9.975, 10.5, 5.99, 20.0
• Arrange these in order of magnitude
3.8, 5.8, 5.99, 6.0, 6.7, 7.0, 9.975, 10.5, 20.0
• The median is Q2 = 6.7 (there are 4 values on ei-
ther side)
• Q1 = 5.9 (median of the 4 smallest values)
• Q3 = 10.2 (median of the 4 largest values)
• IQR = Q3 – Q1 = 10.2 - 5.9 = 4.3
• Example (ranked)
3.8, 5.8, 5.99, 6.0, 6.7, 7.0, 9.975, 10.5, 20.0
• The median is Q1 = 6.7
• Q1 = 5.9 Q3 = 10.2 IQR = Q3 – Q1 = 10.2 - 5.9 =
4.3
Box Plots
5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47
5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47
25th percentile: 11.75 75th percentile: 26
Interquartile range: 26 – 11.75 = 14.25
Other Descriptive Summary Measures
Point
Sx
i=1
i
x= Estimator
Estimate n
• In this case, the sample mean is being used to estimate
m, the population mean
Estimation
• It is quite unlikely that our statistic will be exactly the
same as the population parameter (because we know
that sampling error does occur), but ideally it should be
pretty close to ‘right’, perhaps within some specified
range of the parameter
• We can define this in terms of our statistic falling
within some interval of values around the parameter
value (as determined by our sampling distribution)
• But how close is close enough?
Estimation and Confidence
• When n increases:
1. The distributions becomes more and more normal
2. The spread of the distributions decreases
Source: [Link]
Central Limit Theorem – Example
95%
68%
f(x)
2 2
s s
N N
pr 1.96 x 1.96 0.95
N N
Confidence Intervals for the Mean
• An individual sample mean should, 95% of the time,
lie within of(the
1.96 / ntrue
) mean, μ
pr 1.96 x 1.96 0.95
n n
• Rearrange the expression:
pr x 1.96 x 1.96 0.95
n n
• This tells us that 95% of the time the true mean
should lie within 1.96( / n ) of the sample mean
Confidence Intervals for the Mean
x 89.96, s 51.68
pr x 1.96 x 1.96 0.95
n n
Source: [Link]
Example
pr 89.96 1.96 89.96 1.96 0.95
75 75
51.68 51.68
pr 89.96 1.96 89.96 1.96 0.95
75 75
Example
51.68 51.68
pr 89.96 1.96 89.96 1.96 0.95
75 75
(78.26, 101.66)
s (0.55)(0.45)
sX = (p)(1-p) = = 0.0157
n =
n 1000
14.43 14.43
pr 21.93 2.045 21.93 2.045 0.95
30 30
Confidence intervals & t-distribution
14.43 14.43
pr 21.93 2.045 21.93 2.045 0.95
30 30
• We are 95% sure that the true mean is within the in-
terval (16.54, 27.32)
• More precisely, 95% of confidence intervals con-
structed from samples in this way will contain the true
mean
Hypothesis Testing
• One-sample tests
– One-sample tests for the mean
• Two-sample tests
– Two-sample tests for the mean
179
Hypothesis Testing
• Hypothesis testing
Null hypothesis
• Purpose
Test the viability
• Null hypothesis
Population parameter
Reverse of what the experimenter believes
180
Hypothesis Testing: Steps
1. State the null hypothesis, H0
– H0: μ = μ0
– H0: μ - μ0 = 0
182
Hypothesis Testing – Step 2
or
– HA : μ > μ0 upper-tailed
one-sided (one-tailed)
– HA : μ < μ0 lower-tailed
183
Hypothesis Testing – Step 3
– α = 0.05
– α = 0.01
– Type I error
184
Hypothesis Testing - Errors
185
Hypothesis Testing - Errors
H0 is true H0 is false
186
Hypothesis Testing – Step 4
4. Select a statistical test, and find the test statistic
q - q0
Test statistic = Std. error
x x
z z
/ n s/ n
187
Hypothesis Testing – Step 4
4. Select a statistical test, and find the test statis-
tic
q - q0
Test statistic = Std. error
x x
t t
/ n s/ n
188
Hypothesis Testing – Step 5
189
Two-sided tests Zα/2
190
One-sided tests Zα
191
Hypothesis Testing – Step 6
6. Compare the observed test statistic with
the critical value
192
Hypothesis Testing – Step 6
193
Hypothesis Testing – Step 6
194
Hypothesis Testing – Step 6
HA
195
p-value
• p-value is the probability of getting a value of the test
statistic as extreme as or more extreme than that observed
by chance alone, if the null hypothesis H0, is true.
196
p-value
197
One-Sample z-Tests
198
Next topics:
• Analysis of variance (ANOVA)
• Correlation
• Regression
199