0% found this document useful (0 votes)
13 views199 pages

Advanced Quantitative Methods in Geography

The document outlines a course on Advanced Quantitative Methods and Software in Geography and Environmental Studies, taught by Dr. Zbelo Tesfamariam at Mekelle University for the 2024/2025 academic year. It covers the purpose and use of statistics in geographic research, including descriptive and inferential statistics, and aims to equip students with the skills to apply statistical techniques to real-world geographical problems. Key concepts such as population, sample, central tendency, and measures of dispersion are discussed, along with practical applications of statistical methods in geography.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views199 pages

Advanced Quantitative Methods in Geography

The document outlines a course on Advanced Quantitative Methods and Software in Geography and Environmental Studies, taught by Dr. Zbelo Tesfamariam at Mekelle University for the 2024/2025 academic year. It covers the purpose and use of statistics in geographic research, including descriptive and inferential statistics, and aims to equip students with the skills to apply statistical techniques to real-world geographical problems. Key concepts such as population, sample, central tendency, and measures of dispersion are discussed, along with practical applications of statistical methods in geography.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Advanced Quantitative Methods &

Software in Geography & Environ-


mental Studies

Lecturer: Zbelo Tesfamariam (PhD)


Associate Professor, Department of Geography & Environmen-
tal Studies, Mekelle University
Programme: Summer in Service Post Graduate
Academic year: 2024/2025
Lecture room: GeES- GIS lab

1
Introduction

• Geography is a diverse discipline, that seeks to


understand our world in terms of space and place
• Geographers use quantitative approaches to
describe, understand, and assess geographic
phenomena
• This course will provide you with an advanced to
quantitative approaches in geography & envi-
ronmental studies
2
What is quantitative methods?

• Quantitative methods refers primarily, but not ex-

clusively, to statistics
• Descriptive and inferential (univariate and multi-
variate to a lesser extent) statistical methods, BUT
• Background and theory that you need to make
use of statistics properly in the context of geogra-
phy
3
Goals
• The goals of this course are:
– To help students understand the purpose,
meaning, and use of statistics in geographic
research.
– To introduce basic statistical methods used by
geographers in research.
– To learn how to use computer spreadsheets to
simplify geographic problem solving.
– To learn how to apply statistical techniques
to real geographical problems.
Terminology
• Population
– A collection of items of interest in research
– A complete set of things
– A group that you wish to generalize your re-
search to
– An example – All the trees in a Park
• Sample
– A subset of a population
– The size smaller than the size of a population
– An example – 100 trees randomly selected
from a Park
Terminology
• Representative – An accurate reflection of the
population (A primary problem in statistics)
• Variables – The properties of a population that
are to be measured (i.e., how do parts of the popu-
lation differ?
• Constant – Something that does not vary
• Parameter – A constant measure which describes
the characteristics of a population
• Statistic – The corresponding measure for a
sample
Two Sorts of Statistics
• Descriptive statistics
– To describe and summarize the characteristics
of the sample
– Fall within the class of exploratory techniques
• Inferential statistics
– To infer something about the population from
the sample
– Lie within the class of confirmatory methods
(i) Descriptive Statistics

• Descriptive statistics – Statistics that describe


and summarize the characteristics of a dataset
(sample or population)
• Descriptive methods – Fall within the class of
exploratory techniques
• The most common way of describing a variable
distribution is in terms of two of its properties:
central tendency & dispersion
Descriptive Statistics

• Measures of central tendency


– Measures of the location of the middle or the
center of a distribution
– Mean, median, mode

• Measures of dispersion
– Describe how the observations are distributed
– Variance, standard deviation, range, etc
Measures of Central Tendency – Mean

• Mean – Most commonly used measure of central ten-


dency
• Average of all observations
• The sum of all the scores divided by the number of
scores
• Note: Assuming that each observation is equally signifi-
cant
Mathematical Notation
• The mathematical notation used most often in this
course is the summation notation
• The Greek letter is used as a shorthand way of indi-
cating that a sum is to be taken:
i n

x
i 1
i

The expression is equivalent to:

x1  x2   xn
Summation Notation: Components

refers to where the


sum of terms ends

i n

x
i 1
i
indicates what we
are summing up

indicates we are
taking a sum

refers to where the


sum of terms begins
Summation Notation: Simplification
• A summation will often be written leaving out the
upper and/or lower limits of the summation, as-
suming that all of the terms available are to be
summed

i n n

 x  x  x
i 1
i
i 1
i i
Summation Notation: Examples

Example I: All observations are included in the sum:


1, 2, 3, 4, 5, 6, 7, 8, 9, 10
10

x
i 1
i  x1  x2  x3  x4  x5  x6  x7  x8  x9  x10
1  2  3  4  5  6  7  8  9  10
Example II: Only observations 3 through 5 are included
in the sum: 5

x
i 3
i  x3  x4  x5 3  4  5 12
Summation Notation: Rules

• Rule I: Summing a constant n times yields a result of na:

 a a  a   a na
i 1

• Here we are simply using the summation notation to


carry out a multiplication, e.g.:
5

4
i 1
4  4  4  4  4 4 5 20
Measures of Central Tendency – Mean

Sample mean: Population mean:

n N

 xi x i

x i 1  i 1

n N
Measures of Central Tendency – Mean
• Example I
– Data: 8, 4, 2, 6, 10
5

x i
(8  4  2  6  10)
x i 1
 6
5 5
• Example II
– Sample: 10 trees randomly selected from Battle Park
– Diameter (inches):
9.8, 10.2, 10.1, 14.5, 17.5, 13.9, 20.0, 15.5, 7.8, 24.5
10

x i
(9.8  10.2   24.5)
x 
i 1
14.38
10 10
Measures of Central Tendency – Mean
• Example III
Annual mean temperature (°F)

x 59.70

Monthly mean temperature (°F) at Chapel Hill, NC (2001).


Mean annual precipitation (mm) Example IV

Mean
1198.10 (mm)

Mean annual temperature (°F)

Mean
58.51 (°F)

Chapel Hill, NC
(1972-2001)
Measures of Central Tendency – Mean
• Advantage
– Sensitive to any change in the value of any observation
• Disadvantage
– Very sensitive to outliers

# Tree Height # Tree Height


(m) (m)
1 5.0 6 5.3
2 6.0 7 7.1
3 7.5 8 25.4
4 8.0 9 7.5
5 4.8 10 4.5

Source: [Link] Mean = 6.19 m Mean = 8.10 m


Measures of Central Tendency – Mean
• A standard geographic application of the mean is to lo-
cate the center (centroid) of a spatial distribution
• Assign to each member a gridded coordinate and calcu-
lating the mean value in each coordinate direction -->
Bivariate mean or mean center
• For a set of (x, y) coordinates, the mean center
is calculated as:
(x, y) n n

 xi y i

x i 1 y i 1

n n
Map Coordinates
• Geographic coordinates – The geographic coordinate
system is a system used to locate points on the surface
of the globe (degrees of latitude and longitude)

Geographic coordi-
nates of Chapel Hill,
NC

Lat: 35˚ 54’ 25’’N

Lon: 79˚ 02’ 55’’W


Source: Xiao & Moody, 2004
• Other map coordinates (UTM, state plane, Lambert,
etc)
• e.g., UTM (Universal Transverse Mercator)
– Chapel Hill, NC  X: 676096.67 Y: 3975379.18

Source: [Link]
Weighted Mean
• We can also calculate a weighted mean using some
weighting factor:
e.g. What is the average income of all
n people in cities A, B, and C:

 wi xi City
A
Avg. Income
$23,000 100,000
Population

x i 1
n
B $20,000 50,000

w
C $25,000 150,000

i
i 1 Here, population is the weighting fac-
tor and the average income is the variable
of interest
Weighted Mean Center
• We can also calculate a weighted mean center in
much the same way, by using weights:
n n

For a set of (x, y) coor- w x i i w y i i


dinates, the weighted x  i 1n y  i 1n
mean center
( x , y ) is computed as: w i w
i 1
i
i 1

e.g., suppose we had


the centroids and ar- Here we weight
eas of 3 polygons by area
Measures of Central Tendency – Median
• Median – This is the value of a variable such that half
of the observations are above and half are below this
value i.e. this value divides the distribution into two
groups of equal size
• When the number of observations is odd, the median
is simply equal to the middle value
• When the number of observations is even, we take
the median to be the average of the two values in the
middle of the distribution
Measures of Central Tendency – Median
• Example I
– Data: 8, 4, 2, 6, 10 (mean: 6)

2, 4, 6, 8, 10 median: 6

• Example II
– Sample: 10 trees randomly selected from a Park
– Diameter (inches):
9.8, 10.2, 10.1, 14.5, 17.5, 13.9, 20.0, 15.5, 7.8, 24.5
(mean: 14.38)

7.8, 9.8, 10.1, 10.2, 13.9, 14.5, 15.5, 17.5, 20.0, 24.5

median: (13.9 + 14.5) / 2 = 14.2


# Tree Height # Tree Height
(m) (m)
1 5.0 6 5.3
2 6.0 7 7.1
3 7.5 8 25.4
4 8.0 9 7.5
5 4.8 10 4.5

Source: [Link] Mean = 6.19 m Mean = 8.10 m

# Tree Height # Tree Height median: (6.0 + 7.1) = 6.55


(m) (m)
1 4.5 6 7.1
• Advantage: the value is NOT
2 4.8 7 7.5
3 5.0 8 7.5
affected by extreme values at
4 5.3 9 8.0 the end of a distribution (which
5 6.0 10 25.4 are potentially are outliers)
Measures of Central Tendency – Mode

• Mode - This is the most frequently occurring value


in the distribution
• This is the only measure of central tendency that
can be used with nominal data
• The mode allows the distribution's peak to be lo-
cated quickly
Mean = 6.19 m (without outlier)

Mean = 8.10 m
Source: [Link]
median: (6.0 + 7.1) = 6.55
# Tree Height # Tree Height
(m) (m)
1 4.5 6 7.1 mode: 7.5
2 4.8 7 7.5
3 5.0 8 7.5
4 5.3 9 8.0
5 6.0 10 25.4
30 40 25 50 45

50 55 45 48 61

60 75 70 45 72

24 45 200 205 65

65 39 58 45 65

Landsat ETM+, Chapel Hill (2002-05-24)


(7-4-1 band combination)

24, 25, 30, 39, 40, 45, 45, 45, 45, 45, 48, 50, 50,
55, 58, 60, 61, 65, 65, 65, 70, 72, 75, 200, 205

mean: 63.28 median: 50 mode:


45 mean (without outliers): 51.17
Which one is better: mean, median,
or mode?
• Most often, the mean is selected by default
• The mean's key advantage is that it is sensitive
to any change in the value of any observation
• The mean's disadvantage is that it is very sen-
sitive to outliers
• We really must consider the nature of the data,
the distribution, and our goals to choose prop-
erly
Which one is better: mean, median,
or mode?
• The mean is valid only for interval data or ratio
data.
• The median can be determined for ordinal data as
well as interval and ratio data.
• The mode can be used with nominal, ordinal, in-
terval, and ratio data
• Mode is the only measure of central tendency that
can be used with nominal data
Which one is better: mean, median,
or mode?
• It also depends on the nature of the distribution

Multi-modal distribution Unimodal symmetric

Unimodal skewed Unimodal skewed


Which one is better: mean, median,
or mode?
• It also depends on your goals
• Consider a company that has nine employees with
salaries of 35,000 a year, and their supervisor
makes 150,000 a year.
• If you want to describe the typical salary in the
company, which statistics will you use?
• I will use mode or median (35,000), because it tells
what salary most people get

Source: [Link]
Which one is better: mean, median,
or mode?
• It also depends on your goals
• Consider a company that has nine employees with
salaries of 35,000 a year, and their supervisor
makes 150,000 a year
• What if you are a recruiting officer for the com-
pany that wants to make a good impression on a
prospective employee?
• The mean is (35,000*9 + 150,000)/10 = 46,500 I
would probably say: "The average salary in our
company is 46,500" using mean
Source: [Link]
Which one is better: mean, median,
or mode?
• The mean is valid only for interval data or ratio
data
• The median can be determined for ordinal data as
well as interval and ratio data
• The mode can be used with nominal, ordinal, in-
terval, and ratio data
• Mode is the only measure of central tendency that
can be used with nominal data
• It also depends on the nature of the distribution
and your goals
Some Characteristics of Data
• Not all data is the same. There are some limitations
as to what can and cannot be done with a data set,
depending on the characteristics of the data
• Some key characteristics that must be considered
are:
• A. Continuous vs. Discrete
• B. Grouped vs. Individual
• C. Scale of Measurement
A. Continuous vs. Discrete Data
• Continuous data can include any value (i.e., real
numbers)
– e.g., 1, 1.43, and 3.1415926 are all acceptable values.
– Geographic examples: distance, tree height, amount of
precipitation, etc
• Discrete data only consists of discrete values,
and the numbers in between those values are not
defined (i.e., whole or integer numbers)
– e.g., 1, 2, 3.
– Geographic examples: # of vegetation types,
B. Grouped vs. Individual Data
• The distinction between individual and
grouped data is somewhat self-explanatory, but
the issue pertains to the effects of grouping data
• While a family income value is collected for each
household (individual data), for the purpose of
analysis it is transformed into a set of classes
(e.g., <$10K, $10K-20K, > $20K)
• e.g., elevation (1000m vs. < 500m, 500-1000m,
1000-2000m, > 2000m)
B. Grouped vs. Individual Data

• In grouped data, the raw individual data is catego-


rized into several classes, and then analyzed
• The act of grouping the data, by taking the central
value of each class, as well as the frequency of the
class interval, and using those values to calculate a
measure of central tendency has the potential to in-
troduce a significant distortion
• Grouping always reduces the amount of information
contained in the data
C. Scales of Measurement

• Data is the plural of a datum, which are generated


by the recording of measurements
• Measurements involves the categorization of an
item (i.e., assigning an item to a set of types) when
the measure is qualitative
• or makes use of a number to give something a
quantitative measurement
C. Scales of Measurement

• The data used in statistical analyses can divided


into four types:
1. The Nominal Scale
As we progress through
2. The Ordinal Scale these scales, the types
of data they describe
3. The interval Scale have increasing infor-
4. The Ratio Scale mation content
The Nominal Scale
• Nominal scale data are data that can simply be
broken down into categories, i.e., having to do
with names or types
• Dichotomous or binary nominal data has just
two types, e.g., yes/no, female/male, is/is not,
hot/cold, etc
• Multichotomous data has more than two types,
e.g., vegetation types, soil types, counties, eye
color, etc
• Not a scale in the sense that categories cannot
be ranked or ordered (no greater/less than)
The Ordinal Scale
• Ordinal scale data can be categorized AND can
be placed in an order, i.e., categories that can be
assigned a relative importance and can be ranked
such that numerical category values have
– star-system restaurant rankings
5 stars > 4 stars, 4 stars > 3 stars, 5 stars > 2 stars
• BUT ordinal data still are not scalar in the sense
that differences between categories do not have a
quantitative meaning
– i.e., a 5 star restaurant is not superior to a 4 star restau-
rant by the same amount as a 4 star restaurant is than a
3 star restaurant
The Interval Scale
• Interval scale data take the notion of ranking
items in order one step further, since the distance
between adjacent points on the scale are equal
• For instance, the Fahrenheit scale is an interval
scale, since each degree is equal but there is no
absolute zero point.
• This means that although we can add and sub-
tract degrees (100° is 10° warmer than 90°), we
cannot multiply values or create ratios (100° is
not twice as warm as 50°)
The Ratio Scale
• Similar to the interval scale, but with the addition
of having a meaningful zero value, which allows
us to compare values using multiplication and
division operations, e.g., precipitation, weights,
heights, etc
• e.g., rain – We can say that 2 inches of rain is
twice as much rain as 1 inch of rain because this
is a ratio scale measurement
• e.g., age – a 100-year old person is indeed twice
as old as a 50-year old one
Measures of Central Tendency
• Measurements of Central Tendency
– Mean, median, mode, weighted mean
– Advantages vs. disadvantages
– Which one is better?
• Nature of Data
– A. Continuous vs. discrete
– B. Grouped vs. individual
– C. Scale of measurement
Scales of Measurement

• The data used in statistical analyses can be di-


vided into four types:
1. The Nominal Scale
As we progress through
2. The Ordinal Scale these scales, the types
of data they describe
3. The interval Scale have increasing infor-
4. The Ratio Scale mation content
The Nominal Scale
• Nominal scale data are data that can simply be
broken down into categories, i.e., having to do
with names or types
• Dichotomous or binary nominal data has just
two types, e.g., yes/no, female/male, is/is not,
hot/cold, etc
• Multichotomous data has more than two types,
e.g., vegetation types, soil types, counties, eye
color, etc
• Not a scale in the sense that categories cannot
be ranked or ordered (no greater/less than)
The Ordinal Scale
• Ordinal scale data can be categorized AND can
be placed in an order, i.e., categories that can be
assigned a relative importance and can be ranked
such that numerical category values have
– star-system restaurant rankings
5 stars > 4 stars, 4 stars > 3 stars, 5 stars > 2 stars
• BUT ordinal data still are not scalar in the sense
that differences between categories do not have a
quantitative meaning
– i.e., a 5 star restaurant is not superior to a 4 star restau-
rant by the same amount as a 4 star restaurant is than a
3 star restaurant
The Interval Scale
• Interval scale data take the notion of ranking
items in order one step further, since the distance
between adjacent points on the scale are equal
• For instance, the Fahrenheit scale is an interval
scale, since each degree is equal but there is no
absolute zero point.
• This means that although we can add and sub-
tract degrees (100° is 10° warmer than 90°), we
cannot multiply values or create ratios (100° is
not twice as warm as 50°)
The Ratio Scale
• Similar to the interval scale, but with the addition
of having a meaningful zero value, which allows
us to compare values using multiplication and
division operations, e.g., precipitation, weights,
heights, etc
• e.g., rain – We can say that 2 inches of rain is
twice as much rain as 1 inch of rain because this
is a ratio scale measurement
• e.g., age – a 100-year old person is indeed twice
as old as a 50-year old one
(ii) Measures of dispersion
• Measures of Dispersion
– Range
– Variance
– Standard deviation
– z-score
– Coefficient of variation
• Other Descriptive Summary Measures
– Ratios, Proportions, Percentages
– Rates of change, Location quotients
Frequency & Frequency Distribution
• Frequency is the number of times a variable takes
on a particular value
• A frequency distribution is one of the most
common graphical tools used to describe a single
population
• It is a tabulation of the frequencies of each value
(or range of values)
• There are a wide variety of ways to illustrate fre-
quency distributions, including histograms, rela-
tive frequency histograms, density histograms,
and cumulative frequency distributions.
Why Do We Need Measures of Disper-
sion?
• Measures of central tendency tell us nothing about
the variability / dispersion / deviation / range of val-
ues about the central value. Consider the following
two unimodal symmetric distributions:

Source: Earickson, RJ, and Harlin, JM. 1994. Geographic Measurement and Quantitative Analysis.
USA: Macmillan College Publishing Co., p. 91.
Source: [Link]
Source: [Link]
We Need Both Measures!

• Data sets may have similar central tendencies but


different levels of dispersion
• Data set may have also similar levels of dispersion
but different central tendencies
• Therefore, we need to use a measure of both cen-
tral tendency and dispersion in order to obtain
less ambiguous numerical description of our data
set.
Source: [Link]
Source: [Link]
GTOPO30 is a Global Digital Elevation Model
(DEM) (GTOPO30, 1km)

Source: [Link]
MODIS 16-day NDVI composite (Terra, 08/12/2004-08/27/2004)
(Original data obtained from Global Land Cover Facility, [Link]
Landsat ETM+ image at Chapel Hill, NC (2002-05-24)
(7-4-1)
Monthly mean temperature (°F) at Chapel Hill, NC (2001).
Measures of Dispersion
• In addition to measures of central tendency, we
can also summarize data by characterizing its
variability
• Measures of dispersion are concerned with the
distribution of values around the mean in data:
– Range
– Interquartile range
– Variance
– Standard deviation
– z-scores
Measures of Dispersion - Range
• Range – this is the most simply formulated of all
measures of dispersion
• Given a set of measurements x1, x2, x3, … ,xn-1, xn ,
the range is defined as the difference between the
largest and smallest values:
Range = xmax – xmin
• This is another descriptive measure that is vulner-
able to the influence of outliers in a data set,
which result in a range that is not really descriptive
of most of the data
Measures of Dispersion - Range

• The diurnal temperature range (DTR) is the differ-


ence between the night time low temperature and
the daytime high temperature, usually for a given
day
• Example: July 01, 2005, Chapel Hill
– Highest hourly temperature: 93 (°F)
– Lowest hourly temperature: 72 (°F)
• DTR: 93 – 72 = 21 (°F)
Measures of Dispersion - Range

• Example I Annual mean temperature (°F)


x 59.70

Range = 77.29 – 39.53


= 37.76 (°F)

Monthly mean temperature (°F) at Chapel Hill, NC (2001).


Mean = 6.19 m (without outlier)

Mean = 8.10 m

Source: [Link]

median: (6.0 + 7.1) =


# Tree Height # Tree Height
(m) (m)
6.55
1 4.5 6 7.1 mode: 7.5
2 4.8 7 7.5
3 5.0 8 7.5
4 5.3 9 8.0
range: 25.4 – 4.5 = 20.9 m
5 6.0 10 25.4
30 40 25 50 45

50 55 45 48 61

60 75 70 45 72

24 45 200 205 65

65 39 58 45 65

Landsat ETM+, Chapel Hill (2002-05-24)


(7-4-1 band combination)

24, 25, 30, 39, 40, 45, 45, 45, 45, 45, 48, 50, 50,
55, 58, 60, 61, 65, 65, 65, 70, 72, 75, 200, 205

mean: 63.28 median: 50 mode:


45
mean (without outliers): 51.17 range: 205 – 24 = 181
Measures of Dispersion – Interquar-
tile Range

• Quartiles – We can divide distributions into four


parts each containing 25% of observations
• Percentiles – each contains 1% of all values
• Interquartile range – The difference between the
25th and 75th percentiles
Measures of Dispersion – Interquar-
tile Range
• Quartiles – We can divide distributions into four parts
each containing 25% of observations
• When the data are arranged in order of magnitude (i.e.
they are ranked) the quartiles are 3 numbers which di-
vide the data into four groups each having approxi-
mately the same number of values (25%)
• Procedure for Calculating Quartiles
1. Order the n data values from smallest to largest
2. The 2nd quartile, Q2 is the median of the whole
data set
3. If n is even, the first quartile, Q1, is the median of
the smallest n/2 observations and the third quartile,
Q3, is the median of the largest n/2 observations
If n is odd, Q1 is the median of the smallest ob-
servations, and Q3 is the median of the largest ob-
servations
Interquartile Range

• The interquartile range is defined as IQR = Q3 - Q1


Interquartile Range
• Example – Consider first 9 Commodore prices ( in
$,000)
6.0, 6.7, 3.8, 7.0, 5.8, 9.975, 10.5, 5.99, 20.0
• Arrange these in order of magnitude
3.8, 5.8, 5.99, 6.0, 6.7, 7.0, 9.975, 10.5, 20.0
• The median is Q2 = 6.7 (there are 4 values on ei-
ther side)
• Q1 = 5.9 (median of the 4 smallest values)
• Q3 = 10.2 (median of the 4 largest values)
• IQR = Q3 – Q1 = 10.2 - 5.9 = 4.3
Interquartile Range & Percentiles
• Quartiles divide the ordered data into quarters, but
we can consider any fractions we please
• The most common are "percentiles", where we
take hundredths
• Percentiles – each contains 1% of all values
• The first quartile is thus the 25th percentile, the
median is the 50th percentile and the upper quartile
is the 75th percentile
• Interquartile range – The difference between the
25th and 75th percentiles
Measures of Dispersion – In-
terquartile Range
• With n observations, the 25th percentile is repre-
sented by observation (n+1)/4, when the data have
been ranked from lowest to highest
• The 75 percentile is represented by observation
3(n+1)/4
• These will often not be integers, and interpolation
is used, just as it is for the median when there is an
even number of observations
• Interquartile range is the difference between these
two percentiles
Measures of Dispersion – In-
terquartile Range
Example: Table 1.1 Commuting data (Rogerson, p5)

Ranked commuting times:

5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47

25th percentile is represented by observation (30+1)/4=7.75


75th percentile is represented by observation 3(30+1)/4=23.25
25th percentile: 11.75
75th percentile: 26
Interquartile range: 26 – 11.75 = 14.25
Measures of Dispersion –
Mean Deviation
• Mean Deviation – Once we have calculated the
mean value for a data set, we can assess the differ-
ence between any observation and that mean, and
this is termed the statistical distance:
Statistical distance = xi - x
• If we take the absolute values of these, and sum
for all observations, we have calculated the mean
deviation: n

Mean deviation =
S |
i=1
x i – x |
n
Measures of Dispersion –
Mean Deviation
• Mean deviation -- Why is it necessary to take absolute
values of statistical distances before summing them to
get the mean deviation?
• Because the statistical distances would be both positive
and negative, and when summed using the mean devia-
tion formula, they would sum to zero

1
 ( xi  x)  xi   x  xi  n x  xi  n n  xi 0
Measures of Dispersion –
Variance, Standard Deviation
• As an alternative to taking the absolute values
of the statistical distances, we can square
each deviation before taking their sum, which
yields the sum of squares
n

sum of squares =
S (x
i=1
i – x) 2

• The sum of squares expresses the total square


variation about the mean, and using this value
we can calculate variances and standard devi-
ations for both populations and samples
Measures of Dispersion – Variance
• Variance is formulated as the sum of squares of
statistical distances divided by the population
size or the sample size minus one:
n N

2
 ( xi  x ) 2
 (x  x)
i
2

s  i 1
  i 1
2

n 1 N
Sample variance Population variance
• Note the differences in the two formulae, both in
the notation but also the denominators!
Measures of Dispersion – Variance
• To ensure that the sample variance gives an un-
biased estimate of the true, unknown variance of
the population from which the sample was drawn,
sample variance is computed by taking the sum
of the squared deviations, and then dividing by n –
1, instead of by n
• Unbiased implies that if we were to repeat this
sampling many times, we would find that the aver-
age or mean of our many sample variances would
be equal to the true variance
Chapel Hill Bend
Mean 1198.10 298.07
Variance 36786.34 6737.095

Annual Precipitation (mm) at Chapel Hill, NC and Bend, OR (1972-2001)


Measures of Dispersion – Standard
Deviation
• Standard deviation is equal to the square root of the
variance

 (x  x)i
2

s i 1
n 1
• Compared with variance, standard deviation has a
scale closer to that used for the mean and the original
data
Chapel Hill Bend
Mean 1198.10 298.07
Variance 36786.34 6737.095
Standard Deviation 191.80 82.08

Annual Precipitation (mm) at Chapel Hill, NC and Bend, OR (1972-2001)


Measures of Dispersion – z-score
• Since data come from distributions with different means
and difference degrees of variability, it is common to
standardize observations
• One way to do this is to transform each observation into
a z-score
xi  x
z
s
• May be interpreted as the number of standard devia-
tions an observation is away from the mean
Measures of Dispersion – z-score
Month T (°F) Z-score
J 39.53 -1.56
F 46.36 -1.03
M 46.42 -1.02
A 60.32 0.05
M 66.34 0.51
J 75.49 1.22
J 75.39 1.21
A 77.29 1.36
S 68.64 0.69
O 57.57 -0.16
N 54.88 -0.37
D 48.2 -0.89

Mean = 59.70
Standard deviation = 12.97 Monthly temperature at Chapel Hill, NC (2001)
Other measures of dispersion

• Another Measure of Dispersion


– Coefficient of Variation (CV)
• Histograms
• Skewness
• Kurtosis
• Other Descriptive Summary Measures
Measures of Dispersion – Coefficient
of Variation
• Coefficient of variation (CV) measures the spread of a set
of data as a proportion of its mean.
• It is the ratio of the sample standard deviation to the
sample mean
s
CV  100%
x
• It is sometimes expressed as a percentage
• There is an equivalent definition for the coefficient of
variation of a population
Measures of Dispersion – Coefficient
of Variation
• A standard application of the Coefficient of Varia-
tion (CV) is to characterize the variability of geo-
graphic variables over space or time
• Coefficient of Variation (CV) is particularly ap-
plied to characterize the interannual variability of
climate variables (e.g., temperature or precipita-
tion) or biophysical variables (leaf area index
(LAI), biomass, etc)
Chapel Hill Bend
(A) (B)
Mean 1198.10 298.07

Standard Deviation 191.80 82.08

Coefficient of Variation 0.16 0.28


(CV) (16%) (28%)
Coefficient of Variation (CV)
• It is a dimensionless number that can be used to
compare the amount of variance between popula-
tions with different means

 (x  x) 2 n

2
i
 (x  x)
i
2

s  i 1
s i 1
n 1 n 1
s
CV  100%
x
Measures of Skewness and Kurtosis
• A fundamental task in many statistical analyses is
to characterize the location and variability of a
data set (Measures of central tendency vs. mea-
sures of dispersion)
• Both measures tell us nothing about the shape of
the distribution
• A further characterization of the data includes
skewness and kurtosis
• The histogram is an effective graphical technique
for showing both the skewness and kurtosis of a
data set
Histograms

Fig. 3. Histogram of crown width (m) measured in situ for a random sample
of Quercus robur trees in Frame Wood (n = 63; mean = 9.3 m; SD = 4.64 m).
Source: Koukoulas & Blackburn, 2005. Journal of Vegetation Science: Vol. 16, No. 5, pp. 587–596
Frequency & Distribution
• A histogram is one way to depict a frequency
distribution
• Frequency is the number of times a variable takes
on a particular value
• Note that any variable has a frequency distribution
• e.g. roll a pair of dice several times and record the
resulting values (constrained to being between and
2 and 12), counting the number of times any given
value occurs (the frequency of that value occur-
ring), and take these all together to form a fre-
quency distribution
Frequency & Distribution
• Frequencies can be absolute (when the frequency
provided is the actual count of the occurrences) or
relative (when they are normalized by dividing the
absolute frequency by the total number of observa-
tions [0, 1])
• Relative frequencies are particularly useful if you
want to compare distributions drawn from two dif-
ferent sources (i.e. while the numbers of observa-
tions of each source may be different)
Histograms
• We may summarize our data by constructing his-
tograms, which are vertical bar graphs
• A histogram is used to graphically summarize
the distribution of a data set
• A histogram divides the range of values in a data
set into intervals
• Over each interval is placed a bar whose height
represents the frequency of data values in the in-
terval.
Building a Histogram
• To construct a histogram, the data are first
grouped into categories
• The histogram contains one vertical bar for each
category
• The height of the bar represents the number of
observations in the category (i.e., frequency)
• It is common to note the midpoint of the category
on the horizontal axis
Building a Histogram – Example
• 1. Develop an ungrouped frequency table
– That is, we build a table that counts the number of oc-
currences of each variable value from lowest to high-
est:
TMI Value Ungrouped Freq.
4.16 2
4.17 4
4.18 0
… …
13.71 1
• We could attempt to construct a bar chart from this table,
but it would have too many bars to really be useful
Building a Histogram – Example
• 2. Construct a grouped frequency table
– Select an appropriate number of classes

Class Frequency Percentage


4.00 - 4.99 120
5.00 - 5.99 807
6.00 - 6.99 1411
7.00 - 7.99 407
8.00 - 8.99 87
9.00 - 9.99 33
10.00 - 10.99 17
11.00 - 11.99 22
12.00 - 12.99 43
13.00 - 13.99 19
Building a Histogram – Example
• 3. Plot the frequencies of each class
– All that remains is to create the bar graph

Pond Branch TMI Histogram

48
Percent of cells in catchment

44
40
36
32
28
24
20
16 A proxy for
12
8 Soil Moisture
4
0
4 5 6 7 8 9 10 11 12 13 14 15 16

Topographic Moisture Index


Further Moments of the Distribution
• While measures of dispersion are useful for helping
us describe the width of the distribution, they tell us
nothing about the shape of the distribution

Source: Earickson, RJ, and Harlin, JM. 1994. Geographic Measurement and Quantitative Analysis. USA: Macmil-
lan College Publishing Co., p. 91.
Further Moments of the Distribution

• There are further statistics that describe the


shape of the distribution, using formulae that are
similar to those of the mean and variance
• 1st moment - Mean (describes central value)
• 2nd moment - Variance (describes dispersion)
• 3rd moment - Skewness (describes asymmetry)
• 4th moment - Kurtosis (describes peakedness)
Further Moments – Skewness

• Skewness measures the degree of asymmetry exhibited


by the data
n

 (x  x)
i
3

skewness  i 1
3
ns
• If skewness equals zero, the histogram is symmetric
about the mean
• Positive skewness vs negative skewness
Further Moments – Skewness

Source: [Link]
Further Moments – Skewness
• Positive skewness
– There are more observations below the mean
than above it
– When the mean is greater than the median
• Negative skewness
– There are a small number of low observations
and a large number of high ones
– When the median is greater than the mean
Further Moments – Kurtosis
• Kurtosis measures how peaked the histogram is
n

 (x  x)
i
4

kurtosis  i
4
3
ns
• The kurtosis of a normal distribution is 0
• Kurtosis characterizes the relative peakedness or flat-
ness of a distribution compared to the normal distribu-
tion
Further Moments – Kurtosis
• Platykurtic– When the kurtosis < 0, the frequen-
cies throughout the curve are closer to be equal
(i.e., the curve is more flat and wide)
• Thus, negative kurtosis indicates a relatively flat
distribution
• Leptokurtic– When the kurtosis > 0, there are
high frequencies in only a small part of the curve
(i.e, the curve is more peaked)
• Thus, positive kurtosis indicates a relatively
peaked distribution
Further Moments – Kurtosis

platykurtic leptokurtic

Source: [Link]

• Kurtosis is based on the size of a distribution's


tails.
• Negative kurtosis (platykurtic) – distributions with
short tails
• Positive kurtosis (leptokurtic) – distributions with
relatively long tails
Why Do We Need Kurtosis?

• These two distributions have the same variance,


approximately the same skew, but differ markedly
in kurtosis.
Source: [Link]
How to Graphically Summarize Data?

• Histograms

• Box plots
Functions of a Histogram
• The function of a histogram is to graphically
summarize the distribution of a data set
• The histogram graphically shows the following:
1. Center (i.e., the location) of the data
2. Spread (i.e., the scale) of the data
3. Skewness of the data
4. Kurtosis of the data
4. Presence of outliers
5. Presence of multiple modes in the data.
Functions of a Histogram

• The histogram can be used to answer the fol-


lowing questions:
1. What kind of population distribution do the
data come from?
2. Where are the data located?
3. How spread out are the data?
4. Are the data symmetric or skewed?
5. Are there outliers in the data?
Skewness & Kurtosis: Reference

Source: [Link]
Further Moments – Skewness
• Skewness measures the degree of asymmetry exhibited
by the data n

 (x  x)
i
3

skewness  i 1
3
ns
• If skewness equals zero, the histogram is symmetric
about the mean
• Positive skewness vs negative skewness
• Skewness measured in this way is sometimes referred to
as “Fisher’s skewness”
Further Moments – Skewness

Source: [Link]
n
Mode
Median
 i
( x  x ) 3

skewness  i 1
Mean
ns 3

A B
Median

Mean

n = 26 mean = 4.23 median = 3.5 mode = 8


n

 i
( x  x ) 3

skewness  i 1 3
ns

Value Occurrences Deviation Cubed deviation Occur*Cubed

1 1 (1 – 4.23) = -3.23 (-3.23)3 = -33.70 -33.70


2 4 (2 – 4.23) = -2.23 (-2.23)3 = -11.09 -44.36
3 8 (3 – 4.23) = -1.23 (-1.13)3 = -1.86 -14.89
4 4 (4 – 4.23) = -0.23 (-0.23)3 = -0.01 -0.05
5 3 (5 – 4.23) = 0.77 (+0.77)3 = 0.46 1.37
6 2 (6 – 4.23) = 1.77 (+1.77)3 = 5.54 11.09
7 1 (7 – 4.23) = 2.77 (+2.77)3 = 21.25 21.25
8 1 (8 – 4.23) = 3.77 (+3.77)3 = 53.58 53.58
9 1 (9 – 4.23) = 4.77 (+4.77)3 = 108.53 108.53
10 1 (10 - 4.23)= 5.77 (+5.77)3 = 192.10 192.10
Mean = 4.23 Sum = 294.94
s = 2.27 Skewness = 0.97
n
Mode
Median
 (x  x)
i
3

skewness  i 1 3
Mean
ns

Skewness > 0 (Positively skewed)


n

 i
( x  x ) 3
Mode
skewness  i 1
ns 3 Median

Mean

A B

Skewness < 0 (Negatively skewed)


n

 (x  x) i
3

skewness  i 1
ns 3

Source: [Link]

Skewness = 0 (symmetric distribution)


Skewness – Review
• Positive skewness
– There are more observations below the mean
than above it
– When the mean is greater than the median
• Negative skewness
– There are a small number of low observations
and a large number of high ones
– When the median is greater than the mean
Kurtosis – Review
• Kurtosis measures how peaked the histogram is (Karl
Pearson, 1905)
n

 (x  x)
i
4

kurtosis  i
4
3
ns
• The kurtosis of a normal distribution is 0
• Kurtosis characterizes the relative peakedness or flat-
ness of a distribution compared to the normal distribu-
tion
Kurtosis – Review
• Platykurtic– When the kurtosis < 0, the frequen-
cies throughout the curve are closer to be equal
(i.e., the curve is more flat and wide)
• Thus, negative kurtosis indicates a relatively flat
distribution
• Leptokurtic– When the kurtosis > 0, there are
high frequencies in only a small part of the curve
(i.e, the curve is more peaked)
• Thus, positive kurtosis indicates a relatively
peaked distribution
n

 (x  x)
i
4

kurtosis  i
4
3
ns
Source: [Link]
Source: [Link] (First three)
[Link] (Last)
Box Plots
• We can also use a box plot to graphically sum-
marize a data set
• A box plot represents a graphical summary of
what is sometimes called a “five-number sum-
mary” of the distribution
– Minimum
– Maximum
– 25th percentile
max. 75th
– 75th percentile %-ile
– Median median
25th
• Interquartile Range (IQR) min. %-ile
Rogerson, p. 8.
Box Plots
• Example – Consider first 9 Commodore prices ( in
$,000)
6.0, 6.7, 3.8, 7.0, 5.8, 9.975, 10.5, 5.99, 20.0
• Arrange these in order of magnitude
3.8, 5.8, 5.99, 6.0, 6.7, 7.0, 9.975, 10.5, 20.0
• The median is Q2 = 6.7 (there are 4 values on ei-
ther side)
• Q1 = 5.9 (median of the 4 smallest values)
• Q3 = 10.2 (median of the 4 largest values)
• IQR = Q3 – Q1 = 10.2 - 5.9 = 4.3
• Example (ranked)
3.8, 5.8, 5.99, 6.0, 6.7, 7.0, 9.975, 10.5, 20.0
• The median is Q1 = 6.7
• Q1 = 5.9 Q3 = 10.2 IQR = Q3 – Q1 = 10.2 - 5.9 =
4.3
Box Plots

Example: Table 1.1 Commuting data (Rogerson, p5)

Ranked commuting times:

5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47

25th percentile is represented by observation (30+1)/4=7.75


75th percentile is represented by observation 3(30+1)/4=23.25
25th percentile: 11.75
75th percentile: 26
Interquartile range: 26 – 11.75 = 14.25
Example (Ranked commuting times):

5, 5, 6, 9, 10, 11, 11, 12, 12, 14, 16, 17, 19, 21, 21, 21, 21,
21, 22, 23, 24, 24, 26, 26, 31, 31, 36, 42, 44, 47
25th percentile: 11.75 75th percentile: 26
Interquartile range: 26 – 11.75 = 14.25
Other Descriptive Summary Measures

• Descriptive statistics provide an organization


and summary of a dataset
• A small number of summary measures replaces
the entirety of a dataset
• We’ll briefly talk about other simple descriptive
summary measures
Other Descriptive Summary Measures

• You're likely already familiar with some simple


descriptive summary measures
– Ratios
– Proportions
– Percentages
– Rates of Change
– Location Quotients
Other Descriptive Summary Measures
• Ratios –
# of observations in A
=
# of observations in B
e.g., A - 6 overcast, B - 24 mostly cloudy days
• Proportions – Relates one part or category of data to
the entire set of observations, e.g., a box of marbles
that contains 4 yellow, 6 red, 5 blue, and 2 green gives
a yellow proportion of 4/17 or
colorcount = {yellow, red, blue, green}
ai
acount = {4, 6, 5, 2} proportion 
 ai
Other Descriptive Summary Measures
• Proportions - Sum of all proportions = 1. These
are useful for comparing two sets of data w/dif-
ferent sizes and category counts, e.g., a different
box of marbles gives a yellow proportion of 2/23,
and in order for this to be a reasonable compari-
son we need to know the totals for both samples
• Percentages - Calculated by proportions x 100,
e.g., 2/23 x 100% = 8.696%, use of these should
be restricted to larger samples sizes, perhaps
20+ observations
Other Descriptive Summary Measures
• Location Quotients - An index of relative concentration in
space, a comparison of a region's share of something to the
total
• Example – Suppose we have a region of 1000 Km2 which
we subdivide into three smaller areas of 200, 300, and 500
km2 (labeled A, B, & C)
• The region has an influenza outbreak with 150 cases in A,
100 in B, and 350 in C (a total of 600 flu cases): (Interpreta-
tion: LQ>1= relative concentration, LQ=1, in acccordance
with its share, LQ<1, the area has a less share)
Proportion of Area Proportion of Cases Location Quo-
tient
A 200/1000=0.2 150/600=0.25 0.25/0.2=1.25
B 300/1000=0.3 100/600=0.17 0.17/0.3 =
0.57
C 500/1000=0.5 350/600=0.58 0.58/0.5=1.17
Part Two: Inferential Statistics

• Descriptive statistics (mainly for samples)


• Our objective is to make a statement with reference to
a parameter describing a population
• Inferential statistics does this using a two-part process:
• (1) Estimation (of a population parameter)
• (2) Hypothesis testing
Inferential Statistics

• Estimation (of a population parameter) - The estimation


part of the process calculates an estimate of the parameter
from our sample (called a statistic), as a kind of “guess”
as to what the population parameter value actually is
• Hypothesis testing - This takes the notion of estimation
a little further; it tests to see if a sampled statistic is really
different from a population parameter to a significant
extent, which we can express in terms of the probability
of getting that result
Estimation
• Another term for a statistic is a point estimate, which is
simply an estimate of a population parameter
• The formula you use to compute a statistic is an estima-
tor, e.g.
i=n

Point
Sx
i=1
i
x= Estimator
Estimate n
• In this case, the sample mean is being used to estimate
m, the population mean
Estimation
• It is quite unlikely that our statistic will be exactly the
same as the population parameter (because we know
that sampling error does occur), but ideally it should be
pretty close to ‘right’, perhaps within some specified
range of the parameter
• We can define this in terms of our statistic falling
within some interval of values around the parameter
value (as determined by our sampling distribution)
• But how close is close enough?
Estimation and Confidence

• We can ask this question more formally:


• (1) How confident can we be that a statistic falls within
a certain distance of a parameter
• (2) What is the probability that the parameter is within
a certain range that includes our sample statistic
• This range is known as a confidence interval
• This probability is the confidence level
Confidence Interval & Probability
• A confidence interval is expressed in terms of a range
of values and a probability (e.g. my lectures are be-
tween 60 and 70 minutes long 95% of the time)
• For this example, the confidence level that I used is the
95% level, which is the most commonly used confi-
dence level
• Other commonly selected confidence levels are 90%
and 99%, and the choice of which confidence level to
use when constructing an interval often depends on the
application
Central Limit Theorem
• We have now discussed both the notions of probability
and the way that the normal distribution
• By combining these two concepts, we can go further
and state some expectations about how the statistics
that we derive from a sample might relate to the pa-
rameters that describe the population from which the
sample is drawn
• The approach that we use to construct confidence inter-
vals relies upon the central limit theorem
The Central Limit Theorem
• Suppose we draw a random sample of size n (x1,
x2, x3, … xn – 1, xn) from a population random variable
that is distributed with mean µ and standard deviation
σ
• Do this repeatedly, drawing many samples from the
population, and then calculate the x of each sample
• We will treat the x values as another distribution,
which we will call the sampling distribution of the
mean ( X )
The Central Limit Theorem
• Given a distribution with a mean μ and variance σ 2, the
sampling distribution of the mean approaches a nor-
mal distribution with a mean (μ) and a variance σ2/n as
n, the sample size, increases
• The amazing and counter- intuitive thing about the cen-
tral limit theorem is that no matter what the shape of
the original (parent) distribution, the sampling distri-
bution of the mean approaches a normal distribution
Central Limit Theorem
• A normal distribution is approached very quickly as n
increases
• Note that n is the sample size for each mean and not the
number of samples
• Remember in a sampling distribution of the mean the
number of samples is assumed to be infinite
• Foundation for many statistical procedures because the
distribution of the phenomenon under study does not
have to be normal because its average will be
Central Limit Theorem
• Three different components of the central limit theo-
rem
• (1) successive sampling from a population
• (2) increasing sample size
• (3) population distribution
• Keep in mind that this theorem applies only to the
mean and not other statistics
Central Limit Theorem – Example

• On the right are shown the resulting


frequency distributions each based on
500 means. For n = 4, 4 scores were
sampled from a uniform distribution
500 times and the mean computed
each time. The same method was fol-
lowed with means of 7 scores for n = 7
and 10 scores for n = 10.

• When n increases:
1. The distributions becomes more and more normal
2. The spread of the distributions decreases
Source: [Link]
Central Limit Theorem – Example

The distribution of an average tends to be normal,


even when the distribution from which the average is
computed is decidedly non-normal.
Source: [Link]
Central Limit Theorem & Confi-
dence Intervals for the Mean
• The central limit theorem states that given a distribu-
tion with a mean μ and variance σ2, the sampling dis-
tribution of the mean approaches a normal distribu-
tion with a mean (μ) and a variance σ2/n as n, the
sample size, increases
• Since we know something about the distribution of
sample means, we can make statements about how
confident we are that the true mean is within a given
interval about our sample mean
Standard Error
• The standard deviation of the sampling distribution
of the mean (X) is formulated as:
s
sX =
n
• This is the standard error (the unit of measurement
of a confidence interval, used to express the close-
ness of a statistic to a parameter
• When we construct a confidence interval we are
finding how many standard errors away from the
mean we have to go to find the area under the
curve equal to the confidence level
99.7%

95%

68%

f(x)

-3σ -2σ -1σ μ +1σ +2σ +3σ

P(Z>=2.0) = 0.0228 P(-2<=Z<=+2) = 1 – 2*0.0228 = 0.9544

P(Z>=1.96) = 0.025 P(-1.96<=Z<=+1.96) = 1 – 2*0.025 = 0.95


Confidence Intervals for the Mean
• The sampling distribution of the mean roughly follows a
normal distribution
• 95% of the time, an individual sample mean should lie
within 2 (actually 1.96) standard deviations of the mean

pr   1.96s   x   1.96s  0.95


Confidence Intervals for the Mean

2  2

s  s
N N

pr  1.96s   x   1.96s  0.95

     
pr     1.96   x    1.96   0.95
 N  N 
Confidence Intervals for the Mean
• An individual sample mean should, 95% of the time,
lie within of(the
1.96 / ntrue
) mean, μ

     
pr     1.96   x    1.96   0.95
 n  n 
• Rearrange the expression:

     
pr   x  1.96    x  1.96   0.95
 n  n 
• This tells us that 95% of the time the true mean
should lie within 1.96( / n ) of the sample mean
Confidence Intervals for the Mean

• More generally, a (1- α)*100% confidence interval


around the sample mean is: margin of
Standard error
error
     
pr   x  z    x  z   1  
 n  n 
• Where zα is the value taken from the z-table that is as-
sociated with a fraction α of the weight in the tails
(and therefore α/2 is the area in each tail)
Example
• Income Example: Suppose we take a sample of 75
students from UNC and record their family incomes.
Suppose the incomes (in thousands of dollars) are:
28 29 35 42 ··· 158 167 235

x 89.96, s 51.68
     
pr   x  1.96    x  1.96   0.95
 n  n 
Source: [Link]
Example

     
pr   89.96  1.96    89.96  1.96   0.95
 75   75  

• We don't know σ so we can't use the interval! We will


replace σ by the sample standard deviation s

 51.68   51.68  
pr   89.96  1.96    89.96  1.96   0.95
 75   75  
Example

 51.68   51.68  
pr   89.96  1.96    89.96  1.96   0.95
 75   75  

(89.96 – 1.96*5.97, 89.96 + 1.96*5.97)

(78.26, 101.66)

pr78.26  101.66 0.95


Constructing a Confidence Interval
• 1. Select our desired confidence level (1-α)*100%
• 2. Calculate α and α/2
• 3. Look up the corresponding z-score in a standard
normal table
• 4. Multiply the z-score by the standard error to
find the margin of error
• 5. Find the interval by adding and subtracting this
product from the mean
Constructing a Confidence
Interval - Steps
• Select our desired level of confidence
• Let’s suppose we want to construct an interval
using the 95% confidence level
• Calculate α and α/2
• (1-α)*100% = 95%  α = 0.05, α/2 = 0.025
3. Look up the corresponding z-score
• α/2 = 0.025  a z-score of 1.96
Constructing a Confidence
Interval - Steps
4. Multiply the z-score by the standard error to find
the margin of error
 
Z / 2  1.96 
n n

5. Find the interval by adding and subtracting this


product from the mean
( x  Z / 2 std .error , x  Z / 2 std .error )
 
( x  1.96  , x  1.96  )
n n
Common Confidence Levels and a
values
• For your convenience, here is a table of commonly
used confidence levels, α and α/2 values, and
corresponding z-scores:

(1 - a)*100% α α/2 Zα/2


90% 0.1 0.05 1.645
95% 0.05 0.025 1.96
99% 0.01 0.005 2.58
Constructing a Confidence
Interval - Example
• Suppose we conduct a poll to try and get a sense of the
outcome of an upcoming election with two candidates.
We poll 1000 people, and 550 of them respond that they
will vote for candidate A
• How confident can we be that a given person will cast
their vote for candidate A?
• Select our desired levels of confidence
• We’re going to use the 90%, 95%, and 99% levels
Constructing a Confidence
Interval - Example
• Calculate α and α/2
• Our a values are 0.1, 0.05, and 0.01 respectively
• Our a/2 values are 0.05, 0.025, and 0.005
3. Look up the corresponding z-scores
• Our Za/2 values are 1.645, 1.96, and 2.58
• Multiply the z-score by the standard error to find
the margin of error
• First we need to calculate the standard error
Constructing a Confidence
Interval - Example
• Find the interval by adding and subtracting this
product from the mean
• In this case, we are working with a distribution we have
not previously discussed, a normal binomial distribu-
tion (i.e. a vote can choose Candidate A or B, a binomial
function)
• We have a probability estimator from our sample, where
the probability of an individual in our sample voting for
candidate A was found to be 550/1000 or 0.55
• We can use this information in a formula to estimate the
standard error for such a distribution:
Constructing a Confidence
Interval - Example
• Multiply the z-score by the standard error cont.
• For a normal binominal distribution, the standard er-
ror can be estimated using:

s (0.55)(0.45)
sX = (p)(1-p) = = 0.0157
n =
n 1000

• We can now multiply this value by the z-scores to cal-


culate the margins of error for each conf. level
Constructing a Confidence
Interval - Example
• Multiply the z-score by the standard error cont.
• We calculate the margin of error and add and subtract
that value from the mean (0.55 in this case) to find the
bounds of our confidence intervals at each level of con-
fidence:
Margin Bounds
CI Za/2 of error Lower Up-
per
90% 1.645 0.026 0.524 0.576
95% 1.96 0.031 0.519 0.581
t-distribution
• The central limit theorem applies when the sample size
is “large”, only then will the distribution of means pos-
sess a normal distribution
• When the sample size is not “large”, the frequency dis-
tribution of the sample means has what is known as the
t-distribution
• t-distribution is symmetric, like the normal distribution,
but has a slightly different shape
• The t distribution has relatively more scores in its tails
than does the normal distribution. It is therefore lep-
tokurtic
t-distribution

• The t-distribution or Student's t-distribution is a


probability distribution that arises in the problem of es-
timating the mean of a normally distributed population
when the sample size is small
• It is the basis of the popular Student's t-tests for the
statistical significance of the difference between two
sample means, and for confidence intervals for the dif-
ference between two population means
t-distribution
• The derivation of the t-distribution was first published
in 1908 by William Sealy Gosset. He was not allowed to
publish under his own name, so the paper was written
under the pseudonym Student
• The t-test and the associated theory became well-known
through the work of R.A. Fisher, who called the distribu-
tion "Student's distribution"
• Student's distribution arises when (as in nearly all practi-
cal statistical work) the population standard deviation is
unknown and has to be estimated from the data
Confidence intervals & t-distribution
• The areas under the t-distribution are given in Table
A.3 in Appendix A
• e.g., with a sample size of n = 30, 95% confidence inter-
vals are constructed using t = 2.045, instead of the value
of z = 1.96 used above for the normal distribution
• For the commuting data (textbook):

 14.43   14.43  
pr   21.93  2.045    21.93  2.045   0.95
 30   30  
Confidence intervals & t-distribution

 14.43   14.43  
pr   21.93  2.045    21.93  2.045   0.95
 30   30  

• We are 95% sure that the true mean is within the in-
terval (16.54, 27.32)
• More precisely, 95% of confidence intervals con-
structed from samples in this way will contain the true
mean
Hypothesis Testing

• One-sample tests
– One-sample tests for the mean

– One-sample tests for proportions

• Two-sample tests
– Two-sample tests for the mean

179
Hypothesis Testing
• Hypothesis testing
 Null hypothesis
• Purpose
 Test the viability
• Null hypothesis
 Population parameter
 Reverse of what the experimenter believes
180
Hypothesis Testing: Steps
1. State the null hypothesis, H0

2. State the alternative hypothesis, HA

3. Choose a, our significance level


4. Select a statistical test, and find the observed test
statistic
5. Find the critical value of the test statistic
6. Compare the observed test statistic with the critical
value, and decide to accept or reject H0
181
Hypothesis Testing – Step 1

1. State the null hypothesis (H0)

– H0: μ = μ0

– H0: μ - μ0 = 0

182
Hypothesis Testing – Step 2

2. State the alternative hypothesis


– HA: μ # μ0  two-sided (two-tailed)

or
– HA : μ > μ0 upper-tailed

 one-sided (one-tailed)
– HA : μ < μ0 lower-tailed
183
Hypothesis Testing – Step 3

3. Choose α, our significance level


– It really depends on what we are testing

– α = 0.05

– α = 0.01

– Type I error

184
Hypothesis Testing - Errors

• Type I Error - α error, occurs when we reject


the null hypothesis when we should accept it
• Type II Error - β error, occurs when we ac-
cept the null hypothesis when we should re-
ject it

185
Hypothesis Testing - Errors

H0 is true H0 is false

Accept H0 Correct decision Type II Error (β)


(1-α)

Reject H0 Type I Error (α) Correct decision


(1-β)

186
Hypothesis Testing – Step 4
4. Select a statistical test, and find the test statistic

q - q0
Test statistic = Std. error

x  x 
z  z
/ n s/ n
187
Hypothesis Testing – Step 4
4. Select a statistical test, and find the test statis-
tic
q - q0
Test statistic = Std. error

x  x 
t  t
/ n s/ n
188
Hypothesis Testing – Step 5

5. Find the critical value of the test statistic

– Standard normal table

– Student’s t distribution table

– Two-sided vs. one-sided

189
Two-sided tests  Zα/2

190
One-sided tests  Zα

191
Hypothesis Testing – Step 6
6. Compare the observed test statistic with
the critical value

| Ztest | > | Zcrit | Þ HA -Zcrit H0 Zcrit


| Ztest | £ | Zcrit | Þ H0
HA HA

192
Hypothesis Testing – Step 6

6. Compare the observed test statistic with


the critical value

| Ztest | > | 1.96 | Þ HA -1.96 H0 1.96


| Ztest | £ | 1.96 | Þ H0
HA HA

193
Hypothesis Testing – Step 6

6. Compare the observed test statistic with


the critical value

Ztest > Zcrit Þ HA Zcrit


Ztest £ Zcrit Þ H0 H0
HA

194
Hypothesis Testing – Step 6

6. Compare the observed test statistic with


the critical value

Ztest > 1.645 Þ HA 1.645


Ztest £ 1.645 Þ H0 H0

HA

195
p-value
• p-value is the probability of getting a value of the test
statistic as extreme as or more extreme than that observed
by chance alone, if the null hypothesis H0, is true.

• It is the probability of wrongly rejecting the null hypothe-


sis if it is in fact true
• It is equal to the significance level of the test for which
we would only just reject the null hypothesis

196
p-value

• p-value vs. significance level

• Small p-values  the null hypothesis is unlikely to be


true

• The smaller it is, the more convincing is the rejection of


the null hypothesis

197
One-Sample z-Tests

198
Next topics:
• Analysis of variance (ANOVA)
• Correlation
• Regression

199

You might also like