2/10/2025
GEOG 1005
Map use, reading and interpretation
Class 5
Yanjia Cao, PhD
Assistant Professor
Dept of Geography
yanjiac@[Link]
A few more words on colors from
lecture 4
• We mentioned…
• Desaturated, bright colors are perceived as friendly and
professional
• Desaturated dark colors are perceived as serious and
professional
• Saturated colors are perceived as more exciting and
dynamic
1
2/10/2025
An example…
Saturated dark De-saturated dark
An example…
De-saturated
Saturated darkbright De-saturated dark
2
2/10/2025
About wavelength
Plagiarism
• Please check the university info:
[Link]
• An academic theft – ‘stole’ some intellectual
property and present it as his or her own – Write in
your own words!
• It is the responsibility of all students to keep the
originality of the work, not TA’s job to check
similarity
• Consequences include fail of course and dismissal
from university
3
2/10/2025
Today’s lecture
• Data type
• Classification
Thematic map
• A thematic map displays the spatial pattern of a
theme or attribute
• E.g., a temperature map in a newspaper
• Family income across Hong Kong SAR
• The spatial pattern is key here…
• They include choropleth maps (remember those?)
• Different presentations of thematic maps are
possible
• E.g., dot maps, proportional circles
4
2/10/2025
Use of thematic maps
• Provide specific information about a theme at
particular locations
• Focused purpose
• General information about spatial patterns
• Compare patterns on 2 or more maps
Some questions that have to be
answered
• How many categories are appropriate for a given
set of data?
• What approach should be used to partition the
entities into categories?
• How might the results be effectively presented or
displayed?
10
5
2/10/2025
Levels of Measurement
• Classification discussions are often in relation to
scales of measurement
• Nominal, ordinal, interval and ratio
• Nominal data are typically classified according to
taxonomic principles
• Soil types, climate zones, geological periods
11
Ordinal data
• Ordinal data are classified based on different
disciplinary views, e.g., transportation routes
(primary highway, secondary highway, light-duty
road, unimproved road)
Offers an
ordering of
entities
12
6
2/10/2025
Interval and ratio
• Interval and ratio levels are both linked to quantitative
data
• Interval data, e.g., temperature, year (conceptual
measurement, arbitrary zero)
• Ratio data, e.g., distances and area (quantitative
measurement, zero is not arbitrary)
• Inferential statistics (correlation, regression) hold for
these 2 levels of data
• Most census data relates to these 2 levels
• County Population, % of children under the age of 15
• These kinds of data are commonly portrayed on
choropleth maps
13
Common methods of data
classification
• Six common methods of data classification
• Equal intervals
• Quantiles
• Mean-standard deviation
• Maximum breaks
• Natural breaks
• Optimal
14
7
2/10/2025
• Before beginning your classification task, look at
your data and see whether the data have natural or
a meaningful dividing points that can be used to
partition the data
– e.g., zero point – above or below zero
– Or mean percentage value and split above and below
mean
• Then can apply the following methods to those
partitions
15
Equal intervals
• Each class occupies an equal interval along the
number line
• Determine the class interval or width that each
class occupies
• Determine the upper limit of each class
• Compute by repeatedly adding the class interval to the
lowest value in the data
16
8
2/10/2025
• Specify the class limits to be shown in the legend on the map
17
Equal interval
• Best applied if the data ranges are familiar to the
user of the map, such as temperature bands
• About break points…
• Not all datasets show clear break points
• And if they do, the chance that there are enough
for the number of classes planned is quite
uncertain
18
9
2/10/2025
Pros and Cons
• Equal intervals method is very straightforward for
computation
• Also easy for map users to interpret
• Faster map interpretation
• Require No gaps in the data
• Although depending on the values it is not always easy
to reveal what is the class interval
• Can fail to consider how data are distributed along the
number line
• might not have data for every class interval
• One class might contain two clusters of values that beg to be
separated
19
Practice
• Put the following numbers into 4 categories using
equal interval method
1 2 3 3 4 5 7 8 10 11 12 15 19 19 22 23 25
20
10
2/10/2025
Another Practice
• Put the following numbers into 4 categories using
equal interval method
1 2 3 3 4 5 7 15 15 16 17 18 18 25
22
Quantiles
• Data are rank-ordered and equal numbers of
observations are placed in each class
• Total number of observations is divided by number of
classes
• Ideally want the same number of observations in each
class
• To compute class limits
• Specify the lowest and highest values of members in a class
• Or, can compute a class boundary as an average of the
highest value in a class and the lowest value of the next class
and just show the upper limit of each class
24
11
2/10/2025
• A pre-determined number of classes contains an
equal number of observations
• Four-category quantile classifications are known as
quartile
• Widely used
• Five-category classifications are known as quintile
• Well suited to the display of uniformly distributed data
25
Quantiles
• Quantiles are especially useful for ordinal data
(e.g., political boundaries – national, state,
county…)
• If enumeration units are approximately the same
size, each class will have approximately the same
map area
• This is useful for comparing maps
• Another method that is easy to compute and very
common in GIS
• Solves the problem of empty classes
26
12
2/10/2025
Practice
• Put the following numbers into 4 categories using
quantile (quartile) method
1 2 3 3 4 5 7 8 10 11 12 15 19 19 22 23 25
27
Mean-standard deviation
• This method considers how data are distributed along
the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution
• Need some basic understanding of these statistics
29
13
2/10/2025
Mean-standard deviation
• This method considers how data are distributed along
the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution (appropriate sample size etc.)
• Need some basic understanding of these statistics
30
Mean-standard deviation
• This method considers how data are distributed along
the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution (appropriate sample size etc.)
• Need some basic understanding of these statistics
31
14
2/10/2025
Maximum breaks
• This method considers individual data values and
groups those that are similar
• Or avoids grouping those that are dissimilar
• Raw data are ordered from low to high
• Differences between adjacent values are computed and
the largest of these differences serve as class breaks
• Simple to compute, just subtract adjacent values
• One disadvantage is that by focusing on the largest
breaks, you may miss natural clusters of data along the
number line
32
Natural Breaks
• Histograms are examined visually to determine
natural groupings of data
• Goal is to minimize the distance between data
values in the same class and maximize differences
between classes
• Any high values may be put in a class by themselves
• Deciding how to divide up the classes is subjective
and so can vary between analysts
• This method is the default method of data
classification that ArcGIS uses
33
15
2/10/2025
Natural breaks
34
Using statistics
• The normal distribution is the most appropriate for
using the standard deviation classification, whereas a
uniform distribution is the most suitable for using the
equal interval classification
• A uniform distribution is very rare in a real world
• A variable having a perfectly uniform distribution is not very
interesting from the analytical point of view.
• A normal distribution is one of the most common in a
real world
• The majority of variables, however, have irregular
distributions not similar to either the uniform or
normal ones
35
16
2/10/2025
Recently…
• Head/Tail Breaks classification method
• Designed for variable distributions where the data are heavily
distributed toward the tails
• Reported that Natural Breaks does not work so well for these
kinds of distributions
• To determine break values, distribution is divided into 2
parts: the head (above the mean) and the tail (below
the mean).
• The mean value is the 1st break value. Due to the scaling
properties of the data, the head will also be heavy-tailed
distributed, and can once more be divided into a new head
and tail by a new mean value. This will be the 2nd break
value. Procedure is repeated until the data in the head are no
longer heavy-tailed distributed
36
• On next slide we see the distribution of the %
below poverty using 4 of the classification schemes
• Some cartographers believe that more than one
map should be made for a particular data set to
allow the reader to compare them.
37
17
2/10/2025
Equal Interval Quantile Take a look
at the
histogram
Standard Deviation Natural Breaks
38
Exploratory data analysis
• Involves testing whether distribution of data is
uniform or normal, and if there are any extreme
values or outliers
• Also incorporates visualizing and analyzing the
patterns in distribution of data using such tools as
histograms
39
18
2/10/2025
Which classification method to select?
Clusters
• Always take a look at the
distribution of data in a
histogram:
• Are there obvious clusters
within your data?
• Are there large gaps in your
data range that suggest
nice compact data classes?
• If so, pick that number of
classes and place those
class breaks around those
clusters
40
Which classification method to select?
Mean: the average
of all the
observations
Median: the Middle
value in the data set
Mode: the value that
is found mostly in the
/ Normal
data set
• Equal interval works best when data are relatively
evenly distributed between the minimum and
maximum value, and there are no outliers.
• If data is not evenly distributed, quantile may be a
better method to manifest the pattern.
• Standard deviation is best suited for datasets that
conforms to a normal distribution.
• If the mean skews towards left or right, then the The population density of US is generally
map will largely lack the spatial variation very low, only a few (counties in California
or New York) soar to a very high density
41
19
2/10/2025
Practice
• Which classification method does the following example
data use? (2020 population in 24 Counties in Maryland)
42
Summary
• Some classification methods consider how
data are distributed along the number line
• Natural breaks
• Statistical methods
• Some do not consider the distribution along
number line, but are
• Easy to compute and understand
• Equal intervals
• Allow for comparison between maps
• Quantiles
43
20