0% found this document useful (0 votes)
5 views20 pages

Lecture 5 Classification

The document discusses various methods of data classification in geography, focusing on thematic maps and their interpretation. It covers classification techniques such as equal intervals, quantiles, mean-standard deviation, maximum breaks, and natural breaks, highlighting their pros and cons. Additionally, it emphasizes the importance of understanding data distribution when selecting an appropriate classification method.

Uploaded by

Ashley Kwok
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views20 pages

Lecture 5 Classification

The document discusses various methods of data classification in geography, focusing on thematic maps and their interpretation. It covers classification techniques such as equal intervals, quantiles, mean-standard deviation, maximum breaks, and natural breaks, highlighting their pros and cons. Additionally, it emphasizes the importance of understanding data distribution when selecting an appropriate classification method.

Uploaded by

Ashley Kwok
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2/10/2025

GEOG 1005
Map use, reading and interpretation

Class 5

Yanjia Cao, PhD


Assistant Professor
Dept of Geography
yanjiac@[Link]

A few more words on colors from


lecture 4
• We mentioned…
• Desaturated, bright colors are perceived as friendly and
professional
• Desaturated dark colors are perceived as serious and
professional
• Saturated colors are perceived as more exciting and
dynamic

1
2/10/2025

An example…

Saturated dark De-saturated dark

An example…

De-saturated
Saturated darkbright De-saturated dark

2
2/10/2025

About wavelength

Plagiarism

• Please check the university info:


[Link]
• An academic theft – ‘stole’ some intellectual
property and present it as his or her own – Write in
your own words!
• It is the responsibility of all students to keep the
originality of the work, not TA’s job to check
similarity
• Consequences include fail of course and dismissal
from university

3
2/10/2025

Today’s lecture

• Data type
• Classification

Thematic map

• A thematic map displays the spatial pattern of a


theme or attribute
• E.g., a temperature map in a newspaper
• Family income across Hong Kong SAR
• The spatial pattern is key here…
• They include choropleth maps (remember those?)
• Different presentations of thematic maps are
possible
• E.g., dot maps, proportional circles

4
2/10/2025

Use of thematic maps

• Provide specific information about a theme at


particular locations
• Focused purpose
• General information about spatial patterns
• Compare patterns on 2 or more maps

Some questions that have to be


answered
• How many categories are appropriate for a given
set of data?
• What approach should be used to partition the
entities into categories?
• How might the results be effectively presented or
displayed?

10

5
2/10/2025

Levels of Measurement

• Classification discussions are often in relation to


scales of measurement
• Nominal, ordinal, interval and ratio
• Nominal data are typically classified according to
taxonomic principles
• Soil types, climate zones, geological periods

11

Ordinal data

• Ordinal data are classified based on different


disciplinary views, e.g., transportation routes
(primary highway, secondary highway, light-duty
road, unimproved road)

Offers an
ordering of
entities

12

6
2/10/2025

Interval and ratio

• Interval and ratio levels are both linked to quantitative


data
• Interval data, e.g., temperature, year (conceptual
measurement, arbitrary zero)
• Ratio data, e.g., distances and area (quantitative
measurement, zero is not arbitrary)
• Inferential statistics (correlation, regression) hold for
these 2 levels of data
• Most census data relates to these 2 levels
• County Population, % of children under the age of 15
• These kinds of data are commonly portrayed on
choropleth maps

13

Common methods of data


classification
• Six common methods of data classification
• Equal intervals
• Quantiles
• Mean-standard deviation
• Maximum breaks
• Natural breaks
• Optimal

14

7
2/10/2025

• Before beginning your classification task, look at


your data and see whether the data have natural or
a meaningful dividing points that can be used to
partition the data
– e.g., zero point – above or below zero
– Or mean percentage value and split above and below
mean
• Then can apply the following methods to those
partitions

15

Equal intervals

• Each class occupies an equal interval along the


number line
• Determine the class interval or width that each
class occupies

• Determine the upper limit of each class


• Compute by repeatedly adding the class interval to the
lowest value in the data

16

8
2/10/2025

• Specify the class limits to be shown in the legend on the map

17

Equal interval

• Best applied if the data ranges are familiar to the


user of the map, such as temperature bands

• About break points…


• Not all datasets show clear break points
• And if they do, the chance that there are enough
for the number of classes planned is quite
uncertain

18

9
2/10/2025

Pros and Cons

• Equal intervals method is very straightforward for


computation
• Also easy for map users to interpret
• Faster map interpretation
• Require No gaps in the data
• Although depending on the values it is not always easy
to reveal what is the class interval
• Can fail to consider how data are distributed along the
number line
• might not have data for every class interval
• One class might contain two clusters of values that beg to be
separated

19

Practice

• Put the following numbers into 4 categories using


equal interval method
1 2 3 3 4 5 7 8 10 11 12 15 19 19 22 23 25

20

10
2/10/2025

Another Practice

• Put the following numbers into 4 categories using


equal interval method
1 2 3 3 4 5 7 15 15 16 17 18 18 25

22

Quantiles

• Data are rank-ordered and equal numbers of


observations are placed in each class

• Total number of observations is divided by number of


classes
• Ideally want the same number of observations in each
class
• To compute class limits
• Specify the lowest and highest values of members in a class
• Or, can compute a class boundary as an average of the
highest value in a class and the lowest value of the next class
and just show the upper limit of each class

24

11
2/10/2025

• A pre-determined number of classes contains an


equal number of observations
• Four-category quantile classifications are known as
quartile
• Widely used
• Five-category classifications are known as quintile
• Well suited to the display of uniformly distributed data

25

Quantiles

• Quantiles are especially useful for ordinal data


(e.g., political boundaries – national, state,
county…)
• If enumeration units are approximately the same
size, each class will have approximately the same
map area
• This is useful for comparing maps
• Another method that is easy to compute and very
common in GIS
• Solves the problem of empty classes

26

12
2/10/2025

Practice

• Put the following numbers into 4 categories using


quantile (quartile) method
1 2 3 3 4 5 7 8 10 11 12 15 19 19 22 23 25

27

Mean-standard deviation

• This method considers how data are distributed along


the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution
• Need some basic understanding of these statistics

29

13
2/10/2025

Mean-standard deviation

• This method considers how data are distributed along


the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution (appropriate sample size etc.)
• Need some basic understanding of these statistics

30

Mean-standard deviation

• This method considers how data are distributed along


the number line
• Classes are formed by repeatedly adding or subtracting
the standard deviation from the mean of the data
• It works well with data that are normally distributed
(bell-shaped curve)
• Can be problematic if you are working with raw data
that has not been transformed into a normal
distribution (appropriate sample size etc.)
• Need some basic understanding of these statistics

31

14
2/10/2025

Maximum breaks

• This method considers individual data values and


groups those that are similar
• Or avoids grouping those that are dissimilar
• Raw data are ordered from low to high
• Differences between adjacent values are computed and
the largest of these differences serve as class breaks
• Simple to compute, just subtract adjacent values
• One disadvantage is that by focusing on the largest
breaks, you may miss natural clusters of data along the
number line

32

Natural Breaks

• Histograms are examined visually to determine


natural groupings of data
• Goal is to minimize the distance between data
values in the same class and maximize differences
between classes
• Any high values may be put in a class by themselves
• Deciding how to divide up the classes is subjective
and so can vary between analysts
• This method is the default method of data
classification that ArcGIS uses

33

15
2/10/2025

Natural breaks

34

Using statistics

• The normal distribution is the most appropriate for


using the standard deviation classification, whereas a
uniform distribution is the most suitable for using the
equal interval classification
• A uniform distribution is very rare in a real world
• A variable having a perfectly uniform distribution is not very
interesting from the analytical point of view.
• A normal distribution is one of the most common in a
real world
• The majority of variables, however, have irregular
distributions not similar to either the uniform or
normal ones

35

16
2/10/2025

Recently…

• Head/Tail Breaks classification method


• Designed for variable distributions where the data are heavily
distributed toward the tails
• Reported that Natural Breaks does not work so well for these
kinds of distributions
• To determine break values, distribution is divided into 2
parts: the head (above the mean) and the tail (below
the mean).
• The mean value is the 1st break value. Due to the scaling
properties of the data, the head will also be heavy-tailed
distributed, and can once more be divided into a new head
and tail by a new mean value. This will be the 2nd break
value. Procedure is repeated until the data in the head are no
longer heavy-tailed distributed

36

• On next slide we see the distribution of the %


below poverty using 4 of the classification schemes
• Some cartographers believe that more than one
map should be made for a particular data set to
allow the reader to compare them.

37

17
2/10/2025

Equal Interval Quantile Take a look


at the
histogram

Standard Deviation Natural Breaks

38

Exploratory data analysis

• Involves testing whether distribution of data is


uniform or normal, and if there are any extreme
values or outliers
• Also incorporates visualizing and analyzing the
patterns in distribution of data using such tools as
histograms

39

18
2/10/2025

Which classification method to select?


Clusters
• Always take a look at the
distribution of data in a
histogram:
• Are there obvious clusters
within your data?
• Are there large gaps in your
data range that suggest
nice compact data classes?
• If so, pick that number of
classes and place those
class breaks around those
clusters

40

Which classification method to select?


Mean: the average
of all the
observations
Median: the Middle
value in the data set
Mode: the value that
is found mostly in the
/ Normal
data set

• Equal interval works best when data are relatively


evenly distributed between the minimum and
maximum value, and there are no outliers.
• If data is not evenly distributed, quantile may be a
better method to manifest the pattern.
• Standard deviation is best suited for datasets that
conforms to a normal distribution.
• If the mean skews towards left or right, then the The population density of US is generally
map will largely lack the spatial variation very low, only a few (counties in California
or New York) soar to a very high density

41

19
2/10/2025

Practice

• Which classification method does the following example


data use? (2020 population in 24 Counties in Maryland)

42

Summary

• Some classification methods consider how


data are distributed along the number line
• Natural breaks
• Statistical methods
• Some do not consider the distribution along
number line, but are
• Easy to compute and understand
• Equal intervals
• Allow for comparison between maps
• Quantiles

43

20

You might also like