1
Basics of IMAGE CLASSIFICATION
Introduction
The image classification methods are some of the most important techniques of image processing and
analysis. In the context of image classification, an image consists of pixels, each of which has a definite
spectral characteristic and the pixels that represent similar features of the surface of the earth usually
have similar spectral characteristics, based on which they can be categorized into information classes.
The output of the image classification is a thematic map where the information classes represent land
use/land cover themes. According to Lillesand and Keifer (1994), image classification is defined as the
process of categorizing all pixels in an image or raw remotely sensed satellite data to obtain a given
set of labels or land cover themes. Digital image classification utilizes spectral information of spectral
bands of multispectral images. The spectral information is represented by digital numbers or gray level
or tone of the pixels. Therefore, the digital image classification is also known as spectral pattern
recognition.
Land cover refers to the physical and biological cover over the surface of the earth, which includes
vegetation, water, land, snow field, wasteland etc. The term ‘land use’ is rather complicated and it
includes human intervention on the surface of the earth. Typical land use classes include agricultural
land, plantation, urban settlements, rural settlements etc. However, in practice both the terms are
quite amalgamated, and a holistic land use/land cover mapping would include identification of both
natural and manmade covers of the surface of the earth for the same thematic map.
Image classification techniques can be manual or computer based. Manual image classification can be
done either over hard copy images or in image processing softwares, by identifying the land use/land
cover classes with eyes. The identification of land use/land cover classes in manual method is based
on different image elements such as tone, texture, shape, pattern, shadow etc.
In computer based image classification methods, the spectral classes are identified by the algorithms
of image processing softwares and they are essentially tone based. Common classification methods
are grouped into two broad subdivisions, i.e., supervised and unsupervised image classifications.
In the unsupervised image classification, the spectral classes are clustered first based solely on spectral
characteristics of the pixels and then matched by the analyst with the real world land use/land cover
classes. The analyst usually specifies how many spectral classes or clusters are to be identified in the
beginning of the procedure. In some instances, a particular land use/land cover class may represent
more than one spectral class, identified by the software. In such cases those spectral classes have to
be merged into one spectral class.
2
In the supervised classification the analyst first identifies the representative samples of information
classes with definite spectral characteristics which are also known as training classes. The
identification of training classes requires prior knowledge about the land use/land cover pattern
within the geographical area of the area of interest. Thus, the analyst ‘supervises’ the procedure of
this particular image classification. After proper identification of the training classes, the image
classification software categorize all the pixels of the image into spectrally distinct classes which
resemble to spectral characteristics of the training classes.
The Fig. 1. shows a raw image (a) and the classified image (b) with the index of land use/land cover
classes.
(a) Satellite image (b) Classified image
Fig.1. An example of land use/land cover classification (Zhou et al, 2014)
Unsupervised classification
Unsupervised image classification is a method in which the image interpreting software separates a
large number of unknown pixels in an image based on their reflectance values into classes or clusters
with no direction from the analyst (Tou, Gonzalez 1974). This method is purely based on spectral
characteristics of pixel and no prior knowledge of the characteristics of the land use/land cover classes
of the study are is required. There are two clustering algorithms utilized for unsupervised
classification, which rely essentially on spectral characteristics of pixels. These two algorithms are
known as K-means and ISODATA (Iterative self-organizing data analysis technique).
Both of these algorithms are iterative procedures. At first both of the methods first assign arbitrary
cluster properties as initial values which are evenly distributed in the data space. In the second step,
all the pixels of the image are classified to the closest cluster. In the third step, the mean cluster
properties are selected based on pixel properties of each cluster. After this the second and third steps
are iterated again and again until the change in clustering of pixels is minor. This is measured either
from (i) the distances of the mean cluster properties from the counter part of the previous iteration
or by (ii) the difference in percentages of total pixels in each cluster between iterations.
3
The ISODATA algorithm adds some refinement by splitting and merging the clusters. Merging of
clusters is based on two thresholds which are defined by minimum number of pixels in each cluster
and minimum distance between the mean values of the clusters. A cluster is split into two when the
standard deviation of the cluster exceeds a predefined value or total number of member pixels in that
particular cluster is twice the threshold for the minimum number of pixels.
The K-means algorithm differs from the ISODATA in the fact that the number of clusters remains the
same throughout the iteration.
The output clusters of the unsupervised classification are then visually correlated with the real world
land use/land cover classes by ground truthing. Sometimes a single land use/land cover class may be
represented by more than one spectral cluster in the output thematic map. In that case, these spectral
clusters are grouped into a single one.
The chief advantages of the unsupervised classification are:
(i) No detailed prior knowledge about the area of interest is required.
(ii) Clustering is done by computer based algorithms, thus, chances of human error are less.
(iii) In case of supervised classification, classification is based on identification of training classes.
Therefore, minor land use/land cover classes may be left unrecognized. In unsupervised image
classification distinct spectral classes, be it minor, never left unrecognized.
Despite these advantages of unsupervised classification over supervised classification, it also bears
some disadvantages, notably:
(i) Unsupervised image classification algorithms may recognize distinct spectral classes within
the same information class, because of spatial variation in spectral characters of the same
information class caused by illumination differences.
(ii) The relation between spectral classes and information classes is not constant and change over
space (location) and time (seasonal variation). Therefore, this relationship varies from image
to image.
(iii) Identification of spectral classes in the output of the unsupervised classification may be very
time consuming.
The major steps involved in the unsupervised image classification method shown in the Fig. 1.
Ground truthing,
cluster Accuracy
Image Clustering identification and assessment
grouping
Fig. 2. Major steps involved in unsupervised image classification.
4
6.4 Supervised classification
In supervised image classification, the user first identifies representative samples of the land use/land
cover classes over the digital image. These representative classes are known as ‘training classes’. To
select the training classes over the image, ground truth data with positional accuracy is a prime
requisite in this classification. The image processing software, then, does a statistical characterization
of spectral reflectance of the each training class over the spectral bands of the digital image. This stage
is known as ‘signature analysis’ and the statistical characterization of the reflectance values of the
training classes involves determination of mean, variance and covariance over the spectral bands of
the digital image. The image is then classified by examining the spectral reflectance of each pixel and
aligning the pixel to the training class to which it resembles the most. Since, user’s supervision is
necessary in the initial stage of the classification scheme in identifying the training classes, this
classification is known as ‘supervised classification’.
Sometimes, spectral reflectance some pixels of the digital image may not resemble to those of the
training classes. The analyst can leave an option not to classify such pixels and these pixels remain as
unclassified pixels in the output image.
The classification scheme and steps of the supervised image classification is shown in the Fig. 3.
Fig. 3. Steps in supervised classification
5
There are three statistical methods used for statistical characterization of the spectral reflectance of
the training classes. These methods are minimum distance, parallipiped and maximum likelihood
methods.
Minimum distance method
This is the simplest statistical method of supervised image classification. The minimum distance
method determines the mean DN value (spectral reflectance or pixel brightness) of each training class
over each spectral band of the digital image. In the second step, it assigns the unclassified pixels of
the image to the training class for which it has the minimum Euclidean distance in terms of spectral
reflectance. Thus all pixel of the image are assigned to the nearest training classes. The analyst can
designate a minimum distance threshold for the classification. Pixels which fall beyond the threshold
minimum distance to the training classes would be left as unclassified. This classification method does
not require high processing power of the computer and was very popular few decades back when
computers were not very efficient.
Fig. 4. Plot of digital numbers of two image bands for minimum distance classification
This method is still be a good option when analyst requires classifying a large and high resolution
digital image within a short time.
Parallelepiped method
This parallelepiped classifier is little more complex than the minimum distance classifier. In this
classification the unclassified pixels of the image are assigned to a training class, when their brightness
values fall within a range of the training mean brightness value. A standard deviation from the mean
selects the minimum and maximum value of the class range of each training class. However,
sometimes the parallelepiped classifier creates training class overlap, which causes misclassification
of the some pixels of the image.
6
Maximum likelihood method
It is the most modern and popular classification algorithm of supervised classification, but
computationally complex. The classifier first determines the variance and covariance about the mean
of the training classes. Based on variance and covariance, the classifier, then, determines the Gaussian
probability of the each unknown pixel and assigns it to the class for which it has the highest probability.
Fig. 5. Plot of two bands of the image for Maximum likelihood classification
Accuracy assessment
The accuracy of a classified image is usually evaluated comparing it with ground truth data. The
accuracy of a classification is qualitatively calculated by constructing an ‘error matrix’. The matrix
calculates errors due to omission (exclusion error) and commission (inclusion error) and returns an
overall accuracy. The error matrix tabulates the number of pixels found in a given class (Fig. 6). The
rows represent the pixels classified by the classification algorithm and the columns represent the
pixels in the reference or ground truth data. The omission error calculates the probability of a pixel
being accurately classified in comparison to a reference, while the commission error calculate the
probability of a pixel for accurately representing the class for which it has been assigned by the
classification algorithm. The overall accuracy is calculated by the ration between correctly classified
pixels to the total number of tested (by ground truthing) pixels.
LULC class Ground truth data (tested pixels)
Forest Water Agricultural Rural Urban Row
body land settlement settlement total
7
Forest 70 5 0 13 0 88
Water body 3 55 0 0 0 58
Agricultural 0 0 99 0 0 99
land
Rural 0 0 4 37 0 41
settlement
Urban 0 0 0 0 121 121
settlement
Column total 73 60 103 50 121 407
Overall accuracy= 382/407= 93.86%
Producer’s accuracy (omission error) User’s accuracy (commission error)
Forest=70/73=96 − 4% omission error Forest= 70/88=80 − 20% commission error
Water body=55/60= 92 − 8% omission error Water body=55/88=95 − 5% commission error
Agricultural land=99/103= 96 − 4% omission Agricultural land=99/99=100 − 0% commission
error error
Rural settlement=97/50=74 − 26% omission Rural settlement=37/41=90 − 10% commission
error error
Urban settlement= 121/121=100 − 0% omission Urban settlement=121/121=100 − 0% commission
error error
Fig. 6. Example of an error matrix (modified after Jensen, 1986)