1|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
1. With diagram explain the structure of the human eye?
Figure below shows a simplified horizontal cross section of the human eye.
The eye is nearly a sphere, with an average diameter of approximately 20 mm. Three membranes
enclose the eye: the cornea and sclera outer cover; the choroid; and the retina.
The cornea is a tough, transparent tissue that covers the anterior surface of the eye.
Continuous with the cornea, the sclera is an opaque membrane that encloses the remainder of the
optic globe.
The choroid lies directly below the sclera. This membrane contains a network of blood vessels
that serve as the major source of nutrition to the eye. Even superficial injury to the choroid, often
not deemed serious, can lead to severe eye damage as a result of inflammation that restricts
blood flow. The choroid is heavily pigmented and hence helps to reduce the amount of
extraneous light entering the eye and the backscatter within the optic globe. At its anterior
extreme, the choroid is divided into the ciliary body and the iris.
The iris contracts or expands to control the amount of light that enters the eye. The central
opening of the iris (the pupil) varies in diameter from approximately 2 to 8 mm. The front of the
iris contains the visible pigment of the eye, while the back contains a black pigment.
The lens is made up of concentric layers of fibrous cells and is suspended by fibers that attach to
the ciliary body. It contains 60 to 70% water, about 6% fat, and more protein than any other
tissue in the eye. The lens is colored by a slightly yellow pigmentation that increases with age. In
extreme cases, excessive clouding of the lens, commonly referred to as cataracts, can lead to
poor color discrimination and loss of clear vision.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
2|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
The innermost membrane of the eye is the retina, which lines the inside of the wall’s entire
posterior portion. When the eye is properly focused, light from an object outside the eye is
imaged on the retina. Pattern vision is possible by the distribution of discrete light receptors over
the surface of the retina. There are two classes of receptors: cones and rods.
o The cones in each eye number between 6 and 7 million. They are located primarily in the
central portion of the retina, called the fovea, and are highly sensitive to color. Humans
can resolve fine details with these cones largely because each one is connected to its own
nerve end. Muscles controlling the eye rotate the eyeball until the image of an object of
interest falls on the fovea. Cone vision is called photopic or bright-light vision.
o The number of rods is much larger. Some 75 to 150 million are distributed over the
retinal surface. The larger area of distribution and the fact that several rods are connected
to a single nerve end reduce the amount of details discernible by these receptors. Rods
serve to give a general, overall picture of the field of view. They are not involved in color
vision and are sensitive to low levels of illumination. For example, objects that appear
brightly colored in day-light when seen by moonlight appear as colorless because only
the rods are stimulated. This phenomenon is known as scotopic or dim-light vision.
Blind spot is the region in which no receptors are present.
2. Explain the difference between scotopic and photopic vision.
Refer cones and rodes from question number 1.
3. Explain how an image is formed inside a human eye.
In an ordinary photographic camera, the lens has a fixed focal length, and focusing at various distances
is achieved by varying the distance between the lens and the imaging plane, where the film (or imaging
chip in the case of a digital camera) is located. In the human eye, the converse is true; the distance
between the lens and the imaging region (the retina) is fixed, and the focal length needed to achieve
proper focus is obtained by varying the shape of the lens. The fibers in the ciliary body accomplish this,
flattening or thickening the lens for distant or near objects, respectively. The distance between the center
of the lens and the retina along the visual axis is approximately 17 mm. The range of focal lengths is
approximately 14 mm to 17 mm, the latter taking place when the eye is relaxed and focused at distances
greater than about 3 m.
The figure illustrates how to obtain the dimensions of an image formed on the retina. For example,
suppose that a person is looking at a tree 15 m high at a distance of 100 m. Letting h denote the height of
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
3|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
that object in the retinal image, from the geometry we obtain 15/100 = h/17 or h = 2.55 mm. The
inverted retinal image is focused primarily on the region of the fovea. Perception then takes place by the
relative excitation of light receptors, which transform radiant energy into electrical impulses that
ultimately are decoded by the brain.
4. What do you mean by brightness adaptation?
The range of light intensity levels to which the human visual system can adapt is enormous- on the order
of 1010- from the scotopic threshold to the glare limit. The least amount of light that can be perceived by
the human eye is called the scotopic threshold. Similarly, the maximum amount of light that can be
perceived is called glare limit. The human visual system can adapt to this large dynamic range.
However, the human eye cannon operate over this entire range simultaneously. Instead, the eye adapts to
the average brightness of the region on which it is focused. This range is called the instantaneous range.
This kind of calibration or adjustment is called brightness adaptation.
This adaptation is done by dilation and contraction of the eye by the retina, by controlling
the amount of light falling on the retina. The iris acts as a light regulator. In dark conditions, it dilates
and allows more light. In very bright conditions, it contracts and admits less light. Thus, changes in the
size of the iris provide a fine control over the eye's adaptation by changing its sensitivity level. In bright
light, the retina is less sensitive and in dim light, it is more sensitive. Hence, the overall effective
response of the eye is maintained. There are three types of brightness adaptation-general, local, and
lateral.
General brightness adaptation: It refers to the average brightness of the scene. For example,
when we enter a dark room, we cannot see anything initially. Then slowly the eye increases its
sensitivity, and we can see things in the dark. Brightness is also influenced by the contrast of the
light.
Local brightness adaptation: When we view a scene, we view the objects one by one. The eye
quickly adjusts its brightness adaptation before moving to the next object. This transition is so
rapid that we often fail to take notice of it. If the transition is not smooth, it leads to the ‘after
image’ effect. This effect is felt when we see a darker image after gazing at a bright object for a
long period.
Lateral brightness adaptation: Sensitivity changes in the local areas of an object are associated
with similar changes in the background. This leads to the concept of simultaneous contrast.
5. What do you mean by subjective brightness?
Subjective brightness (intensity as perceived by the human visual system) is a logarithmic function of
the light intensity incident on the eye.
6. Explain simultaneous contrast.
Simultaneous contrast is related to the fact that a region’s perceived brightness does not depend simply
on its intensity. Vision is also influenced by the background.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
4|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
In the above figure, all the center squares have exactly the same intensity. However they appear to the
eye to become brighter as the background gets darker. A more familiar example is a piece of paper that
seems white when lying on a desk, but can appear totally black when used to shield the eyes while
looking directly at a bright sky. This is due to the fact that human perception is sensitive to the luminous
contrast with respect to the background rather than the absolute luminance value.
7. What is Weber Ratio?
Let the foreground object intensity be Io and the intensity of the background be Ib. If
the difference between the intensity of the foreground and that of the background is large, such as a
white object in a black background and vice versa, the human visual system has
no problem in identifying the foreground. However, if both are same, it cannot recognize
the foreground. Let the foreground intensity be changed slowly. The minimum amount
of change in intensity that produces a noticeable variation in the sensory experience of
perception is called the difference threshold, which is also known as the ‘just noticeable
difference’. This is the difference between the intensities of the foreground and background objects. For
example, if the intensity of the background is 100 and that of the foreground is l20, the difference
threshold is 20.
Let ΔI= Io - Ib. If ΔI is not large enough, the subject says “no”, indicating no perceivable change. As ΔI
gets stronger, the subject may give a positive response of “yes”, indicating a perceived change. Finally,
when ΔI is strong enough, the subject will give a response of “yes” all the time. The quantity ΔI c/Ib,
where ΔIc is the increment of illumination discriminable 50% of the time with background illumination
Ib, is called the Weber ratio. A small value of ΔIc/Ib means that a small percentage change in intensity is
discriminable. This represents “good” brightness discrimination. Conversely, a large value of ΔI c/Ib
means that a large percentage change in intensity is required. This represents “poor” brightness
discrimination.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
5|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
The brightness discrimination is poor (the Weber ratio is large) at low levels of illumination, and it
improves significantly (the Weber ratio decreases) as background illumination increases. At low levels
of illumination, vision is carried out by the rods, whereas at high levels (showing better discrimination)
vision is the function of cones.
8. What do you mean by brightness discrimination?
Explain from question number 7.
9. Explain the image formation model with respect to the terms illumination and reflectance.
Image is a two-dimensional function of the form f(x, y). When an image is generated from a physical
process, its intensity values are proportional to energy radiated by a physical source (e.g.,
electromagnetic waves). As a consequence, f(x, y) must be nonzero and finite; that is,
0 < f(x,y) < ∞
The function f(x, y) may be characterized by two components:
1. Illumination: The amount of source illumination incident on the scene being viewed denoted by
i(x,y), and
2. Reflectance: The amount of illumination reflected by the objects in the scene denoted by r(x,y).
The two functions combine as a product to form f(x, y):
f(x, y) = i(x, y) . r(x, y)
where
0<i(x, y)<
and
0<r(x, y)<1.
The above equation indicates that reflectance is bounded by 0 (total absorption) and 1 (total reflectance).
The nature of i(x, y) is determined by the illumination source, and r(x, y) is determined by the
characteristics of the imaged objects.
10. Explain the different types of images.
The colors that humans perceive in an object are determined by the nature of the light reflected from the
object. A body that reflects light relatively balanced in all visible wavelengths appears white to the
observer. However, a body that favors reflectance in a limited range of the visible spectrum exhibits
some shades of color.
Light that is void of color is called monochromatic (or achromatic) light. The only attribute of
monochromatic light is its intensity or amount. Because the intensity of monochromatic light is
perceived to vary from black to grays and finally to white, the term gray level is used commonly
to denote monochromatic intensity. The range of measured values of monochromatic light from
black to white is usually called the gray scale, and monochromatic images are frequently
referred to as gray-scale images. The intensity range is denoted as [0, Lmin ] where 0 signifies
black, Lmin signifies white and all other intensities in between are shades of gray varying from
black to white.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
6|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
Image with only two possible intensity level is known as binary image or black and white
image. Here intensity level is denoted by [0, 1] where 0 is black and 1 is white.
Chromatic(color) light generates chromatic or colour images. In addition to frequency, three
basic quantities are used to describe the quality of a chromatic light source: radiance, luminance,
and brightness. Radiance is the total amount of energy that flows from the light source, and it is
usually measured in watts(W). Luminance, measured in lumens(lm), gives a measure of the
amount of energy an observer perceives from a light source. For example, light emitted from a
source operating in the far infrared region of the spectrum could have significant energy
(radiance), but an observer would hardly perceive it; its luminance would be almost zero.
Finally, brightness is a subjective descriptor of light perception that is practically impossible to
measure. It embodies the achromatic notion of intensity and is one of the key factors in
describing color sensation.
11. Explain image sensing and acquisition in detail.
The images are generated by the combination of an “illumination” source and the reflection or
absorption of energy from that source by the elements of the “scene” being imaged. Depending on the
nature of the source, illumination energy is reflected from, or transmitted through, objects. An example
in the first category is light reflected from a planar surface. An example in the second category is when
X-rays pass through a patient’s body for the purpose of generating a diagnostic X-ray film. In some
applications, the reflected or transmitted energy is focused onto a photoconverter (e.g. a phosphor
screen), which converts the energy into visible light. Electron microscopy and some applications of
gamma imaging use this approach.
Three principal sensor arrangements used to transform illumination energy into digital images. The
incoming energy is transformed into a voltage by the combination of input electrical power and sensor
material that is responsive to the particular type of energy being detected. The output voltage waveform
is the response of the sensor(s), and a digital quantity is obtained from each sensor by digitizing its
response.
Image acquisition using a single sensor:
Perhaps the most familiar sensor of this type is the photodiode, which is constructed of silicon materials
and whose output voltage waveform is proportional to light. The use of a filter in front of a sensor
improves selectivity. For example, a green (pass) filter in front of a light sensor favors light in the green
band of the color spectrum. As a consequence, the sensor output will be stronger for green light than for
other components in the visible spectrum. In order to generate a 2-D image using a single sensor, there
has to be relative displacements in both the x- and y-directions between the sensor and the area to be
imaged.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
7|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
Image acquisition using sensor strips:
A geometry that is used much more frequently than single sensors consists of an in-line arrangement of
sensors in the form of a sensor strip. The strip provides imaging elements in one direction. Motion
perpendicular to the strip provides imaging in the other direction. This is the type of arrangement used in
most flat bed scanners.
Image acquisition using sensor arrays:
Numerous electromagnetic and some ultrasonic sensing devices frequently are arranged in an array
format. This is also the predominant arrangement found in digital cameras. A typical sensor for these
cameras is a CCD array. CCD sensors are used widely in digital cameras and other light sensing
instruments. The response of each sensor is proportional to the integral of the light energy projected onto
the surface of the sensor, a property that is used in astronomical and other applications requiring low
noise images. Noise reduction is achieved by letting the sensor integrate the input light signal over
minutes or even hours. Because the sensor array is two-dimensional, its key advantage is that a complete
image can be obtained by focusing the energy pattern onto the surface of the array. Motion obviously is
not necessary.
12. Explain image sampling and quantization in detail.
Though there are numerous ways to acquire images, but our objective is to generate digital images from
sensed data. The output of most sensors is a continuous voltage waveform whose amplitude and spatial
behavior are related to the physical phenomenon being sensed. To create a digital image, we need to
convert the continuous sensed data into digital form. This involves two processes: sampling and
quantization. An image may be continuous with respect to the x- and y-coordinates, and also in
amplitude. To convert it to digital form, we have to discretize the function in both coordinates and in
amplitude. Digitizing the coordinate values is called sampling. Digitizing the amplitude value is called
quantization.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
8|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
(a) Continuous image. (b) A scan line from A to B in the continuous image, (c) Sampling and quantization.
(d) Digital scan line.
Figure (a) shows a continuous image f that we want to convert to digital form. The one-dimensional
function in figure (b) is a plot of amplitude (intensity level) values of the continuous image along the
line segment AB. The random variations are due to image noise. To sample this function, we take
equally spaced samples along line AB, as shown in figure (c). The samples are shown as small white
squares superimposed on the function. The set of these discrete locations gives the sampled function.
However, the values of the samples still span (vertically) a continuous range of intensity values. In order
to form a digital function, the intensity values also must be converted (quantized) into discrete
quantities. The right side of figure (c) shows the intensity scale divided into eight discrete intervals,
ranging from black to white. The continuous intensity levels are quantized by assigning one of the eight
values to each sample. The assignment is made depending on the vertical proximity of a sample to a
vertical tick mark. The digital samples resulting from both sampling and quantization are shown in
figure (d). Starting at the top of the image and carrying out this procedure line by line produces a two-
dimensional digital image. It is implied that, in addition to the number of discrete levels used, the
accuracy achieved in quantization is highly dependent on the noise content of the sampled signal.
13. Define dynamic range and contrast of the image.
Sometimes, the range of values spanned by the gray scale is referred to informally as the dynamic range.
This is a term used in different ways in different fields. Here, we define the dynamic range of an
imaging system to be the ratio of the maximum measurable intensity to the minimum detectable
intensity level in the system. As a rule, the upper limit is determined by saturation and the lower limit by
noise. Basically, dynamic range establishes the lowest and highest intensity levels that a system can
represent and, consequently, that an image can have.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
9|Page Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
Closely associated with this concept is image contrast, which we define as the difference in intensity
between the highest and lowest intensity levels in an image. When an appreciable number of pixels in an
image have a high dynamic range, we can expect the image to have high contrast. Conversely, an image
with low dynamic range typically has a dull, washed-out gray look.
14. What is the size of an image?
The size of the image is the number of bits required to store a digitized image, which is given by
b=MxNxk
where M and N are the number of pixels along the X and the Y co-ordinates of the image, and L =2 k are
the number of quantization levels, which gives k as the number of bits needed to represent each pixel
intensity.
15. Define spatial and intensity resolution.
Intuitively, spatial resolution is a measure of the smallest discernible detail in an image. Quantitatively,
spatial resolution can be stated in a number of ways, with line pairs per unit distance, and dots (pixels)
per unit distance being among the most common measures.
A widely used definition of image resolution is the largest number of discernible line pairs per unit
distance (e.g., 100 line pairs per mm). Dots per unit distance is a measure of image resolution used
commonly in the printing and publishing industry. Measures of spatial resolution must be stated with
respect to spatial units.
Intensity resolution similarly refers to the smallest discernible change in intensity level. Based on
hardware considerations, the number of intensity levels usually is an integer power of two, The most
common number is 8 bits, with 16 bits being used in some applications in which enhancement of
specific intensity ranges is necessary. Unlike spatial resolution, which must be based on a per unit of
distance basis to be meaningful, it is common practice to refer to the number of bits used to quantize
intensity as the intensity resolution.
16. What do you mean by false contouring?
The effect caused by the use of an insufficient number of intensity levels in smooth areas of a digital
image, is called false contouring, so called because the ridges resemble topographic contours in a map.
False contouring generally is quite visible in images displayed using 16 or less uniformly spaced
intensity levels.
17. What do you mean by isopreference curve?
The woman’s face is representative of an image with relatively little detail; the picture of the cameraman
contains an intermediate amount of detail; and the crowd picture contains, by comparison, a large
amount of detail.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
10 | P a g e Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
Sets of these three types of images were generated by varying N and k, and observers were then asked to
rank them according to their subjective quality. Results were summarized in the form of so-called
isopreference curves in the Nk-plane. Each point in the Nk-plane represents an image having values of
N and k equal to the coordinates of that point. Points lying on an isopreference curve correspond to
images of equal subjective quality. It was found in the course of the experiments that the isopreference
curves tended to shift right and upward, but their shapes in each of the three image are similar to those
given in figure below. This is not unexpected, because a shift up and right in the curves simply means
larger values for N and k, which implies better picture quality.
The key point of interest in the context of the present discussion is that isopreference curves tend to
become more vertical as the detail in the image increases. This result suggests that for images with a
large amount of detail only a few intensity levels may be needed. For example, the isopreference curve
in corresponding to the crowd is nearly vertical. This indicates that, for a fixed value of N, the perceived
quality for this type of image is nearly independent of the number of intensity levels used (for the range
of intensity levels shown in figure). It is of interest also to note that perceived quality in the other two
image categories remained the same in some intervals in which the number of samples was increased,
but the number of intensity levels actually decreased. The most likely reason for this result is that a
decrease in k tends to increase the apparent contrast, a visual effect that humans often perceive as
improved quality in an image.
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
11 | P a g e Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
18. Explain image interpolation.
Interpolation is a basic tool used extensively in tasks such as zooming, shrinking, rotating, and
geometric corrections. Image resizing (shrinking and zooming) are basically image resampling methods.
Fundamentally, interpolation is the process of using known data to estimate values at unknown
locations.
We begin the discussion of this topic with a simple example. Suppose that an image of size 500 * 500
pixels has to be enlarged 1.5 times to 750 * 750 pixels. A simple way to visualize zooming is to create
an imaginary 750 * 750 grid with the same pixel spacing as the original, and then shrink it so that it fits
exactly over the original image. Obviously, the pixel spacing in the shrunken 750 * 750 grid will be less
than the pixel spacing in the original image. To perform intensity-level assignment for any point in the
overlay, we look for its closest pixel in the original image and assign the intensity of that pixel to the
new pixel in the 750 * 750 grid. When we are finished assigning intensities to all the points in the
overlay grid, we expand it to the original specified size to obtain the zoomed image.
The method just discussed is called nearest neighbor interpolation because it assigns to each new
location the intensity of its nearest neighbor in the original image. This approach is simple but it has the
tendency to produce undesirable artifacts, such as severe distortion of straight edges. For this reason, it
is used infrequently in practice.
A more suitable approach is bilinear interpolation, in which we use the four nearest neighbors to
estimate the intensity at a given location. Let (x, y) denote the coordinates of the location to which we
want to assign an intensity value (think of it as a point of the grid described previously), and let v(x, y)
denote that intensity value. For bilinear interpolation, the assigned value is obtained using the equation
v(x, y) = ax + by + cxy + d
where the four coefficients are determined from the four equations in four unknowns that can be written
using the four nearest neighbors of point (x, y). Bilinear interpolation gives much better results than
nearest neighbor interpolation, with a modest increase in computational burden.
19. What do you mean by neighbors of a pixel?
A pixel p at coordinates (x, y) has four horizontal and vertical neighbors whose coordinates are given by
(x + 1, y), (x - 1, y), (x, y + 1), (x, y - 1)
This set of pixels, called the 4-neighbors of p, is denoted by N4(p). Each pixel is a unit distance from (x,
y).
The four diagonal neighbors of p have coordinates
(x + 1, y + 1), (x + 1, y - 1), (x - 1, y + 1), (x - 1, y - 1)
and are denoted by ND(p). These points, together with the 4-neighbors, are called the 8-neighbors of p,
denoted by N8(p). Some of the neighbor locations in ND(p) and N8(p) will fall outside the image if (x,
y) is on the border of the image.
20. What do you mean by adjacency of pixels?
Let V be the set of intensity values used to define adjacency. In a binary image, V = {1} if we are
referring to adjacency of pixels with value 1. We consider three types of adjacency:
a. 4-adjacency. Two pixels p and q with values from V are 4-adjacent if q is in the set N4(p).
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
12 | P a g e Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
b. 8-adjacency. Two pixels p and q with values from V are 8-adjacent if q is in the set N8(p).
c. m-adjacency (mixed adjacency). Two pixels p and q with values from V are m-adjacent if
i) q is in N4(p), or
ii) q is in ND(p) and the set N4(p) ∩ N4(q) has no pixels whose values are from V.
Mixed adjacency is a modification of 8-adjacency. It is introduced to eliminate the ambiguities that often
arise when 8-adjacency is used. For example, consider the pixel arrangement shown below for V = {1}.
The three pixels at the top of show multiple (ambiguous) 8-adjacency, as indicated by the dashed lines.
This ambiguity is removed by using m-adjacency.
21. What is a digital path?
A (digital) path (or curve) from pixel p with coordinates (x, y) to pixel q with coordinates (s, t) is a
sequence of distinct pixels with coordinates
(x0, y0), (x1, y1),......, (xn, yn)
where (x0, y0) = (x, y), (xn, yn) = (s, t), and pixels (xi, yi) and (xi-1, yi-1) are adjacent for 1 ≤ i ≤ n. In this
case, n is the length of the path. If (x0, y0) = (xn, yn), the path is a closed path. We can define 4-, 8-, or m-
paths depending on the type of adjacency specified.
22. Explain connectivity between pixels.
Let S represent a subset of pixels in an image. Two pixels p and q are said to be connected in S if there
exists a path between them consisting entirely of pixels in S. For any pixel p in S, the set of pixels that
are connected to it in S is called a connected component of S. If it only has one connected component,
then set S is called a connected set.
23. Explain what do you mean by region.
Let R be a subset of pixels in an image. We call R a region of the image if R is a connected set. Two
regions, Ri and Rj are said to be adjacent if their union forms a connected set. Regions that are not
adjacent are said to be disjoint.
The boundary (also called the border or contour) of a region R is the set of points that are adjacent to
points in the complement of R. Said another way, the border of a region is the set of pixels in the region
that have at least one background neighbor. Here again, we must specify the connectivity being used to
define adjacency.
24. What are the distance measures between pixels?
For pixels p, q, and z, with coordinates (x, y), (s, t), and (v, w), respectively, D is a distance function or
metric if
(a) D(p, q) ≥ 0 (D(p, q) = 0 iff p = q),
(b) D(p, q) = D(q, p), and
(c) D(p, z) ≤ D(p, q) + D(q, z).
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE
13 | P a g e Digital Image and Video Processing (PE-EC702B)
MODULE 1: DIGITAL IMAGE PROCESSING SYSTEMS
The Euclidean distance between p and q is defined as
For this distance measure, the pixels having a distance less than or equal to some value r from (x, y) are
the points contained in a disk of radius r centered at (x, y).
The D4 distance (called the city-block distance) between p and q is defined as
In this case, the pixels having a D4 distance from (x, y) less than or equal to some value r form a
diamond centered at (x, y). For example, the pixels with D4 distance ≤ 2 from (x, y) (the center point)
form the following contours of constant distance:
The pixels with are the 4-neighbors of (x,y).
The D8 distance (called the chessboard distance) between p and q is defined as:
In this case, the pixels with distance from (x, y) less than or equal to some value r form a square
centered at (x,y) . For example, the pixels with D8 distance ≤ 2 from (x, y) (the center point) form the
following contours of constant distance:
The pixels with D8 = 1 are the 8-neighbors of (x,y).
PREPARED BY: DWAIPAYAN GHOSH, ASSISTANT PROFESSOR OF ECE