0% found this document useful (0 votes)
10 views28 pages

Retinal Vessel Segmentation via FDM

This paper presents a novel approach for retinal vessel segmentation using a resource-efficient unsupervised technique that combines Fourier decomposition and Gabor transform, achieving high accuracy across multiple datasets. The proposed method outperforms existing techniques in terms of accuracy, sensitivity, and resource usage, making it suitable for various biomedical applications. The study highlights the importance of preprocessing and compares the effectiveness of different segmentation methods, demonstrating the robustness of the proposed approach.

Uploaded by

bvkarthik2711
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views28 pages

Retinal Vessel Segmentation via FDM

This paper presents a novel approach for retinal vessel segmentation using a resource-efficient unsupervised technique that combines Fourier decomposition and Gabor transform, achieving high accuracy across multiple datasets. The proposed method outperforms existing techniques in terms of accuracy, sensitivity, and resource usage, making it suitable for various biomedical applications. The study highlights the importance of preprocessing and compares the effectiveness of different segmentation methods, demonstrating the robustness of the proposed approach.

Uploaded by

bvkarthik2711
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Multimedia Tools and Applications (2024) 83:85871–85898

[Link]

Image decomposition based segmentation of retinal vessels

Anumeha Varma1 · Monika Agrawal1

Received: 21 September 2023 / Revised: 28 June 2024 / Accepted: 28 August 2024 /


Published online: 9 September 2024
© The Author(s), under exclusive licence to Springer Science+Business Media, LLC, part of Springer Nature 2024

Abstract
Retinal vessel segmentation has various applications in the biomedical field. This includes
early disease detection, biometric authentication using retinal scans, classification and others.
Many of these applications rely critically on an accurate and efficient segmentation technique.
In the existing literature, a lot of work has been done to improve the accuracy of the segmen-
tation task, but it relies heavily on the amount of data available for training as well as the
quality of the images captured. Another gap is observed in terms of the resources used in these
heavily trained algorithms. This paper aims to address these gaps by using a resource-efficient
unsupervised technique and also increasing the accuracy of retinal vessel segmentation using
the Fourier decomposition method (FDM) along with the Gabor transform for image signals.
The proposed method has an accuracy of 97.39%, 97.62%, 95.34%, and 96.57% on DRIVE,
STARE, CHASE_DB1, and HRF datasets, respectively. The sensitivities were found to be
88.36%, 88.51%, 90.37%, and 79.07%, respectively. A separate section makes a detailed
comparison of the proposed method with several well-known methods and an analysis of the
efficiency of the proposed method. The proposed method proves to be efficient in terms of
time and resource requirements.

Keywords Image decomposition · Convolutional neural network ·


Denoising autoencoder neural network · Multi-scale wavelet transform ·
Retinal fundus image · Vessel segmentation · Deep learning ·
Two dimensional Fourier decomposition method · Gabor transform

1 Introduction

The retinal vessel network is the only vascular structure that is visible, and even the micro-
vessels can be easily observed through simple methods [1, 2]. Segmentation of these vessels
can help in the early diagnosis of certain diseases, like diabetic retinopathy (DR), glaucoma,
age-related macular degeneration (AMD), cardiovascular diseases, and others, which show
a direct effect on the blood vessels [3, 4]. Other applications include retinal scan based

B Anumeha Varma
crz198693@[Link]
Monika Agrawal
[Link]@[Link]
1 CARE, IIT Delhi, Hauz Khas, New Delhi 110016, Delhi, India

123
85872 Multimedia Tools and Applications (2024) 83:85871–85898

biometric authentication, laser surgeries with the aid of computer vision, and classification
of arteries and veins, which are further informative of other eye diseases like hypertension
and atherosclerosis, which change the artery-to-vein ratio (AVR) in some cases [2, 3, 5–9].
When done manually, the accuracy of the vessel segmentation is subject to skill and human
error [3]. It is also a very tedious task. With algorithmic segmentation methods, the accuracy
faces certain other issues in terms of the varying size of the vessels in the image, low contrast,
and noisy data acquisition [1].
The objective of this work is to efficiently and accurately perform retinal vessel segmen-
tation using the proposed FDM + Gabor transform methods. We also try to apply multiple
well-known methods, which are then compared with the proposed method in terms of accu-
racy, specificity, sensitivity, timing efficiency, thin vessel retention, and many more. The
images taken from various publicly available datasets were passed through several prepro-
cessing steps to enhance them for better segmentation. These enhanced images are trained
with a supervised method using UNet and a self-supervised method using denoising autoen-
coders, while for unsupervised extraction, image decomposition techniques are used. The
proposed method falls into the unsupervised category.

1.1 Related work

A raw fundus image would not give good results with segmentation due to the low, non-
uniform contrast of the vessel structure as compared to the background [10]. To increase the
accuracy of the classification of a pixel as a vessel or non-vessel pixel, the fundus images
need to be processed before being input to the supervised or unsupervised model. Some
commonly used preprocessing techniques include green channel extraction (GE), contrast-
limited adaptive histogram equalisation (CLAHE), and filtering. Soares et al. [11] visualised
the RGB (Red, Green, Blue) channels separately and found the red and the blue channels
were blurry and noisy. The green channel possesses better contrast and signal-to-noise ratio
(SNR) than red and blue channels [10, 11].
Swathi et al. [12] mentioned different types of filters that can be applied to the images
for preprocessing, like the Gaussian filter, which removes noise by averaging neighbouring
pixels. It works as a low-pass filter (LPF) and is used for deblurring images. Budak et al. [13]
applied Gaussian and median filtering for shade correction. The median filter is typically
used to counter salt and pepper noise. It’s good for enhancing edges and smoothening, but
it negatively affects the details in the image. Adaptive median filtering solves this issue. Lei
et al. [14] suggest that a single technique might not be sufficient in most cases, and some
pipelines are created in order to apply multiple techniques in a bundle rather than having
to figure out the best combination of techniques for all datasets. They have proposed and
compared the performance of five classical pipelines, which are a combination of multiple
processing methods like RGB to grayscale conversion, filtering, CLAHE, GE, histogram
equalization, and gamma correction [14]. Zhou et al. [15] applied EMD and gained a 2 dB
enhancement of SNR when compared to the wavelet threshold method and 3 dB compared
to the band-pass filter.
Once the preprocessing is done, the images can be input into the segmentation network,
which should be efficient in terms of time, true acceptance, and false rejection rates. In this
work, several supervised, self-supervised, and unsupervised methods are experimented on
four publicly available datasets: DRIVE, STARE, CHASE_DB1, and HRF. The unsupervised
methods are image decomposition methods like wavelet transform, EMD, VMD, and FDM.
In multiple works, EMD is used for vessel segmentation and analysis of various features from

123
Multimedia Tools and Applications (2024) 83:85871–85898 85873

the retinal images, like age-related macular degeneration (AMD) lesions, pixel variations,
circinate exudates, intensity profile, noise, and retinal texture [16–20]. Lahmiri et al. [21]
have used VMD for retinal image analysis for hemorrhage classification. Here, we used
EMD and VMD directly on the preprocessed images for segmentation and compared the
results. Another decomposition method, FDM, is used, which was proposed by Singh et al.
[22] and was later used in different works like the analysis of electroencephalogy (EEG) and
electrocardiogram (ECG) signals for different applications, automated alcoholism detection,
and recognition of hand movement from surface electromyogram (sEMG) signals [23–28].
Pattern classification is heavily used in the segmentation domain. It can be supervised,
self-supervised, or unsupervised. Tolias et al. [29] discuss fuzzy clustering as a type of
unsupervised learning. The starting point is taken in the OD, and then the pixels are clustered
according to their probabilities of belonging to vessels or non-vessel clusters. The supervised
method includes the use of convolutional neural networks with the UNet framework as used
by Ronneberger et al. [30], and the self-supervised method includes denoising autoencoder
neural networks as used by Larrazbal et al. [31] and Chen et al. [32] for biomedical images.
Maqsood et al. [33] use a 3D convolutional neural network (CNN) based model to detect
features from retinal images. A cascaded residual attention U-Net (CRAUNet), which uses
DropBlock regularization and multi-scale fusion channel attention (MFCA) for sorting out
overfitting issues and enhancing the performance, was suggested by Ding et al. [34], while
Yakut et al. [35] used a fusion of deep learning method and unsupervised method. TUnet-
LBF, which is a combination of transformer Unet and local binary energy function model,
was proposed by Zhang et al. [36] for both course and fine segmentation. Jayachandran et al.
[37] suggested a multidimensional neural network that works on the transformer principle.
Other segmentation techniques seen in the literature are matched filtering as applied by
Odstrčilík et al. [38] to improve vessel segmentation, while Kumar et al. [39] have used
Laplacian of Gaussian (LoG) filters along with matched filters for segmentation. Kadri et al.
[40], Li et al. [41], and Sreejini et al. [42] used the multi-scale matched filtering approach,
and Gao et al. [43] applied the matched filter only for preprocessing with a UNet. Chaudhuri
et al. [44] also talk about using a matched filter for segmentation. The vessel pixel intensity
is used to get a filter kernel, which extracts the vessel network. The vessels are considered
darker than the background because they reflect less light because of the blood, and they are
taken to be piece-wise linear. Another popular method is applying mathematical morphology
techniques like top hat transformation for vessel extraction. Sigursson et al. [45] applied the
top-hat transformation for feature extraction, while Hassan et al. [46] and Abbadi et al.
[47] used it for vessel segmentation after some image enhancement procedures. Rodrigues
et al. [48] suggested summation of top hat transforms for contrast enhancement and noise
reduction. Morphological operations were used in combination with matched filters by Jiang
et al. [49] and with topological extractors by Zana et al. [50].
Fundus images are easy to capture, cost-effective, non-invasive, and safe [51]. These are
widely used. There are other types of imaging of the retina too. Fluorescein angiography
(FA) is taken by inserting a sodium fluorescein dye in the patient’s arm before capturing the
image, due to which the blood appears brighter. It covers a larger field of view (FOV) than
a regular fundus image [52]. Rodrigues et. al [53] use feature extraction and region growing
on FA and scanning laser ophthalmoscope (SLO) images for retinal vessel segmentation.
Optical coherence tomography angiography (OCTA) is non-invasive, like fundus imaging.
It is also high-resolution [54]. Ultra-widefield (UWF) fundus photography (FP) captures a
higher FOV of 200° as compared to 30-50° of normal fundus images.
While we are talking about the different types of images of the retina, we should note
that medical image segmentation is not limited to retinal images only. There are various

123
85874 Multimedia Tools and Applications (2024) 83:85871–85898

other medical images, the segmentation of whose features helps with the diagnosis of var-
ious diseases. Bennazzouz et al. [55] make use of a modified U-Net, which helps with the
segmentation of white blood cells. Gua et al. [56] have worked with three-dimensional car-
diac images, and Roy et al. [57] have used three-dimensional magnetic resonance images for
detecting brain tumors.

1.2 Motivation and contributions

The motivation behind this work is to approach the retinal vessel segmentation problem from
multiple perspectives so as to develop a robust algorithm that works well with segmenting
any given retinal image. There are some common issues with retinal vessel segmentation,
such as the precise extraction of very light and thin vessels, the quality of the scan, or the
complexity of a technique that gives good performance.
The main contributions of this work are listed below:
• The proposed method is data-independent and works well with any number of images in
a dataset.
• We proposed a way to use one-dimensional FDM for image decomposition in combina-
tion with the Gabor transform, which performed considerably better than the rest of the
methods experimented with.
• The key highlight of this work is that a time- and resource-efficient method for image
segmentation was suggested.
• The proposed approach shows improved accuracy, sensitivity, specificity, precision, and
F1 score in segmenting retinal vessels when compared to existing state-of-the-art methods
for the DRIVE and HRF datasets. Other datasets were also tested with the proposed
algorithms, and the performance was found to be comparable with the state of the art.
The sensitivity was improved for all four datasets.
The rest of this paper is structured as follows: Section 2 discusses the methodology, which
covers the datasets used, preprocessing methods, and the segmentation technique. Sections 3
and 4 are dedicated to the results and discussions, and conclusions of the paper, respectively.
Section 5 mentions the future scope of the work, followed by appendix and references.

2 Methodology

The retinal images, when captured raw, become harder to process and achieve satisfactory
segmentation results. We have processed four publicly available retinal datasets with certain
preprocessing techniques that are widely known for enhancing the images for further process-
ing. These processed images are segmented using the proposed FDM + Gabor technique,
which are then postprocessed using some basic operations like spurring and binarization.
First, we’ll discuss the datasets used in this work.

2.1 Datasets used

• Digital Retinal Images for Vessel Extraction (DRIVE)- This dataset [58] consists of 40
images taken randomly from a population of 400. 20 images are for training and 20 for
testing, with their respective ground truths. Out of 40, 7 show mild DR. The images are
Joint Photographic Experts Group (JPEG) compressed with 565x584x3 pixels.

123
Multimedia Tools and Applications (2024) 83:85871–85898 85875

• STructured Analysis of the Retina (STARE)- This dataset [59] was made with 400 raw
images, out of which 20 were hand-labelled for vessel segmentation by two experts.
These images are in portable pixmap (PPM) format and have a size of 605x700x3 pixels.
• High-Resolution Fundus (HRF)- This dataset [60] has 45 images of size 2336x3504x3
pixels. The images are divided into sets of three, with each set consisting of one glaucoma,
DR, and healthy retinal image, respectively.
• CHASE_DB1- This dataset [61], consists of 14 left and 14 right retinas of children. The
images are JPEG with 960x999x3 pixels. It is a part of the Child Heart and Health Study
in England (CHASE) dataset.

2.2 Preprocessing steps

In this work, several preprocessing techniques were applied to both the training and test
images. The basic preprocessing steps applied to all the images are as follows:

2.2.1 Green channel extraction

The green channel gives the highest vessel-background contrast in RGB, as discussed earlier.
So we have only used the green channel for the segmentation task. The green channel is
extracted by only considering the second layer of a three-dimensional image. Let I be the
original image (Fig. 1(a)) and IG E be the green channel image (Fig. 1(b)).

2.2.2 Region of interest (ROI) extraction

The ROI of the green channel image is extracted by making the pixels outside the retinal
region zero. This is done so as to reduce the computational complexity by focusing only on
the retina pixels as compared to the background. In some datasets like DRIVE, the masks
specifying the retinal region were already given (Fig. 1(c)), in which case they can be straight
away multiplied after normalization for ROI extraction. If the masks are not available, a simple
method is to assume a small threshold, t, below which all the pixels are made zero. t can be

Fig. 1 Preprocessing steps on a DRIVE dataset image

123
85876 Multimedia Tools and Applications (2024) 83:85871–85898

selected by observing the histogram and choosing the first zero crossing. The ROI extracted
image (Fig. 1(d)) can be given as:

IG E (i, j), if IG E (i, j) ≥ t
I R O I (i, j) = (1)
0, otherwise
where i and j are integers less than or equal to the height and width of the green channel
image, respectively, and t is the above-mentioned threshold.

2.2.3 Colour normalisation

Colour normalisation is used to compensate for the varying illumination in affected regions
of an image. This is done by histogram matching the given image with a good contrast image.
The image ’[Link]’ from the STARE database is a good candidate for the reference
image in terms of good contrast and a healthy retinal image. The vessels of the said image are
clear and defined visually, and the same image was used with all the images in all datasets. We
could use colour normalisation on a coloured image, and all three channels could be handled
separately and then combined. The green channel would still give the best vessel-background
contrast (Fig. 1(e)). Here, we have first separated the green channel and only processed that
further. The colour normalisation was done using the imhistmatch command in MATLAB.

2.2.4 CLAHE

CLAHE is applied to enhance the contrast of the image. While adaptive histogram equal-
isation (AHE) also improves contrast and enhances edges, it overamplifies noises in large
homogeneous intensity areas. CLAHE solves the noise amplification issue by using a thresh-
old value on the histogram before transformation. The accuracy is the ratio of correctly
classified pixels to total pixels after classification. The method can be represented mathemat-
ically as:
255 
Q
P= H (i) (2)
8∗8
i=0

where P and Q are the new and old pixel intensity values, respectively. i is a particular pixel
intensity value, and H (i) gives the number of pixels of that intensity. Now, to solve the noise
amplification issue, a certain threshold is set, and pixels are redistributed after clipping the
histogram at that threshold [62, 63].
Here, CLAHE was achieved using the adapthisteq command on MATLAB to give I E N
(Fig. 1(f)).

2.2.5 Thin vessel extraction

Since the thin and thick vessels of a retinal image lie in different regions of the image
histogram, it is best to extract them separately and combine them later. For thin vessel
extraction, first the images are converted into small patches of size 64x64 to better focus on
smaller details. A top-hat transformation with a ball-shaped structuring element is applied to
these small patches. It is a type of mathematical morphology technique that is achieved by
applying a structuring element to extract some details from the image. It works on dilation
and erosion [64]. Here we have adopted erosion followed by dilation, which is then subtracted

123
Multimedia Tools and Applications (2024) 83:85871–85898 85877

from the complemented original image (I E N ) (Fig. 1(g)). The structuring element was taken
to be ball-shaped and of small size. The blocks are then combined, and the resultant image
is termed Ithin (Fig. 1(h)).

2.2.6 Thick vessel extraction

Here we simply used the top hat transformation with a bigger ball-shaped element as com-
pared to the thin vessel case. The size of the element should be comparable to the diameter of
the thickest vessel [46, 65]. It was applied to the complemented green channel image (I R O I )
(Fig. 1 (i)) to give IT hick (Fig. 1(j)). This focuses better on the bigger elements, i.e., thicker
vessels.

2.2.7 Filter blur removal

The final preprocessed image could be enhanced further by removing the blurred background.
Xu et al. [66] used a 25x25-sized median filter on the image to get a blurred image, I Med
(Fig. 1(k)), which could be subtracted from the original image for normalization. We have
used the same filter, but reduced the intensity by some factor, s, and subtracted it from the
summation of IT hick and IT hin to get the final image I Pr e (Fig. 1(l)). s was decided visually
so as not to lose any vessel information.

I Pr e = IT hick + IT hin − I Med /s (3)

Here, IT hick and IT hin are the thick and thin vessel extracted images, respectively. I Med
is the median filtered image, subtracted to remove the filter blur, and s is the normalisation
factor.

2.3 Segmentation

In the proposed method, we are utilising unsupervised learning using image decomposition
techniques. These methods are used because they separate the structural and textural parts of
an image, which helps in focusing on the desirable features of the image, for example, edges,
or in our case, blood vessels [67]. These methods also have the added advantage of requiring
less time and resource requirements as compared to supervised methods, while also being
on par with their performance. This has been explored in Section 3.7.
There are several image-decomposition techniques in the literature. The wavelet transform
has several advantages in terms of energy concentration in an efficient manner or detailed
analysis of multi-resolution images [68]. This method converts the image into four sub-
images: low-low (LL), low-high (LH), high-low (HL), and high-high (HH), out of which
LL is the approximation matrix and the rest are details [68]. For wavelet transform, several
types of wavelet families are employed, depending on the application at hand. In this work,
Gabor is chosen as the mother wavelet. In the next section, we compare the performance with
another wavelet, CDF 9/7. Gabor filters are good for recognition and characteristic extraction
but suffer from high dimensionality and high processing requirements. Blood vessels can be
considered as directional edges with Gaussian profiles, and hence Gabor wavelets, which
can be represented as complex modulated Gaussian, are an optimal choice for blood vessel

123
85878 Multimedia Tools and Applications (2024) 83:85871–85898

segmentation [44, 69]. For feature extraction, a bank of Gabor filters is used with different
frequencies and orientations [70]. It can be described mathematically as:
1
g(x) = ex p( jk0 x − |Ax|2 ) (4)
2
where A is a 2x2 matrix which is diag( −1/2 , 1) with  > 1 for specifying anisotropic nature.
k0 is a vector which defines the frequency of the wavelet [11, 69].
Next, we have also applied FDM to the preprocessed images. FDM is based on the Fourier
transform (FT), which satisfies the Dirichlet conditions and includes non-stationary and
nonlinear signals. This method is good for spectrum and time-frequency analysis of the
signals [22].
The FDM is decomposed into a set of sinusoidal functions called Fourier intrinsic band
functions (FIBFs), whose frequency bands are consecutive. These FIBFs must satisfy the
conditions specified by Singh et al. [22] to categorise a signal as an FIBF:
• The signal should be the summation of all the FIBFs and the mean of the signal.
• The FIBFs should be zero-mean and orthogonal.
• The signal could be written analytically as:
yi (t) + j ŷi (t) = ai (t)ex p( jφi (t)) (5)
where the instantaneous frequency and phase are related by,
d
ωi (t) = φi (t) ≥ 0, ∀t (6)
dt
and the amplitude, ai , is positive for all t, ŷi (t) is the Hilbert transform of the FIBF yi (t).
φi (t) and ωi (t) are the phase and frequency terms.
FDM is a desirable decomposition method because mode mixing does not affect it, and
the decomposed FIBFs are not dependent on the local extrema distribution. FDM is also not
affected by the end effects of the signal boundaries and discontinuities, as seen with other
decomposition methods like empirical mode decomposition (EMD) [22].
Bi et al. [71] used the multiband discrete wavelet method to decompose the image into
subimages, which are converted to one-dimensional signal,s and one-dimensional EMD
is applied to all these signals. After discarding the residue, the signals are reshaped into
subimages according to the scheme used for breaking them. The new set of subimages are
watermarked as per the application stated and reconstructed using Mallat’s multiband discrete
wavelet reconstruction method.
In a similar way, we have first decomposed the image into subimages using a two-
dimensional discrete multiwavelet transform (DMWT) with the ’DB2’ mother wavelet and
then further processed it with FDM. DMWT is further described in Appendix C. FDM has
been suggested by Singh et al. [22] for the analysis of nonlinear and non-stationary signals.
They have also suggested multivariate FDM. Here we are working on the multidimensional
aspect.
The following steps are suggested to implement FDM on image signals (Fig. 2):
1. Pad the preprocessed image or the input image to make all sides divisible by 8.
2. Apply DMWT and decompose the image into 16 subimages of equal size.
3. Convert the subimages to one-dimensional vectors, either row- or column-wise, and
combine the vectors to create one big vector, which serves as the one-dimensional signal
to be input to the FDM algorithm.
4. Obtain the Fast Fourier Transform (FFT) of the signal.

123
Multimedia Tools and Applications (2024) 83:85871–85898 85879

Fig. 2 Steps for applying FDM

5. Decompose the signals into FIBFs.


6. Add the FIBFs to create a single vector, and discard the residue.
7. Convert the signal back to the image form by reversing the first three steps.

The image processed with the above-mentioned unsupervised methods is postprocessed


with morphological operations, and Otsu thresholding is applied for binarization to achieve
the final segmented image. The overall scheme for segmentation is shown in Fig. 3.

Fig. 3 Overall scheme for segmentation

123
85880 Multimedia Tools and Applications (2024) 83:85871–85898

3 Results and discussions

In this section, we have compared the proposed method with other well-known methods,
which can be categorised under supervised, unsupervised, and semi-supervised learning
methods. First, we’ll discuss the performance metrics that we have considered to compare
the several techniques.

3.1 Performance metrics

Performance metrics for segmentation evaluation are as follows [63]: Here, TP = True Posi-
tive, TN = True Negative, FP = False Positive, FN = False Negative.
• Accuracy: It depends on how many pixels are classified correctly as vessel or non-vessel.
TP +TN
Acc = (7)
T P + T N + FP + FN
• Sensitivity: The portion of positive result, also known as recall. This should be high for
vessel extraction [63].
TP
Sen = (8)
T P + FN
• Specificity: The portion of negative result.
TN
Spe = (9)
T N + FP
• Precision: Based on the correctly predicted positives.
TP
Pr e = (10)
T P + FP
• F1 Score: The harmonic mean of the sensitivity and the precision.
2 ∗ Pr e ∗ Sen
Pr e = (11)
Pr e + Sen

Table 1 Wavelet Transform on preprocessed DRIVE dataset


Method Scaling Power of 2 Accuracy Specificity Sensitivity

Gabor 1 0.7652 0.7727 0.6838


2 0.9361 0.9304 0.5948
3 0.9186 0.9304 0.8019
4 0.7899 0.7753 0.9343
5 0.6225 0.5909 0.9335
6 0.4577 0.4121 0.9517
CDF9/7 1 0.4510 0.3891 0.9941
2 0.4536 0.3921 0.9931
3 0.4242 0.3921 0.9124
4 0.2925 0.3747 0.8772
5 0.2930 0.2382 0.8341
6 0.2930 0.2382 0.8341
The bold entries highlight the highest value of the column

123
Multimedia Tools and Applications (2024) 83:85871–85898 85881

Table 2 Image decomposition Method Accuracy Specificity Sensitivity


methods on preprocessed DRIVE
dataset EMD 0.8098 0.8011 0.8859
VMD 0.4545 0.3957 0.9853
FDM 0.9602 0.9691 0.8612
The bold entries highlight the highest value of the column

3.2 Unsupervised learning methods

First, we will compare the performance of the Gabor wavelet with CDF 9/7. CDF 9/7 wavelet
is good for image compression and has a linear phase [72]. CDF stands for Cohen-Daubechies-
Feauveau and factors 9 and 7 mention the lengths of high- or low-pass filters. The filter
coefficients were suggested in the JPEG 2000 compression standard, given by Taubman et
al. [73]. Both types of wavelets (CDF 9/7 and Gabor) were scaled with 6 different powers of 2,
from 1-6 (scaling values: 1, 2, 4, 8, 16, 32). Table 1 shows the effect of using Gabor and CDF
wavelets with different scaling on the DRIVE dataset images. It can be observed that using
Gabor wavelet with scaling 4 gives the highest accuracy (93.61%) and specificity (93.04%)
amongst both methods and Gabor with scaling 6 gives the highest sensitivity (80.19%).
Higher sensitivity could be seen for the CDF 9/7 wavelet transform (Table 1), but since
accuracy and specificity were too low, which gives undesirable results, we ignore this result.
Next, we have applied EMD, VMD, and FDM to the preprocessed images. For the steps
of implementation of EMD and VMD, see Appendices A and B. VMD is more adaptive as
compared to EMD as it is capable of decomposing the signal around a particular (center) fre-
quency [74]. The signal gets decomposed into multiple frequency- and amplitude-modulated
waves [75]. But EMD has a smaller number of frequency components per mode and works
better with low-frequency noises and non-stationary signals as compared to VMD [75].
EMD suffers from various difficulties; for example, it involves a lot of mathematical analy-
sis, dependence on stopping criteria, sifting amount, convergence, and others. These issues
could be solved using FDM instead. Table 2 shows a comparison based on accuracy between
the three decomposition methods, and we find the highest accuracy to be 96.02% and the
highest specificity to be 96.91%, as in the case of FDM. The highest sensitivity (86.12%)
was found with the EMD method.
Finally, we use postprocessing techniques like removing small, unconnected pixels and
spurs from the best results of unsupervised learning techniques. The accuracy after postpro-
cessing is reported in Table 3.

Table 3 Post processing on unsupervised learning results


Method Scaling Power of 2 Accuracy Specificity Sensitivity

FDM NA 0.9728 0.9871 0.8149


Gabor 2 0.9567 0.9834 0.6838
Gabor 3 0.9260 0.9391 0.7966
Gabor+FDM 2 0.9739 0.9863 0.8366
Gabor+FDM 3 0.9336 0.9464 0.8072
The bold entries highlight the highest value of the column

123
85882 Multimedia Tools and Applications (2024) 83:85871–85898

Figure 4 shows the results of applying FDM and Gabor transform to the DRIVE dataset
test image 19. The FDM-processed image shows a more accurate segmentation, while the
Gabor retains more thin vessels. We have combined them both to achieve higher accuracy.

Fig. 4 a) FDM segmented DRIVE 19 image. b) Gabor with scaling 4 segmented Drive 19 image. c) FDM +
Gabor scaling 4 segmented DRIVE 19 image. d) Ground truth DRIVE 19 image. e) Overlapping image of c
and d (Red is the ground truth and green is our result)

123
Multimedia Tools and Applications (2024) 83:85871–85898 85883

Fig. 5 a) Post processed, FDM denoised, and autoencoder segmented DRIVE 1 image. b) Ground truth Drive
1 image

3.3 Self supervised method - denoising autoencoders

Autoencoders are unsupervised in nature, but they can be used as self-supervised networks,
which use similar techniques as supervised networks.
The 20 training images of the DRIVE dataset were passed through preprocessing and
unsupervised methods to get 20 processed images. 19000 random patches of size 48x48
were extracted from these processed images and trained with the denoising autoencoders.
Similarly, the test images were processed and converted to patches.
Autoencoders, when used in denoising form, clean the input image [31, 76]. They are
based on encoder-decoder combinations. The encoder encodes the input image to a lower
dimensionality. In our work, we have selected 3x3 convolutional layers with a sigmoid acti-
vation function and 2x2 max pooling layers for the encoder. The decoder consists of 3x3
convolutional layers with a sigmoid activation function and 2x2 up-sampling layers.
Larrazabal et al. [31] suggested the use of denoising autoencoders as a postprocessing tool.
Here, we are using them as a vessel segmentation tool instead. It could also be considered a
postprocessing tool for the unsupervised methods, but the timing requirements increase and
the results are still comparable with the proposed unsupervised method alone.
Preprocessed images, which are further decomposed using FDM, show the best accuracy
of 96.96% in the case of denoising autoencoder-based vessel segmentation (Fig. 5(a), Table 4).
The highest specificity was 98.83% for Gabor wavelet transform with scaling 1 (Table 4) and
the highest sensitivity of 86.21% for FDM (Table 5).

Table 4 Post processed autoencoder segmentation results


Method Scaling Power of 2 Accuracy Specificity Sensitivity

EMD NA 0.9662 0.9878 0.7528


FDM NA 0.9696 0.9795 0.8598
Gabor 1 0.9616 0.9883 0.6980
Gabor 2 0.9646 0.9848 0.7647
Gabor 3 0.9646 0.9846 0.7672
The bold entries highlight the highest value of the column

123
85884 Multimedia Tools and Applications (2024) 83:85871–85898

Table 5 Segmentation using denoising autoencoders on the DRIVE dataset


Method Scaling Power of 2 Accuracy Specificity Sensitivity

EMD NA 0.9663 0.9878 0.7544


VMD NA 0.2002 0.9794 0.8802
FDM NA 0.9697 0.9794 0.8621
Gabor 1 0.9618 0.9882 0.7002
Gabor 2 0.9646 0.9847 0.7663
Gabor 3 0.9646 0.9845 0.7676
Gabor 6 0.8534 0.9131 0.0791
The bold entries highlight the highest value of the column

Table 6 UNet segmentation on the DRIVE dataset


Method Scaling Power of 2 Accuracy Specificity Sensitivity

EMD NA 0.9630 0.9710 0.7952


VMD NA 0.7218 0.6966 0.9573
FDM NA 0.9642 0.9765 0.8411
Gabor 1 0.9303 0.9666 0.5720
Gabor 2 0.9690 0.9752 0.9002
Gabor 3 0.9675 0.9777 0.8541
Gabor 6 0.8810 0.9311 0.4836
CDF9/7 1 0.9730 0.9813 0.8817
CDF9/7 6 0.8626 0.8850 0.6532
The bold entries highlight the highest value of the column

Fig. 6 a) Post processed CDF9/7 UNet DRIVE 1 image. b) Ground truth Drive 1 image

123
Multimedia Tools and Applications (2024) 83:85871–85898 85885

Table 7 Post processed UNet segmentation results


Method Scaling Power of 2 Accuracy Specificity Sensitivity

EMD NA 0.9635 0.9811 0.7906


FDM NA 0.9646 0.9771 0.8405
Gabor 2 0.9691 0.9755 0.8986
Gabor 3 0.9676 0.9780 0.8528
CDF9/7 1 0.9731 0.9816 0.8792
The bold entries highlight the highest value of the column

3.4 Supervised method - UNet

The U-Net architecture used in this work consists of 2x2 and 3x3 convolutional layers along
with a ReLU, 2x2 max pooling layers, and upsampling layers, as mentioned in [30]. Out of the
19000 patches, as used with the denoising autoencoders, 90% patches are used for training,
and the rest are used for the validation set. The test images are taken from the test folder of
the dataset and converted to another set of 19000 test patches, which are later combined to
form the original-sized segmented images.
The code for patch extraction and UNet implementation can be found at [Link]
orobix/retina-unet. The network accepts the processed DRIVE image patches and trains on
them in a supervised manner to reduce training loss over 150 epochs. The stochastic gradient
descent (SGD) optimizer is used for the cross-entropy loss function (Table 6).
UNet gave an overall better performance than denoising autoencoders, and CDF 9/7, with
a scaling of 1 for processing, gave the highest accuracy of 97.31% (Fig. 6(a), Table 7). Highest
specificity was 98.16% for CDF 9/7 wavelet transform with scaling 1 (Table 7) and highest
sensitivity was 90.02% for Gabor wavelet transform with scaling 4 (Table 6).

Table 8 Image decomposition methods on other datasets


Dataset Method Scaling Power of 2 Acc Spe Sen

STARE FDM NA 0.9736 0.9840 0.8493


(Unsupervised) Gabor 1 0.8366 0.8504 0.6570
Gabor 2 0.9586 0.9780 0.7058
Gabor+FDM 2 0.9762 0.9846 0.8759
Gabor+FDM 3 0.9399 0.9474 0.7992
HRF FDM NA 0.9556 0.9721 0.7201
(Unsupervised) Gabor 2 0.8455 0.8824 0.4072
Gabor+FDM 2 0.9555 0.9721 0.7202
CHASE_DB1 FDM NA 0.9534 0.9635 0.7753
(Unsupervised) Gabor 1 0.8509 0.8587 0.7047
Gabor 2 0.8036 0.8119 0.6563
Gabor 3 0.9075 0.9160 0.7488
Gabor+FDM 2 0.9630 0.9562 0.8216
Gabor 3 0.8995 0.9044 0.8079
The bold entries highlight the highest value of the column

123
85886 Multimedia Tools and Applications (2024) 83:85871–85898

Fig. 7 a) Proposed unsupervised method based segmentation on Stare 0163 image. b) Ground truth Stare 0163
image

3.5 STARE, HRF, CHASE_DB1 datasets

The proposed unsupervised method was applied to other datasets, and the Gabor with scaling
4 and FDM combination gave the best accuracy with the STARE and HRF datasets and scaling
3, with the CHASE_DB1 dataset (Table 8). Figures 7, 8, 9 show the results of unsupervised
method on STARE, HRF, and CHASE_DB1 datasets, respectively (Table 8).
dummy

3.6 Supervised methods for other datasets

For supervised methods (self-supervised and supervised), as seen with the DRIVE dataset,
the CDF 9/7 wavelet transform with UNet was chosen for further processing the preprocessed
images.
The results on the HRF dataset can be seen in Fig. 10. Here, out of 20 images, four were
taken as tests and the rest as training data. The maximum accuracy was found to be 96.21%
(Table 9).

3.6.1 Leave one out cross validation

In the case of smaller datasets with no clear distinction between training and test data, a
method, known as K-fold cross validation (CV) is employed. In this method the training data
is divided into k folds, and each fold is considered as the test data once, while the other k-1
folds are taken as the training data. If we consider all of the dataset images as a test case once,

Fig. 8 a) Proposed unsupervised method based segmentation on HRF 07_g image. b) Ground truth HRF 07_g
image

123
Multimedia Tools and Applications (2024) 83:85871–85898 85887

Fig. 9 a) FDM method based segmentation on Chase 11R image. b) Ground truth Chase 11R image

that is, the number of instances is equal to the number of folds, it is a special case of K-fold
CV, known as leave one out cross validation (LOO-CV). In the case of smaller datasets, it is
better to employ LOO-CV for enhanced accuracy [77].
For the STARE dataset, since there is division between training and test images and the
images are only 20, we have applied the LOO-CV method to the testing. 20 folds were
considered, with one image as the test image per iteration and 20 iterations in total. Due to
the increased requirement and unavailability of resources, the number of epochs were reduced
to 50, and the total number of patches for 19 images per iteration were taken to be 1900.
Table 10 shows the results of all iterations, and the best accuracy was found to be 96.72%,
the best specificity was 98.39%, and the best sensitivity was 95.00%. Figure 11 shows the
image with the best accuracy. This method could also be used with HRF and CHASE_DB1
datasets, but due to the increased sizes of images and increase in the number of images, the
resource requirement increased further.

3.7 Time and resource requirement

Unsupervised methods take much less time and resources as compared to self-supervised
and supervised methods (Table 11). Self-supervised and supervised methods perform better
in terms of sensitivity.
*PC details: 32 GB RAM, 64-bit OS, 11th Gen Intel(R) Core (TM) i7-11700 @ 2.50GHz
processor. **HPC IITD: GPU: NVIDIA V100 (32GB, 5120 CUDA cores), CPU: 2x Intel
Xeon G-6148 (20 cores, 2.4 GHz) "Skylake", RAM: 96GB.

Fig. 10 a) Supervised method based segmentation on HRF 15_h image. b) Ground truth HRF 5_h image

123
85888 Multimedia Tools and Applications (2024) 83:85871–85898

Table 9 Supervised method Dataset Acc (%) Sen (%) Spe (%)
based segmentation on other
datasets Stare 96.72 91.01 97.18
HRF 96.21 79.29 97.72

Table 10 K-Fold cross validation Test Image Accuracy Specificity Sensitivity


on the stare dataset with CDF 9/7
processing and UNet im0001 0.9371 0.9499 0.7895
segmentation
im0002 0.9404 0.9604 0.6594
im0003 0.8226 0.8145 0.9510
im0004 0.9561 0.9839 0.6094
im0005 0.9328 0.9406 0.8545
im0044 0.9511 0.9556 0.8915
im0077 0.9609 0.9640 0.9220
im0081 0.9672 0.9718 0.9101
im0082 0.9623 0.9660 0.9186
im0139 0.8050 0.7929 0.9441
im0162 0.7535 0.7401 0.9278
im0163 0.8459 0.8371 0.9500
im0235 0.7863 0.7744 0.9072
im0236 0.8538 0.8457 0.9351
im0239 0.9615 0.9780 0.7875
im0240 0.9068 0.9142 0.8417
im0255 0.8870 0.8854 0.9035
im0291 0.9053 0.9041 0.9274
im0319 0.8618 0.8600 0.9002
im0324 0.9360 0.9458 0.8003
The bold entries highlight the highest value of the column

Fig. 11 a) Supervised method based segmentation on STARE im0081 image. b) Ground truth STARE im0081
image

123
Multimedia Tools and Applications (2024) 83:85871–85898 85889

Table 11 Approximate time and resource requirements on 20 images of DRIVE dataset


Time Required for
Method processing 20 images Resources employed

Preprocessing + FDM +
Gabor + Postprocessing 1 minute 46 seconds CPU*
Preprocessing + FDM 22 seconds CPU*
15 minutes 46 secs
(150 epochs)
Preprocessing + FDM:
27 seconds
Training:
Preprocessing + 14 mins 44 secs
FDM + Test: HPC**: 2 GPUs, MPI
Autoencoder 35 secs (approx) processors and CPUs
02:02:01 hours
(150 epochs)
Preprocessing + CDF 9/7:
34 seconds
Training:
Preprocessing 1 hour 55 mins 59 secs
CDF 9/7 + Test: HPC**: 2 GPUs, MPI
UNet 6 minutes (approx) processors and CPUs

Fig. 12 An overall comparison on the accuracy of different segmentation techniques on the DRIVE dataset

123
85890 Multimedia Tools and Applications (2024) 83:85871–85898

Fig. 13 An overall comparison on the sensitivity of different segmentation techniques on the DRIVE dataset

Fig. 14 An overall comparison on the specificity of different segmentation techniques on the DRIVE dataset

123
Multimedia Tools and Applications (2024) 83:85871–85898 85891

Table 12 Comparison with the state-of-the-art segmentation methods


No. Author Year Dataset Acc (%) Sen (%) Spe (%) Pre (%) F1 (%)

1 Jiang et al. 2019 Drive 97.09 82.46 98.90 - 82.46


[78] Stare 97.81 85.79 99.56 - 84.92
HRF - - - - -
Chase_DB1 97.21 78.39 98.94 - 80.62
2 Wang et al. 2020 Drive 95.81 79.91 98.13 - 82.93
[79] Stare 96.73 81.86 98.44 - 83.79
HRF 96.54 78.03 98.43 - 81.91
Chase_DB1 96.70 82.39 98.13 - 80.74
3 Wu et al. 2020 Drive 95.82 79.96 98.13 86.18 82.14
[7] Stare 96.72 79.63 98.63 86.97 82.98
HRF - - - - -
Chase_DB1 96.88 90.03 98.80 88.01 83.69
4 Khan et al. 2021 Drive 96.10 81.25 97.63 - -
[80] Stare 95.86 80.78 97.21 - -
HRF - - - - -
Chase_DB1 95.78 80.12 97.30 - -
5 Mahapatra 2022 Drive 96.05 70.20 98.44 81.24 75.31
et al. [81] Stare 96.01 68.46 98.02 74.40 71.29
HRF (G) 96.58 71.12 98.43 70.28 70.70
Chase_DB1 - - - - -
6 Liu et al. 2022 Drive 95.61 79.85 97.91 - 82.29
[82] Stare 95.67 79.63 97.92 - 81.72
HRF - - - - -
Chase_DB1 96.72 80.20 97.94 - 82.36
7 Proposed 2023 Drive 97.39 88.36 99.62 94.05 84.15
Stare 97.62 88.51 98.97 88.02 85.07
HRF 96.57 79.07 99.42 92.65 78.03
Chase_DB1 96.30 90.37 96.35 51.48 63.20

3.8 Comparison of methods

Figures 12, 13, and 14 show an overall comparison on the accuracy, sensitivity, and specificity,
respectively, of the DRIVE dataset test images. Here we have only discussed the results for
which all three parameters are above 80%.
Table 12 shows a comparison of proposed methods with current state of art.

4 Conclusion

This work suggests a time- and resource-efficient method for accurate vessel segmentation.
The best method, based on the five performance metrics and processing requirements-wise, on
the DRIVE dataset was FDM in combination with the Gabor transform and some generic pre-

123
85892 Multimedia Tools and Applications (2024) 83:85871–85898

and post-processing methods. The proposed method works well with other datasets as well.
The achieved sensitivity is higher than the state of the art for all the datasets. The supervised
method performs well as per the metrics as well, but the time and resource requirements need
to be considered.
Since the method is not data-dependent and uses much less resources than well-known
methods, it can be very well implemented in practical scenarios where bulk data is to be
handled. This would be a great tool of assistance to the medical community by automating
the generation of an accurate vessel map that can be used for the diagnosis or monitoring of
various eye diseases. Being said that, there is always going to be a need for a human in the
loop for medical issues. The work simply contributes as a tool of assistance for faster and
accurate vessel segmentation.

5 Future scope

• Accuracy can be further improved by formulating two-dimensional FDM instead of


converting the image to a one-dimensional signal first and then applying one-dimensional
FDM.
• The high-performing segmentation method can be employed in further applications like
retinal scan-based biometric authentication, retinal disease diagnosis, and many more to
check the real-time performance.
• The variable parameters for morphological operations and postprocessing need to be
set separately for different datasets. These parameters can be optimised in an automatic
fashion.
• The testing has been done on publicly available datasets. While these datasets also suffer
from some quality issues like irregular illumination or noise, the study could be further
extended to some datasets captured by non-professionals to test the performance of the
proposed algorithm in a low-resource settings.

Appendix A Empirical mode decomposition

It decomposes the image into multiple intrinsic mode functions (IMFs), which have zero
average envelopes and their local extrema are equal to or one less or more than their zero
crossings [15, 83, 84]. Linderhed [85] suggested the use of EMD for image compression,
which was first proposed by Huang et al. [86] for the analysis of nonlinear and non-stationary
time series. The steps for achieving 2D EMD are as follows [85, 87]:
1. First, the local extrema are found for the input image.
2. For finding both the envelopes, upper and lower, cubic spline interpolation is applied to
the local extrema.
3. Then the mean of these envelopes is found according to:
(emax (i, j) + emin (i, j))
E max (i, j) = (A1)
2
where emax and emin are the upper and lower envelopes, respectively.
4. The mean is then subtracted from the input signal.
5. The mean is checked to be close to zero, or else the first 4 steps are repeated with the
signal from step 4 as input to get the current IMF.

123
Multimedia Tools and Applications (2024) 83:85871–85898 85893

6. The residue is found by subtracting the current IMF from the input signal. Then the next
IMF is found, till we get a residue with no extrema left, or the mean is close to zero.

Appendix B Variational mode decomposition

VMD breaks down the signal from low to high frequencies as opposed to EMD [88]. VMD
was proposed by Dragomiretskiy et al. [75], and the steps to implement VMD are as follows
[75, 85]:
• The Hilbert transform is applied to the signal obtain a frequency spectrum that is unilat-
eral.
• The frequency spectrum of each band is mixed with the exponential of the center fre-
quency to shift the spectrum to the baseband.
• Estimation of bandwidth is done by Gaussian smoothening implemented by taking the
squared L2 -norm of the gradient.
• As per the above steps, the VMD problem is given as:
     2
 
min ∂t δ(t) + j ∗ u k (t) e− jωk t  , (B2)
{u k },{ωk }  πt 
k 2

s.t., uk = f (B3)
k

Here, u k are the modes. f is the input signal. t is the time period, k is the number of the
mode, and ωk is the center frequency [75].

Appendix C Discrete multiwavelet transform

DMWT is similar to discrete wavelet transform (DWT), but has multiple mother wavelets
and scaling functions [89]. The scaling functions can be given as a vector:
 := (φ1 , φ2 ...φr )T (C4)
where r is the multiplicity of DMWT, φx is the scaling function for a wavelet, and  is the
overall scaling vector [89]. The vector for multi-wavelets is:
:= (ψ1 , ψ2 ...ψr )T (C5)
where ψx is the unscaled wavelet, is the multi-wavelet vector, and r is the number of
wavelets [89].
The multiwavelets are represented by matrix impulse responses, gi (n), which are r x r
matrices, with transfer functions G i (z), where i is the channel of the filter bank used for
implementing the wavelet. The transform matrix, U, is created using these matrix impulse
responses. Then, for implementing two-dimensional DMWT:
Y = U XU T (C6)
Here Y is the transformed image, X is the input image, and U is the transformation matrix
[89]. This Y image can be recovered to X by applying the inverse discrete multiwavelet
transform (IDMWT). Kromka et al. [90] created a toolbox for DMWT and IDMWT for
multiple mother wavelets.

123
85894 Multimedia Tools and Applications (2024) 83:85871–85898

Data Availability The four publicly available datasets, used in the study, can be found at:
DRIVE: [Link]
STARE: [Link]
HRF: [Link]
CHASE_DB1: [Link] [Link]

Declarations

Human and animal rights No animal or human experiments were conducted as part of this research.

Informed consent Informed consent was obtained from all individual participants included in the study.

Conflict of Interest The authors declare that they have no conflict of interest.

References
1. Xiuqin P, Zhang Q, Zhang H, Li S (2019) A fundus retinal vessels segmentation scheme based on the
improved deep learning u-net model. IEEE Access 7:122634–122643. [Link]
2019.2935138
2. Abrámoff MD, Garvin MK, Sonka M (2010) Retinal imaging and image analysis. IEEE Rev Biomed Eng
3:169–208. [Link]
3. Fraz MM, Remagnino P, Hoppe A, Uyyanonvara B, Rudnicka AR, Owen CG, Barman SA (2012)
Blood vessel segmentation methodologies in retinal images-a survey. Comput Methods Programs Biomed
108(1):407–433
4. L Srinidhi C, Aparna P, Rajan J (2017) Recent advancements in retinal vessel segmentation. J Med Syst
41, 1–22
5. Aleem S, Sheng B, Li P, Yang P, Feng DD (2018) Fast and accurate retinal identification system: Using
retinal blood vasculature landmarks. IEEE Trans Industr Inf 15(7):4099–4110
6. Kanski JJ, Bowling B (2015) Kanski’s Clinical Ophthalmology E-book: a Systematic Approach. Elsevier
Health Sciences, ???
7. Wu Y, Xia Y, Song Y, Zhang Y, Cai W (2020) Nfn+: a novel network followed network for retinal vessel
segmentation. Neural Netw 126:153–162
8. Oliveira A, Pereira S, Silva CA (2018) Retinal vessel segmentation based on fully convolutional neural
networks. Expert Syst Appl 112:229–242
9. Narasimhan K, Neha V, Vijayarekha K (2012) Hypertensive retinopathy diagnosis from fundus images
by estimation of avr. Procedia engineering 38:980–993
10. Leopold HA, Orchard J, Zelek J, Lakshminarayanan V (2017) Use of gabor filters and deep networks in
the segmentation of retinal vessel morphology. In: Imaging, Manipulation, and Analysis of Biomolecules,
Cells, and Tissues XV, vol 10068, p 100680. International Society for Optics and Photonics
11. Soares JV, Leandro JJ, Cesar RM, Jelinek HF, Cree MJ (2006) Retinal vessel segmentation using the 2-d
gabor wavelet and supervised classification. IEEE Trans Med Imaging 25(9):1214–1222
12. Swathi C, Anoop B, Dhas DAS, Sanker SP (2017) Comparison of different image preprocessing methods
used for retinal fundus images. In: 2017 Conference on Emerging Devices and Smart Systems (ICEDSS),
pp 175–179. IEEE
13. Budak Ü, Şengür A, Guo Y, Akbulut Y, Vespa LJ (2017) A novel approach based on image processing
algorithms for microaneurysm candidate detection. In: 2017 International artificial intelligence and data
processing symposium (IDAP), pp 1–4. IEEE
14. Lei G, Xia Y, Zhang W, Chen D, Wang D (2020) Comparative analysis of pre-process pipelines for
automatic retinal vessel segmentation. In: 2020 39th Chinese Control Conference (CCC), pp 3216–3220.
IEEE
15. Zhou M, Xia H, Zhong H, Zhang J, Gao F (2019) A noise reduction method for photoacoustic imaging
in vivo based on emd and conditional mutual information. IEEE Photonics J 11(1):1–10
16. Mookiah MRK, Acharya UR, Fujita H, Koh JE, Tan JH, Chua CK, Bhandary SV, Noronha K, Laude A,
Tong L (2015) Automated detection of age-related macular degeneration using empirical mode decom-
position. Knowl-Based Syst 89:654–668

123
Multimedia Tools and Applications (2024) 83:85871–85898 85895

17. Parashar D, Agrawal DK (2022) Classification of glaucoma stages using image empirical mode decom-
position from fundus images. J Digit Imaging 1–10
18. Lahmiri S, Boukadoum M (2014) Automated detection of circinate exudates in retina digital images using
empirical mode decomposition and the entropy and uniformity of the intrinsic mode functions. Biomed
Eng Biomed Tech 59(4):357–366
19. Marrugo AG, Vargas R, Chirino M, Millán MS (2015) On the illumination compensation of retinal
images by means of the bidimensional empirical mode decomposition. In: 11th International symposium
on medical information processing and analysis, vol 9681, pp 85–91. SPIE
20. Shamaee Z, Mivehchy M (2023) Dominant noise-aided emd (demd): Extending empirical mode decom-
position for noise reduction by incorporating dominant noise and deep classification. Biomed Signal
Process Control 80:104218
21. Lahmiri S, Shmuel A (2017) Variational mode decomposition based approach for accurate classification
of color fundus images with hemorrhages. Opt Laser Technol 96:243–248
22. Singh P, Joshi SD, Patney RK, Saha K (2017) The fourier decomposition method for nonlinear and
non-stationary time series analysis. Proceedings of the Royal Society A: Mathematical, Physical and
Engineering Sciences 473(2199):20160871
23. Zheng J, Huang S, Pan H, Tong J, Wang C, Liu Q (2021) Adaptive power spectrum fourier decomposition
method with application in fault diagnosis for rolling bearing. Measurement 183:109837
24. Fatimah B, Singh P, Singhal A, Pachori RB (2020) Detection of apnea events from ecg segments using
fourier decomposition method. Biomed Signal Process Control 61:102005
25. Fatimah B, Javali A, Ansar H, Harshitha BG, Kumar H (2020) Mental arithmetic task classification using
fourier decomposition method. In: 2020 International conference on communication and signal processing
(ICCSP), pp 0046–0050. [Link]
26. Mehla VK, Singhal A, Singh P (2020) A novel approach for automated alcoholism detection using fourier
decomposition method. J Neurosci Methods 346:108945
27. Fatimah B, Singh P, Singhal A, Pachori RB (2021) Hand movement recognition from semg signals using
fourier decomposition method. Biocybernetics Biomed Eng 41(2):690–703
28. Singh P, Srivastava I, Singhal A, Gupta A (2019) Baseline wander and power-line interference removal
from ecg signals using fourier decomposition method. In: Machine intelligence and signal analysis, pp
25–36. Springer, ???
29. Tolias YA, Panas SM (1998) A fuzzy vessel tracking algorithm for retinal images based on fuzzy clustering.
IEEE Trans Med Imaging 17(2):263–273
30. Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmen-
tation. In: Medical image computing and computer-assisted intervention (MICCAI). LNCS, vol 9351, pp
234–241. Springer, ???. (available on arXiv:1505.04597 [[Link]]). [Link]
Publications/2015/RFB15a
31. Larrazabal AJ, Martínez C, Glocker B, Ferrante E (2020) Post-dae: anatomically plausible segmentation
via post-processing with denoising autoencoders. IEEE Trans Med Imaging 39(12):3813–3820
32. Chen M, Shi X, Zhang Y, Wu D, Guizani M (2021) Deep feature learning for medical image analysis
with convolutional autoencoder neural network. IEEE Trans Big Data 7(4):750–758. [Link]
1109/TBDATA.2017.2717439
33. Maqsood S, Damaševičius R, Maskeliūnas R (2021) Hemorrhage detection based on 3d cnn deep learning
framework and feature fusion for evaluating retinal abnormality in diabetic patients. Sensors 21(11).
[Link]
34. Dong F, Wu D, Guo C, Zhang S, Yang B, Gong X (2022) Craunet: A cascaded residual attention u-net
for retinal vessel segmentation. Comput Biol Med 147:105651
35. Yakut C, Oksuz I, Ulukaya S (2022) A hybrid fusion method combining spatial image filtering with
parallel channel network for retinal vessel segmentation. Arab J Sci Eng 1–14
36. Zhang H, Ni W, Luo Y, Feng Y, Song R, Wang X (2023) Tunet-lbf: Retinal fundus image fine segmentation
model based on transformer unet network and lbf. Comput Biol Med 159:106937
37. Jayachandran A, Kumar SR, Perumal T (2023) Multi-dimensional cascades neural network models for
the segmentation of retinal vessels in colour fundus images. Multimed Tools Appl 1–17
38. Odstrčilík J, Jan J, Gazárek J, Kolář R (2009) Improvement of vessel segmentation by matched filtering in
colour retinal images. In: World Congress on Medical Physics and Biomedical Engineering, September
7-12, 2009, Munich, Germany, pp 327–330. Springer
39. Kumar D, Pramanik A, Kar SS, Maity SP (2016) Retinal blood vessel segmentation using matched filter
and laplacian of gaussian. In: 2016 International conference on signal processing and communications
(SPCOM), pp 1–5. [Link]

123
85896 Multimedia Tools and Applications (2024) 83:85871–85898

40. Kadry S, Rajinikanth V, Damaševičius R, Taniar D (2021) Retinal vessel segmentation with slime-mould-
optimization based multi-scale-matched-filter. In: 2021 Seventh international conference on bio signals,
images, and instrumentation (ICBSII), pp 1–5. [Link]
41. Li Q, You J, Zhang D (2012) Vessel segmentation and width estimation in retinal images using multiscale
production of matched filter responses. Expert Syst Appl 39(9):7600–7610
42. Sreejini K, Govindan V (2015) Improved multiscale matched filter for retina vessel segmentation using
pso algorithm. Egyptian Inf J 16(3):253–260
43. Gao X, Cai Y, Qiu C, Cui Y (2017) Retinal blood vessel segmentation based on the gaussian matched filter
and u-net. In: 2017 10th International congress on image and signal processing, biomedical engineering
and informatics (CISP-BMEI), pp 1–5. [Link]
44. Chaudhuri S, Chatterjee S, Katz N, Nelson M, Goldbaum M (1989) Detection of blood vessels in retinal
images using two-dimensional matched filters. IEEE Trans Med Imaging 8(3):263–269
45. Sigurðsson EM, Valero S, Benediktsson JA, Chanussot J, Talbot H, Stefánsson E (2014) Automatic retinal
vessel extraction based on directional mathematical morphology and fuzzy classification. Pattern Recogn
Lett 47:164–171
46. Hassan G, El-Bendary N, Hassanien AE, Fahmy A, Snasel V et al (2015) Retinal blood vessel segmentation
approach based on mathematical morphology. Procedia Comput Sci 65:612–622
47. El Abbadi NK, Al Saadi EH (2013) Blood vessels extraction using mathematical morphology. J Comput
Sci 9(10):1389
48. Rodrigues J, Bezerra N (2016) Retinal vessel segmentation using parallel grayscale skeletonization algo-
rithm and mathematical morphology. In: 2016 29th SIBGRAPI conference on graphics, patterns and
images (SIBGRAPI), pp 17–24. [Link]
49. Jiang Z, Yepez J, An S, Ko S (2017) Fast, accurate and robust retinal vessel segmentation system. Biocy-
bernetics Biomed Eng 37(3):412–421
50. Zana F, Klein J-C (2001) Segmentation of vessel-like patterns using mathematical morphology and cur-
vature evaluation. IEEE Trans Image Process 10(7):1010–1019. [Link]
51. Shen Z, Fu H, Shen J, Shao L (2021) Modeling and enhancing low-quality retinal fundus images. IEEE
Trans Med Imaging 40(3):996–1006. [Link]
52. Ding L, Bawany MH, Kuriyan AE, Ramchandran RS, Wykoff CC, Sharma G (2020) A novel deep learning
pipeline for retinal vessel detection in fluorescein angiography. IEEE Trans Image Process 29:6561–6573.
[Link]
53. Rodrigues EO, Conci A, Liatsis P (2020) Element: Multi-modal retinal vessel segmentation based on a
coupled region growing and machine learning approach. IEEE J Biomed Health Inform 24(12):3507–
3519. [Link]
54. Hu K, Jiang S, Zhang Y, Li X, Gao X (2022) Joint-seg: Treat foveal avascular zone and retinal vessel
segmentation in octa images as a joint task. IEEE Trans Instrum Meas 71:1–13. [Link]
TIM.2022.3193188
55. Benazzouz M, Benomar ML, Moualek Y (2022) Modified u-net for cytological medical image segmen-
tation. Int J Imaging Syst Technol 32(5):1761–1773
56. Guo S, Liu X, Zhang H, Lin Q, Xu L, Shi C, Gao Z, Guzzo A, Fortino G (2023) Causal knowledge fusion
for 3d cross-modality cardiac image segmentation. Inf Fusion 99:101864
57. Roy S, Maji P (2023) Tumor delineation from 3-d mr brain images. SIViP 17(7):3433–3441
58. Staal JJ, Abramoff MD, Niemeijer M, Viergever MA, Ginneken B (2004) Ridge based vessel segmentation
in color images of the retina. IEEE Trans Med Imaging 23(4):501–509
59. Hoover AD, Kouznetsova V, Goldbaum M (2000) Locating blood vessels in retinal images by piecewise
threshold probing of a matched filter response. IEEE Trans Med Imaging 19(3):203–210. [Link]
10.1109/42.845178
60. Budai A, Bock R, Maier A, Hornegger J, Michelson, G (2013) Robust vessel segmentation in fundus
images. Int J Biomed Imaging 2013
61. Fraz MM, Remagnino P, Hoppe A, Uyyanonvara B, Rudnicka AR, Owen CG, Barman SA (2012) An
ensemble classification-based approach applied to retinal blood vessel segmentation. IEEE Trans Biomed
Eng 59(9):2538–2548. [Link]
62. Li K, Qi X, Luo Y, Yao Z, Zhou X, Sun M (2020) Accurate retinal vessel segmentation in color fundus
images via fully attention-based networks. IEEE J Biomed Health Inform 25(6):2071–2081
63. Upadhyay K, Agrawal M, Vashist P (2020) Wavelet based fine-to-coarse retinal blood vessel extraction
using u-net model. In: 2020 International conference on signal processing and communications (SPCOM),
pp 1–5. IEEE
64. Bai X, Zhou F, Xue B (2012) Image enhancement using multi scale image features extracted by top-hat
transform. Optics & Laser Technology 44(2):328–336

123
Multimedia Tools and Applications (2024) 83:85871–85898 85897

65. Khawaja A, Khan TM, Naveed K, Naqvi SS, Rehman NU, Junaid Nawaz S (2019) An improved retinal
vessel segmentation framework using frangi filter coupled with the probabilistic patch based denoiser.
IEEE Access 7:164344–164361. [Link]
66. Xu L, Luo S (2010) A novel method for blood vessel detection from retinal images. Biomed Eng Online
9(1):1–10
67. Mun H, Yoon G-J, Song J, Yoon SM (2021) Scalable image decomposition. Neural Comput Appl
33(15):9137–9151
68. Bnou K, Raghay S, Hakim A (2020) A wavelet denoising approach based on unsupervised learning model.
EURASIP Journal on Advances in Signal Processing 2020(1):1–26
69. Upadhyay K, Agrawal M, Vashist P (2021) U-net based multi-level texture suppression for vessel seg-
mentation in low contrast regions. In: 2020 28th European Signal Processing Conference (EUSIPCO),
pp 1304–1308. [Link]
70. Alekseev A, Bobe A (2019) Gabornet: Gabor filters with learnable parameters in deep convolutional
neural network. In: 2019 International conference on engineering and telecommunication (EnT), pp 1–4.
IEEE
71. Bi N, Sun Q, Huang D, Yang Z, Huang J (2007) Robust image watermarking based on multiband wavelets
and empirical mode decomposition. IEEE Trans Image Process 16(8):1956–1966. [Link]
TIP.2007.901206
72. Yiu Q-M, Xie S-L (2005) Arithmetic shift method suitable for vlsi implementation to cdf 9/7 discrete
wavelet transform based on lifting scheme. In: 2005 International Conference on Machine Learning and
Cybernetics, vol 8, pp 5241–52448. [Link]
73. Taubman DS, Marcellin MW (2002) Jpeg 2000: standard for interactive imaging. Proc IEEE 90(8):1336–
1357. [Link]
74. Chattoraj S, Vishwakarma K (2018) Classification of histopathological breast cancer images using iterative
vmd aided zernike moments & textural signatures. arXiv preprint arXiv:1801.04880
75. Dragomiretskiy K, Zosso D (2013) Variational mode decomposition. IEEE Trans Signal Process
62(3):531–544
76. Vincent P, Larochelle H, Lajoie I, Bengio Y, Manzagol P-A, Bottou L (2010) Stacked denoising autoen-
coders: Learning useful representations in a deep network with a local denoising criterion. J Mach Learn
Res 11(12)
77. Wong T-T (2015) Performance evaluation of classification algorithms by k-fold and leave-one-out cross
validation. Pattern Recogn 48(9):2839–2846
78. Jiang Y, Tan N, Peng T, Zhang H (2019) Retinal vessels segmentation based on dilated multi-scale convo-
lutional neural network. IEEE Access 7:76342–76352. [Link]
79. Wang D, Haytham A, Pottenburgh J, Saeedi O, Tao Y (2020) Hard attention net for automatic retinal
vessel segmentation. IEEE J Biomed Health Inform 24(12):3384–3396. [Link]
2020.3002985
80. Khan TM, Khan MA, Rehman NU, Naveed K, Afridi IU, Naqvi SS, Raazak I (2022) Width-wise vessel
bifurcation for improved retinal vessel segmentation. Biomed Signal Process Control 71:103169
81. Mahapatra S, Agrawal S, Mishro PK, Pachori RB (2022) A novel framework for retinal vessel seg-
mentation using optimal improved frangi filter and adaptive weighted spatial fcm. Comput Biol Med
147:105770
82. Liu Y, Shen J, Yang L, Bian G, Yu H (2023) Resdo-unet: a deep residual network for accurate retinal
vessel segmentation from fundus images. Biomed Signal Process Control 79:104087
83. Rehman N, Mandic DP (2009) Empirical mode decomposition for trivariate signals. IEEE Trans Signal
Process 58(3):1059–1068
84. Flandrin P, Rilling G, Goncalves P (2004) Empirical mode decomposition as a filter bank. IEEE Signal
Process Lett 11(2):112–114
85. Linderhed A (2002) 2d empirical mode decompositions in the spirit of image compression. In: Wavelet
and independent component analysis applications IX, vol 4738, pp 1–8. SPIE
86. Huang NE, Shen Z, Long SR, Wu MC, Shih HH, Zheng Q, Yen N-C, Tung CC, Liu HH (1998) The
empirical mode decomposition and the hilbert spectrum for nonlinear and non-stationary time series
analysis. Proceedings of the Royal Society of London. Series A: mathematical, physical and engineering
sciences 454(1971), 903–995
87. Lakshmi MD, Murugan SS, Padmapriya N, Somasekar M (2019) Texture analysis on side scan sonar
images using emd, xcs-lbp and statistical co-occurrence. In: 2019 International symposium on ocean
technology (SYMPOL), pp 91–97. IEEE
88. Maji U, Pal S (2016) Empirical mode decomposition vs. variational mode decomposition on ecg signal
processing: A comparative study. In: 2016 International conference on advances in computing, commu-
nications and informatics (ICACCI), pp 1129–1134. IEEE

123
85898 Multimedia Tools and Applications (2024) 83:85871–85898

89. Tham JY, Shen L, Lee SL, Tan HH (2000) A general approach for analysis and application of discrete
multiwavelet transforms. IEEE Trans Signal Process 48(2):457–464. [Link]
90. Kromka J, Kováč O, Šaliga J (2022) Multiwavelet toolbox for matlab. In: 2022 32nd Inter-
national conference radioelektronika (RADIOELEKTRONIKA), pp 01–05. [Link]
RADIOELEKTRONIKA54537.2022.9764952

Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and
institutional affiliations.

Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under
a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted
manuscript version of this article is solely governed by the terms of such publishing agreement and applicable
law.

Anumeha Varma is a Ph.D. student at CARE, IIT Delhi. She has an M. Tech. and B. Tech. degree in Elec-
tronics and Communication from IIT (ISM), Dhanbad, India (2019) and NIT Raipur, India (2017).

Monika Agrawal is a Professor at CARE, IIT Delhi. She has a Ph.D. in Electrical engineering from IIT Delhi
(2000) and an M. Tech. and B. Tech. in ECE from NIT Kurukshetra, India (1995 and 1993).

123

Common questions

Powered by AI

Despite advances, challenges persist in accounting for the variability in vessel size and ensuring clear distinction from the background due to low contrast and noise. While preprocessing attempts to address these through techniques like CLAHE and filtering, the non-uniform contrast inherent in raw fundus images continues to present obstacles . Furthermore, the complexity introduced by the high dimensionality of methods like Gabor filters can lead to high processing demands, while the effectiveness of individual techniques often hinges on the quality of the initial image input .

The combination of FDM with Gabor transforms is effective as FDM provides a robust framework for spectrum and time-frequency analysis by decomposing signals into Fourier intrinsic band functions (FIBFs), satisfying orthogonal and zero-mean conditions that are ideal for handling nonlinear signals . Gabor filters enhance this by extracting detailed frequency and directional information, crucial for identifying vessel-like structures. Together, they balance detailed feature extraction with efficient signal decomposition, leading to accurate segmentation outcomes .

Image decomposition techniques excel in unsupervised learning by explicitly separating the structural and textural components of an image. This separation enables the focus on desirable features, such as edges and blood vessels, by simplifying complex details into more manageable sub-images. Wavelet transforms, for example, isolate low-frequency approximation components from high-frequency detail components, thus accentuating features pertinent to vessel visualization . This aids in direct analysis and accurate segmentation of these features without extensive pre-labeling.

Unsupervised learning approaches, like using image decomposition techniques, can focus on separating structural and textural parts of an image efficiently, requiring less time and resources than supervised methods. This approach is nearly on par with the performance of supervised methods and is less resource-intensive while providing flexibility in feature extraction such as edges and blood vessels . This makes it suitable for large-scale or resource-constrained applications where supervised methods may not be practical.

Preprocessing steps in the study begin with green channel extraction (GE), exploiting the superior contrast of the green channel over red and blue. This is followed by contrast-limited adaptive histogram equalisation (CLAHE), which enhances image contrast. Noise is reduced through Gaussian filtering, acting as a low-pass filter, and further refinements are made using median filters to tackle salt-and-pepper noise, though with careful adaptation to preserve details . Collectively, these steps ensure images are well-prepared for accurate and efficient segmentation, counteracting initial limitations of raw image data.

The high-performing segmentation method proposed could be expanded to retinal scan-based biometric authentication, retinal disease diagnosis, and other areas requiring real-time performance analysis. Further enhancements could include the introduction of 2D Fourier Decomposition Method (FDM) to improve accuracy. Additionally, the method's adaptability to handle variability in parameter settings could optimize it for different datasets, allowing for effective performance in low-resource settings or non-professional image captures .

Preprocessing techniques enhance the contrast and signal-to-noise ratio (SNR) of retinal images, making it easier to distinguish vessels from the background. Techniques like green channel extraction and contrast-limited adaptive histogram equalisation (CLAHE) improve visibility by utilizing the green channel's better contrast compared to the red and blue channels, which are typically blurry and noisy . Filtering methods such as Gaussian filters reduce noise, aiding in clearer image analysis . These steps ensure better results in subsequent segmentation processes.

Supervised learning methods, while capable of producing high-performance metrics in retinal vessel segmentation, are resource-intensive, demanding substantial time and computational resources . They typically require large annotated datasets for effective training, necessitating significant efforts in data preparation. Although these methods provide high accuracy and adaptability to diverse image types, the logistical burden associated with resource and data requirements can limit their practical applicability, especially in large-scale or real-time diagnostic scenarios .

Gabor filters are beneficial for retinal vessel segmentation because they excel in feature extraction from images, especially for identifying directional edges such as blood vessels. They efficiently capture both the frequency and orientation information through complex modulated Gaussian profiles, making them optimal for the task . However, Gabor filters suffer from high dimensionality and extensive processing requirements, which can be limiting factors in terms of computational demand .

The proposed FDM + Gabor method demonstrates superior sensitivity and resource efficiency compared to existing state-of-the-art methods. It achieves higher sensitivity across multiple datasets, including DRIVE, offering practical implementations in scenarios requiring resource conservation, such as bulk data processing . While the proposed method aligns closely with state-of-the-art methods in accuracy, its time- and resource-saving advantages make it more applicable for large-scale operations or scenarios with limited computational resources, highlighting its practical utility in medical diagnostics .

You might also like