0% found this document useful (0 votes)
6 views14 pages

Missing Modality Paper

The document presents a novel approach for brain tumor segmentation using an Edge-aware Discriminative Feature Fusion Based Transformer U-Net (EA-DFFTU-Net) that effectively handles missing MRI modalities. The method combines local feature extraction via a ResNet-50 encoder with global context learning through a Transformer network, utilizing an Edge Feature Module to enhance edge representations. The proposed model aims to improve segmentation accuracy compared to existing techniques by integrating multi-scale features and addressing the challenges posed by absent modalities.

Uploaded by

dr.afzalsitara
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views14 pages

Missing Modality Paper

The document presents a novel approach for brain tumor segmentation using an Edge-aware Discriminative Feature Fusion Based Transformer U-Net (EA-DFFTU-Net) that effectively handles missing MRI modalities. The method combines local feature extraction via a ResNet-50 encoder with global context learning through a Transformer network, utilizing an Edge Feature Module to enhance edge representations. The proposed model aims to improve segmentation accuracy compared to existing techniques by integrating multi-scale features and addressing the challenges posed by absent modalities.

Uploaded by

dr.afzalsitara
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Applied Soft Computing Journal 161 (2024) 111709

Contents lists available at ScienceDirect

Applied Soft Computing


journal homepage: [Link]/locate/asoc

Brain tumor segmentation with missing MRI modalities using edge aware
discriminative feature fusion based transformer U-net
B. Jagadeesh a, G. Anand Kumar b, *
a
Professor, Department of Electronics and Communication Engineering, Anil Neerukonda Institute of Technology and Sciences (Autonomous), Visakhapatnam, India
b
Associate Professor, Department of Cyber Security & IoT, School of Engineering, Malla Reddy University, Hyderabad, India

H I G H L I G H T S

• Features are extracted using ResNet-50 encoder.


• Vision transformer is used to capture global contexts.
• EFM is used to acquire edge attention representations.
• DFFM is used to effectively fuse the multiscale features.

A R T I C L E I N F O A B S T R A C T

Keywords: Brain tumor segmentation is an essential task for medical diagnosis and treatment planning. Multi-modal MRI
Brain tumor provides complementary information that is essential for accurate segmentation of brain tumors, but missing
Transformer network modality images are a common problem in clinical practice. Existing segmentation methods often fail to generate
Missing modalities
accurate object boundaries and selectively fuse the tumor region, resulting in unreliable segmentation masks. In
Feature extraction
this work, we propose an Edge-aware Discriminative Feature Fusion Based Transformer U-Net (EA-DFFTU-Net)
Segmentation
to segment the brain tumor effectively even in the absence of modalities. First, the MRI input data is pre-
processed, and features are then extracted using a ResNet-50 encoder, which learns local spatial. We employ
an Edge Feature Module (EFM) to acquire edge attention representations. These features are then transferred to
the Discriminative Feature Fusion based Transformer U-Net (DFFTU-Net) that learns global contexts and captures
multiscale features. We use a Discriminative Feature Fusion Module (DFFM) in the decoder of the DFFTU-Net to
effectively fuse the multiscale features in order to obtain accurate segmentation. The performance of the pro­
posed EA-DFFTU-Net segmentation method was determined by evaluating it and comparing the obtained results
with those of existing brain tumor segmentation techniques.

1. Introduction various glioma subregions [3,4]. The tumour core, also known as the
tumour without peritumoral edoema, is highlighted on T1 and T1c. T2
The brain tumor is recognized as one of the most dangerous and and FLAIR emphasise the complete tumour, which is the tumour with
deadly diseases globally, and it is crucial to detect brain tumor earlier for peritumoral edoema. T1c also shows an enhancing tumour core, which
clinical evaluation and treatment planning [1]. A common imaging is an area of the tumour core with hyperintensity. The time, availability
method for evaluating brain tumors is magnetic resonance imaging of the scanners, cost, and patient comfort prevent most patients from
(MRI) because without using radiation, it provides a strong soft tissue having access to the entire variety of multi-modality MRI scans [5].
contrast [2]. There are multiple MRI modalities, namely T2-weighted, Despite numerous techniques being developed for brain tumor seg­
T1-weighted, contrast-enhanced T1-weighted (T1c), and FLAIR (Flu­ mentation, the process is still difficult and challenging especially when
id-Attenuation Inversion Recovery), based on various contrasts and certain modalities are absent. This is because gliomas vary for patients in
functional aspects. Different imaging modalities reveal various tissue size, shape, and texture, variable intensity range, and poor contrast in
architectures and provide additional information for the analysis of MRI [6]. The process of manually segmenting brain tumors by experts is

* Corresponding author.
E-mail addresses: bjagadeesh76@[Link] (B. Jagadeesh), anandlife@[Link] (G. Anand Kumar).

[Link]
Received 27 April 2023; Received in revised form 13 March 2024; Accepted 26 April 2024
Available online 10 May 2024
1568-4946/© 2024 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

both expensive and time-consuming [7]. One of the most essential • Discriminative feature fusing Transformer U-Net (DFFTU-Net) is
components for attaining accurate segmentation is edge information. introduced to capture the global contexts and segment the tumor
Existing medical image segmentation approaches use local gradient region in which we employ a DFFM to find the most informative
representations to identify the edge regions and the obtained closed loop features and effectively fuse them.
area is detected as objects [8,9]. These methods are accurate only for • The EADFFTU-Net approach is shown to have better performance
simple structural images. The lack of object-level information and reli­ than existing approaches.
ance on local edge representations in various edge detection algorithms
results in trivial segmentation areas and discontinuous borders [10]. The paper’s remaining sections are outlined as follows:
In medical image, segmentation CNNs have administered a variety of Section 2 offers a literature review, Section 3 describes the proposed
applications since the emergence of deep learning [11]. The conven­ methodology, Section 4 provides the simulation setup and the compar­
tional encoder-decoder-based network U-Net has shown excellent seg­ ative analysis and Section 5 presents the conclusion.
mentation potential among the different CNN variants [12]. Improved
performance in medical image segmentation has been observed with the 2. Related work
introduction of models such as 3D UNet [13], Res-UNet [14], and
UNet++ [15]. CNNs typically increase the receptive field size as the Below are a few recent publications that focus on brain tumor
network depth increases. Attention mechanisms were also introduced segmentation.
[16,17]. However, this strategy has limited efficacy in capturing global A strategy for segmenting brain tumours with missing modalities was
relationships. This causes problems such as loss of global context and not put forth by Zhou et al. [21]. To overcome the problem due to missing
developing long-range dependencies. It also makes the model more modalities, they used a multi-modal segmentation network on the basis
complicated and causes overfitting [18]. of U-Net architecture which integrates fusion block and correlation
Transformer, [19] which successfully enhances network perfor­ model. The latent correlation representations (LCR) obtained from the
mance in the Natural Language Processing (NLP) space, has recently correlation model are fused by the fusion block that the correlation
proved valuable in computer vision. The Vision Transformer (ViT) al­ model discovers between modalities through the attention mechanism
gorithm segments an image into multiple patches, making them com­ which is then decoded to get the segmentation result. However, this
parable. and then the Transformer’s self-attention mechanism is used on approach did not take into account the edge information that is crucial
each patch to capture the long-range dependencies. Limitations that are for precise segmentation.
caused by the networks using convolutional layers are overcome by ViT Yu et al. [22] proposed SA-LuT-Net which is built around backbone
[20]. There are few studies using Transformers for medical image seg­ networks such as the 3D Unet model and DMFNet. A LuT module and a
mentation, but they neglect the local representations which are essential segmentation module make up the network. Propagation of the LuT
for refining the tumor details as they use a transformer encoder to module with the segmentation module learns an intensity mapping
extract features. Also, the existing methods are not robust to missing function. This method gives better segmentation results but it is
modalities and they mainly focus on the primary region and don’t time-consuming.
consider the edge information, which leads to inaccurate segmentation Wu et al. [23] introduced a generative adversarial network that is
of tumor regions. Additionally, the most essential features of several symmetric driven (SD-GAN). To represent the variation (symmetry) of
modalities that are very pertinent to brain tumors are not fused. So, the normal brains the network was trained to learn a non-linear mapping
features of the tumor region may not get fused properly and produce of brain images. Then the trained model is used in normal brain
inaccurate results. Moreover, it cannot capture long-range correlations. reconstruction and brain tumor segmentation based on increased
Some drawbacks of the existing segmentation approaches include in­ reconstruction errors caused by their asymmetry. This method is hard to
accuracy, overfitting, and low significant performance. To overcome train.
this issue Edge Aware Discriminative Feature Fusion-based Transformer Yan et al. [24] proposed SEResU-Net, an improved U-Net model that
U-Net (EADFFTU-Net) is proposed. links deep ResNet and the SENet. The network extracts additional
Three phases, namely pre-processing, feature extraction, and seg­ feature information and overcomes network degradation. It also pre­
mentation, make up the proposed model. In the phase of pre-processing, vents information loss. It also resolves the challenge of tiny brain tu­
the input data is pre-processed using intensity normalization and N4ITK. mour’s low segmentation accuracy but failed to consider edge
The intensity variations in images are normalized using intensity information and not robust to missing modalities.
normalization and the distortion in MR data is corrected using N4ITK. Huang et al. [25] put forward GCAUNet, which may effectively
After pre-processing, the multi-level features are extracted from the pre- utilise the low-level fine information of tumour areas to enhance the
processed data using the ResNet50 encoder. Here, each input modality is segmentation of brain tumors. They introduced a detail recovering path
processed by an individual encoder to extract features. We develop an (DR) and multiscale input path (MI) to retrieve the precise features of
EFM, which is placed after each encoder block to learn and preserve the brain tumours and to obtain multiscale context data. Moreover, GCA is
edge characteristics of each modality. The obtained multi-level features used to draw attention to the important feature groups and channels.
are fed into the dual path encoder of the proposed DFFTU-Net to capture This method solves the problem of feature loss but is not robust to
global contexts and multiscale features. Effective fusion of multi-scale missing modalities.
features from different modalities is achieved and unwanted features The network’s performance is constrained by employing just 3D
are discarded using the proposed DFFM. Finally, decoding the fused convolution as a module of feature extraction. By considering this issue
features yields the final segmentation result. Xiao et al. [26] proposed a MVHS-Net. Modified 3D UNet introduces the
The proposed EADFFTU-Net framework is characterized by the MVHS block and multi-view fusion convolution block as its structural
following key contributions: components. To gather multi-scale as well as multi-view information,
these modules are introduced. This method decreases the amount of
• To achieve the best segmentation result even in the case of missing irrelevant character information but is not robust to missing modalities.
modalities EADFFTU-Net model is proposed that learns both local Zhang et al. [27] proposed an MSMANet that developed an enhanced
features and global contexts. reception module to gather and extract useful information from multiple
• A new module called the Edge Feature Module (EFM) is proposed to receptive fields. MA technique is utilized in place of the U-Net’s skip
learn edge representations while preserving regional edge connection. This gradually reduces the semantic gap and enhances the
properties. shallow features. Various scale multi-level feature aggregation is
increased in this approach. In order to enhance network identification

2
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

and convergence capabilities, MSMANet additionally uses an attention useful life (RUL) prediction in machinery health assessment. The
technique. This method aims only at the primary areas and discards the method extracts global features from multiple sensors and utilizes
edge information. adversarial learning for generalized feature extraction. Li et al. [37]
Liew et al. [28] introduced the CASPIAN network that uses a noisy introduced the first attempt to use event-based vision data for machine
student curriculum learning paradigm. A multidimensional approach is fault diagnosis. They proposed a vibration event representation and
adapted that integrates multiplanar and multiscale information to employed a deep convolutional neural network model for processing.
improve spatial information inside the network. The proposed Additionally, they introduced an event data augmentation method and a
CASPIANET++ combines channel attention with multiplanar and mul­ deep representation clustering method. Table 1 presents an overview of
tiscale spatial attention for accurate segmentation but it is the discussed works.
time-consuming and not robust to missing modalities. The majority of currently used techniques rely on a single modality
Rai et al. [29] proposed UnetResNext-50 (hybrid deep CNN model) and are therefore not effective when a modality is absent. Also, it mainly
with a lot of parameters and layers for accurate brain tumor segmen­ focuses on the primary region and did not consider the edge informa­
tation. Unet and ResNext are combined to form UnetResNext-50 in tion, which leads to inaccurate segmentation of tumor regions. It also
which U-net serves as the backbone and it uses the ResNext-50 encoder. fails to fuse the most essential features of several modalities that are
This technique solves the gradient degradation problem by employing a highly pertinent to brain tumors. So, features other than tumor region
cardinality skip connection but it did not consider edge information. may get fused and produce inaccurate results. Moreover, the use of CNN
Tan et al. [30] proposed ACU-Net, a technique for multimodal brain for segmentation cannot capture long-range correlations. Some draw­
tumor segmentation. Here, the deep separable convolutional layers are backs of the existing segmentation approaches include inaccuracy,
used instead of U-Net in which the mapped convolutional channel’s overfitting, and low significant performance. To overcome these prob­
appearance and spatial correlation are differentiated. To increase the lems EADFFTU-Net is proposed.
capacity of feature propagation and convergence speed residual skip
connections have been included. To improve the segmentation accuracy, 3. Proposed EADFFTU-Net methodology
they insert the active contour model, that ensures proper fitting between
boundary’s inside and outside. This method is computationally efficient, This paper proposes Edge Aware Discriminative Feature Fusion
but it doesn’t give accurate segmentation results. based Transformer U-Net (EADFFTU-Net) for brain tumor segmentation
Syazwany et al. [31] proposed MM-BiFPN to overcome the problems in the absence of MRI modalities. Fig. 1 shows the architectural layout of
in existing methods that ignore the interdependence of the various the proposed method. The framework of the proposed model carries
modalities. The features of each modality are extracted by specific en­ three phases namely pre-processing, feature extraction, and segmenta­
coders to learn the complex relationship among them. The segmentation tion. At the phase of pre-processing, intensity normalization and N4ITK
accuracy is enhanced by the use of the BiFPN layer. This is achieved by are used to pre-process the input data. The intensity variations in images
learning the multiscale features and cross-modality relationships. This are normalized using intensity normalization and the distortion in MR
method works better with missing modalities but discards the edge data is corrected using N4ITK. After pre-processing, multi-level features
features and has high computational overhead. that includes local features are extracted from the pre-processed data
Luo et al. [32] proposed a light-weight HDC-Net, a pseudo-3D using the ResNet50 encoder. Each input modality undergoes feature
network with the architecture of 3D U-Net in which HDC modules are extraction using the individual encoder here. We develop an EFM, which
used in place of 3D convolution modules to extract multiscale and is placed after each encoder block to learn and preserve the edge char­
Multiview spatial features more efficiently. In the first stage, periodic acteristics of each modality. The obtained multi-level features are fed
down-shuffling is used as downsampling because it maintains the entire into the dual path encoder of the proposed DFFTU-Net to capture global
information of the image and is parameter-free. View decoupled contexts and multiscale features With the proposed DFFM, multi-scale
convolution in the spatial domain reduce computational complexity but features from various modalities are fused effectively while discarding
it discards the edge information. unwanted features. Finally, decoding the fused features yields the final
Micallef et al. [33] proposed a modified U-Net++ model. In terms of segmentation result.
loss function, deep supervision employing method and number of con­
volutional blocks it varies from the U-Net++ model. This network model 3.1. Pre-processing
has five depth levels the last of which connects the network’s decoder.
Instance normalization and a filter map resolution of 16 are also Pre-processing entails adding relevant and helpful information to the
included. This method is not robust to missing modalities. dataset and removing irrelevant data. In the realm of computer vision
Aghalari et al. [34] proposed an improved U-Net based architecture. pre-processing is essential, particularly when analysing medical images,
Three TPR-based UNet models such as TPR-Encoder-Unet, TPR-Deco­ since erroneous data or undesired information might cause a highly
der-UNet and TPR-Encoder-Decoder-UNet were introduced to enhance effective algorithm to function poorly. Initially, we used N4ITK, and
the effectiveness of segmentation. Initially, TPR blocks are added to the then intensity normalization was performed to pre-process the input
pipeline for down sampling, which is called TPRE-UNet. Then, the initial data. Correcting for bias-field inhomogeneities is essential for improving
method is repeated for up-sampling to create the TPRD-UNet model. the quality of MRI images, as the smooth, low-frequency signal can
Finally, PR blocks are added on both UNet paths, and a model called degrade them. To achieve this, the N4ITK [38] algorithm is used. The
TPRED-UNet (TPR-Encoder-Decoder-UNet) is created. This model process is detailed as follows:
exploit both local and global features simultaneously but it is not robust The initial step involves utilizing the image formation model, which
to missing modalities and not considered edge information. is provided in Eq. (1)
Hao et al. [35] introduced a generalized pooling method (GP) by
B(x) = m(x)e(x) + n(x) (1)
fusing average pooling and maximum pooling with adaptive weights for
the segmentation of brain tumors. This study aims at making the pooling where B corresponds to the observed image, m represents the uncor­
operations better. This eliminates the requirement of choosing average rupted image, e stands for the bias field, and n signifies noise, assumed to
or maximum pooling for CNN model down-sampling. GP creates pooling conform to a Gaussian distribution and exhibit independence. Upon
using a variety of expressions on the basis of input images or feature
introducing the notation m ̂ = log m and working within a noise-free
maps. This method solves the problem in conventional pooling methods
setting, the model adopts the subsequent expression as in Eq. (2):
but it does not preserve all the spatial information.
Li et al. [36] proposed a deep learning-based method for remaining

3
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

Table 1
Overview of the reviewed works.
Method Year Pre-processing Deep learning model Dataset Metrics
for segmentation

Modelling the latent multi source 2021 N41TK and intensity U-Net [19] BraTS 2018 and 2019 Dice score, Sensitivity, Hausdroff
correlation and fusion [21] normalization dataset Distance (HD),
SA-LuT-Net [22] 2021 Zero-mean, Unit Standard SA-LuT-Net BraTS 2018 and 2019 Dice score, Hausdorff95
Deviation (STD)normalization dataset
SD-GAN [23] 2021 - SD-GAN BRATS 2012 and 2018 precision, sensitivity, Dice score
SEResU-Net [24] 2022 Data standardization SEResU-Net BraTS 2018 and 2019 Dice similarity coefficient (DSC),
dataset specificity, sensitivity, HD
GCAUNet [25] 2021 Iterative thresholding, zero- GCAUNet BraTS 2017 and BraTS DSC, sensitivity, HD
score normalization 2018 dataset
MVHS-Net [26] 2021 - MVHS-Net BraTS 2018 dataset DSC, Hausdorff95
MSMANET [27] 2021 standardization MSMANET BraTS 2018 dataset DSC, sensitivity, specificity,
Hausdorff95
CASPIANET++ [28] 2021 - CASPIANET++ BraTS 2019 and BraTS Dice score, HD
2020 dataset
UnetResext-50 [29] 2021 Resizing, global pixel UnetResext-50 The Cancer Genome Dice score, Jaccard index, F1-score,
normalization, augmentation Atlas (TCGA) dataset precision, recall and accuracy
ACU-Net [30] 2021 - ACU-Net BraTS 2015, BraTS Dice coefficient, Intersection of Union
2018, BraTS 2019 (IoU), precision, recall, Hausdorff95
dataset
MM-BiFPN [31] 2021 - MM-BiFPN MICCAI BraTS2018 and Dice coefficient, sensitivity, positive
MICCAI BraTS2020. predicted value (PPV), Jaccard
similarity index
HDC-Net [32] 2020 - HDC-Net BraTS 2018 and BraTS Dice sore, Hausdorff95
2017 dataset
U-Net++ [33] 2021 Intensity normalization U-Net++ BraTS 2019 dataset Dice coefficient, HD
Modified UNet [34] 2021 Image normalization TPRED-UNet BraTS 2018 dataset Accuracy, dice similarity coefficient,
sensitivity, specificity, PPV
Generalized Pooling [35] 2021 - Unet, FCN8, UNet++ GP- BraTS 2018 and BraTS Dice similarity coefficient, PPV,
CNN 2019 dataset sensitivity

Fig. 1. The overall architecture of the proposed EADFFTU-Net.

4
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

b(x) = m(x)
̂ ̂ + e(x) (2)
In accordance with Tustison, the iterative methodology of N4ITK is
expressed in Eq. (3):
n− 1 n− 1
H ∗ {m
̂ − E[ m|
̂ m̂ ]}
̂n = m
m ̂ n− 1 − (3)
̂e 2r

The smoothing factor is represented by H∗. An alternative iterative


scheme is introduced for achieving convergence, where the estimation
2
of the total bias field e approaches zero ( ̂e r →0). This insight into the
nested structure guides the computation of the complete bias field. Eq.
(4) provides the computation of the complete bias field.
n

̂n = ̂
m b− ̂e ir (4)

Consequently, the total bias field approximation during the nth


iteration emerges as the outcome of adding together the initial n residual
bias fields as expressed in Eq. (5).
n

̂e ne = ̂e ir (5)
i=1

Large intensity variations that occur due to the use of various scan­
ners and examination parameters at various times lead to performance
degradation. Hence, to normalize the intensity variations intensity
normalization is used. This is performed by subtracting the gray value of Fig. 2. Feature extraction using Resnet50 with EFM.
the highest frequency from each voxel’s intensity and then dividing the
resulting deviation. This newly adjusted deviation, denoted as ̃ σ , applies (7)
to our MR image denoted as D, which consists of a set of voxels {d1 ,d2 ,d3 ,
…, dN }. The intensity value of each voxel, dk , is represented as Tk . The Fmʹ = F(Fm ; θm ), m ∈ {1, 2, 3, 4} (7)
calculation of ̃ σ is achieved using the Eq. (6)
where, Fmʹ represents the reduced feature maps of each encoder
√̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅
/ ̅
√ N block, F denotes 1 × 1 convolution function, and θm is the respective
√∑
σ=
̃ √ (Tk − T) ̂ 2
N (6) parameter. Edge features of varying granularities can be obtained by
k=1 subtracting average pooling values of various sizes from its local con­
volutional feature maps. We specify P pooling operations in order to
where T̂ signifies the gray value of the highest frequency. Additionally, maintain generality as given in equation (8).
to enable the MR images to undergo processing similar to regular im­ ( )
Fm,l = Fʹm − avg p Fʹm , p ∈ {1, …, P} (8)
(p)
ages, we linearly transform their intensity range to span from 0 to 255.
th th
where F(p)
m signifies the edge features of the current m stage with the p
3.2. Feature extraction pooling operation, andavg p is the relevant average pooling operation.
The edge features are integrated by concatenating them with the fea­
Local features are robust to image variations and also helps in dis­ tures of the current stage, followed by merging them with a convolution
tinguishing the tumor and non tumor region. With the ability to extract operation. This is expressed in Eq. (9).
both high and low-level features, it can learn more complex image ( ( ) )
Fm,l = F C F(1)
m,l , …, F m,l , F m ; θm,l
(P) ʹ
(9)
representations, the Resnet50 encoder is employed for the feature
extraction process.
where, Fm,l represents the feature maps which is the output of the
Fig. 2 shows the feature extraction process using Resnet50 with EFM.
EFM at the encoder network’s present stage, C denotes the concatena­
The multi-level features from the pre-processed data are extracted using
tion process, θm,l refers to the corresponding parameter. The design of
ResNet50 encoder. We adopt individual encoders for each modality
extracting edge features offers an effective means of improving the
rather than fusing it to be robust in case of missing modalities. This
representation capability of the relevant level. By integrating edge in­
design explicitly takes into account the relationships between the
formation at multiple granularities, the edge features are improved. The
different modalities and specifically extracts features from each mo­
output maps are subsequently processed by the multi-task module to
dality. We employ EFM after each encoder block to learn and preserve
enhance the extraction of finer features.
the edge characteristics of each modality.

3.2.1. Edge feature module 3.3. Segmentation


Edge information give essential constraints to direct feature extrac­
tion for segmentation. We propose EFM, an efficient and simple feature With ResNet50 and EFM employed for feature extraction, DFFTU-
extraction strategy for extracting edge features in each step of the Net is capable of utilizing the extracted multi-level features. For seg­
encoder network in order to acquire a reliable edge information. Fig. 3 mentation, the proposed DFFTU-Net uses a dual encoder structure that
shows the structure of proposed EFM. It also offers an edge-attention captures multi-scale information from the extracted multi-level feature
representation to direct the segmentation process. At first, the features map, where both encoders are fed with different scale features. The
of each modality Fm , m ∈ {1, 2, 3, 4} is extracted from the respective proposed DFFM aims to identify regions relevant to brain tumors by
encoder block and are fed as input to the EFM. And to squeeze the extracting the most significant features from various modalities. The
feature maps Fm , we use 1 × 1 convolution function. It is defined in Eq. decoder of DFFTU-Net comprises three upsampling stages. Moreover,

5
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

Fig. 3. : Structure of proposed EFM.

each stage includes two Swin Transformer blocks followed by a patch- window-based multi-head self-attention (SW-MSA). Non-overlapping
expanding module to acquire a large feature map. blocks of a specified size are used to divide the input volume in
The encoder has a Swin Transformer block and a patch merging W-MSA, while SW-MSA connects unrelated adjacent blocks using a
module. We use Swin Transformer block [39] because it incorporates a shifted window mechanism.
cleverly designed shifted window operation that sets it apart from other The feature map generated from the resnet50 is converted into non-
transformer-based models. Fig. 4 shows the internal connections of swin overlapping patches. To be more specific, we use patch size ps = 4 to
transformer block. Within the Swin Transformer Block, there are several obtain small-scale feature (fss ) and use ps = 8 those results in a high-
components, including a LayerNorm layer for normalization, residual resolution feature which is large-scale (fls ). Next, two separate encoder
connection, a MSA (multi-head self-attention module) and MLPs of two branches with several down sampling layers are fed with fss and fls . The
layer with activation functions. The MSA module in the swin trans­ feature dimension is increased at each step while the feature map’s
former Block contains two types of attention mechanisms: resolution is decreased through the use of a Swin Transformer layer and
window-based multi-head self-attention (W-MSA) and shifted a patch merging module [16]. The input patch is partitioned into four

Fig. 4. Illustration of proposed DFFTU-Net.

6
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

segments, which are then combined using the patch merging module. leading to more accurate segmentation results, especially in complex
Thus, the multi-scale features are able to be represented hierarchically. medical imaging scenarios. The proposed EADFFTU-Net offers a holistic
During the first three stages of downsampling, the large-scale branch approach to brain tumor segmentation by combining advanced feature
produces outputs with resolutions that are twice as large as those of the extraction techniques, discriminative feature fusion, and state-of-the-art
small-scale branch. In comparison to the small-scale branch, the Transformer architecture. Compared to traditional methods that may
large-scale branch has an additional downsampling stage. rely solely on handcrafted features or shallow learning architectures, our
The decoder of DFFTU-Net comprises three upsampling stages. method achieves superior performance in terms of segmentation accu­
Moreover, each stage includes Swin Transformer Block and a Patch racy and robustness.
Expanding module to acquire a large feature map. The output from the The multi-scale features captured by the dual encoders are selec­
last Swin Transformer Block of the encoder is fed into the patch tively fused by DFM and then the fused features are transmitted into the
expanding module. The patch expanding module [37] has a rearrange decoder via skip connections.
operation and a linear layer in which the feature resolution is expanded The proposed DFFM consists of spatial as well as channel attention
by a rearrange operation, while the feature dimension is doubled by a followed by feature fusion as depicted in Fig. 5. The two features from
linear layer. the double encoder branch are denoted as Fl ∈ EC×(H×W) , and
Fs ∈ EC×(H/2×W/2) . The unsampled decoder feature is represented as
3.3.1. Discriminative feature fusion module Fde ∈ EC×(H×W) . We initially perform a transposed convolution layer on
We propose a Discriminative Feature Fusion Module, that integrate
Fs ∈ EC×(H×W) due to its poor resolution in order to produce the upsam­
the multi-scale features effectively. The proposed method leverages
pled feature FUs ∈ EC×(H×W) . We concatenate Fde with FUs and Fl separately
discriminative feature fusion to effectively integrate multi-scale features
during the spatial attention phase. This allows us to create two spatial
from different modalities. Unlike traditional methods that may rely on
attention maps A1 ∈ EH×W and A2 ∈ EH×W by feeding the concatenated
simple fusion techniques like concatenation or summation, our method
features into point-wise convolution method. Then, the spatial attention
selectively fuses features using spatial and channel attention mecha­
maps A1 and A2 are scaled into [0,1] using the sigmoid function. Spatial-
nisms. This allows the model to focus on relevant features while dis­
wise multiplication is performed between spatial attention maps and the
carding irrelevant ones, leading to more accurate segmentation results.
input features for minimizing the noisy features while escalating rele­
One of the major benefits of our method is its ability to preserve edge
filt
characteristics during feature fusion. The Edge Feature Module (EFM) is vant ones. The filtered features Fl and Ffilt s are subjected to global
specifically designed to capture edge features at different granularities average pooling during the channel attention phase to extract channel
and integrate them into the segmentation process. This ensures that the context descriptors Dl ∈ EC×1 and Ds ∈ EC×1 . Following that, we merge
model can accurately delineate tumor boundaries, which is crucial for them as Fd = [Dl , Ds ] ∈ E2C×1 and apply a Softmax function to the logits
precise segmentation. Traditional methods may struggle when faced across each channel as given in Eq. (10).
with missing modalities or incomplete data. Our method addresses this j j
challenge by employing individual encoders for each modality, allowing eDl eDs
ajm = j , bjm = j (10)
the model to extract features independently from each input. This design
j j
Dl
e +e Ds Dl
e + eDs
choice ensures robustness to missing modalities and enables the model
where ajm and bjm correspond to the jth element of their respective
to effectively utilize whatever information is available.
By incorporating Transformer architecture, particularly the Swin channel attention maps. Here we have ajm + bjm = [Link] jth element of Dl
Transformer blocks, our method benefits from their ability to capture and Ds are denoted by Djl and Djs . By combining the features from both
long-range dependencies and global context information. This enhances branches, we obtain an informative feature which is expressed in Eq.
the model’s understanding of spatial relationships within the input data, (11).

Fig. 5. Illustration of proposed DFFM.

7
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

F = am .Ffilt volumes, including skull stripping, co-registration to a common


l + bm .F s
filt
(11)
anatomical template, and interpolation to a resolution of 1mm3 .
where F represents the fused feature map. Finally, is concatenated and
fed as input to the next upsampling stage. 4.2. Evaluation metrics

This study evaluated the proposed method’s effectiveness using


3.4. Loss function
several metrics, including the specificity, Dice Similarity Coefficient
(DSC), Hausdroff Distance and sensitivity which are widely regarded as
To facilitate network training, the following overall loss function is
trustworthy metrics for brain tumor segmentation. The calculation for­
employed. It is expressed in Eq. (12).
mulas for these metrics are provided below, with T +ve , F+ve , T − ve and F− ve
L = Ld + λLr (12) representing correctly identified positive sample regions, incorrectly
identified positive sample regions, and incorrectly identified negative
For our experiment, we fix the trade-off parameter λ to 1, which
sample regions, respectively.
balances the weight of each component. The dice loss function is
a) DSC
employed to evaluate the segmentation performance by calculating the
DSC is a metric used to assess how well the segmentation prediction
overlap among the results of predicted and ground truth. Dice loss is
resembles the ground truth.
expressed in Eq. (13),
∑P ∑Q 2T+ve
DSC = (15)
j=1 uij vij + e
.
i=1
Ld = 1 − 2 ∑P ∑Q (13) F+ve + 2T+ve + F− ve
i=1 j=1 uij + vij + e
b) Sensitivity:
where P represents the collection of all examples, while Q denotes the Sensitivity, also known as the T+ve or recall rate, that measures how
collection of classes. The presumption that pixel i corresponds to tumour well the segmentation experiment can identify the region of interest. It
class j is represented by the value of uij ,while represents the actual counts the number of positive voxels in the actual backdrop.
probability by adding a small constant, which prevents division by zero. T+ve
To prevent division by zerovij uses a tiny constant called e to represent sensitivity = . (16)
T +ve+ F− ve
the real probability. We do reconstruction by aligning each recon­
structed image with its associated input image using Mean Absolute c) Specificity:
Error (MAE) as expressed in Eq. (14) Specificity, also referred to as the T− ve , is a measure of the ability of
m
the segmentation experiment to accurately distinguish pixels outside the

Lr = MAE‖Hi (f) − yi ‖1 (14) region of interest. This is accomplished by measuring the ground truth
i=1 segmentation’s negative voxel fraction.
In this task, mindicates the number of modalities, that is set to 4. the T +ve
reconstruction path is denoted by R, f denotes fused representation by specificity = . (17)
T+ve + F+ve
the last-layer fusion block, and yi denotes the input modality.
d) HD
4. Data and methods In the context of segmentation evaluation, HD is employed to gauge
the correlation among two sets of points by determining the great
4.1. Dataset variation among the predicted and ground truth segmentation results.
[ ] { [ ] }
We employed two fundamental datasets in our study, such as the HD = max dAB , dBA = max maxmind a, b , ×maxmin d[a, b]
a∈A b∈B b∈B a∈A
BRATS 2019 and BRATS 2020 datasets. The BRATS 2019 dataset com­
(18)
prises a training set encompassing 335 cases and a validation set
featuring 125 cases, all with concealed ground-truth data. In both sets, Smaller values of HD correspond to higher segmentation accuracy.
every case is equipped with four distinct image modalities: T1, T1c, T2,
and FLAIR. In accordance with the challenge’s guidelines, the four 5. Experiment results
intratumor structures such as edema, enhancing tumor, necrotic, and
non-enhancing tumor core have been consolidated into three mutually In this study, we conduct a series of comparative experiments to
inclusive tumor regions: (a) The whole tumor region (WT) that showcase the effectiveness of our proposed methodology while drawing
Encompassing all tumor tissues. (b) The tumor core region (TC) comparisons with alternative approaches. We initiate a set of experi­
comprising the enhancing tumor, necrotic and non-enhancing tumor ments aimed at assessing the significance of our proposed components,
core. (c) The enhancing tumor region (ET). providing concrete evidence that their incorporation results in improved
The BRATS 2020 dataset, on the other hand, is characterized by a segmentation performance. Moving to next section, we engage in a
training set with 369 cases and a validation set containing 125 cases. We thorough comparison between our method and cutting-edge techniques
had to rely solely on the training data in our study due to the absence of across all modalities. Furthermore, we dedicated to demonstrate the
ground truth images in the validation dataset. Our model’s performance proposed model’s resilience to missing data. Furthermore, we establish
was evaluated using a five-fold cross-validation methodology, with 240 that it outperforms the state-of-the-art approaches. Lastly, we delve into
cases allocated for training and 55 cases reserved for testing purposes. qualitative experiment results, reaffirming that our method yields
Within the dataset, there are 369 cases. Each case includes data from promising segmentation outcomes across both complete and incomplete
four MRI modalities: T2, T1ce, FLAIR, and T1. Radiologists have modalities.
approved annotations for the lesion area in each case, which is divided
into the peritumoral edema (ED), the necrotic and non-enhancing tumor 5.1. Simulation setup
core (NCR/NET), and the GD-enhancing tumor (ET). These areas can be
clustered into three sub-regions: the enhancing tumor which is repre­ We simulate our proposed method in the PYTHON simulation plat­
sented as ET, the tumor core (TC=ET+NCR), and the whole tumor form. Our model is constructed using Keras and runs on a single Nvidia
(WT=ET+NCR+ED). A series of pre-processing steps were applied to all GPU, specifically the Quadro P5000 (16 G). For model optimization, we

8
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

relied on the Adam optimizer. A learning rate decay scheme was inte­ Table 3
grated, reducing the rate by 0.5 when no improvement was observed Evaluation of our proposed method on BRATS 2019 training set.
within a 10-epoch window. The simulation parameters are provided in Metric Method WT TC ET
Table 2.
DSC Baseline 87.2 77.3 69.6
Additionally, early stopping measures were introduced to prevent Baseline+EFM 88.6 78.0 71.3
over-fitting, terminating training in cases where the validation loss Baseline+DFFM 88.9 77.9 72.4
failed to improve over 50 epochs. A random split is employed to divide Proposed 90.2 81.5 73.7
the dataset into 80% for training and 20% for testing. HD Baseline 8.0 11.5 8.8
Baseline+EFM 6.5 7.7 7.2
Baseline+DFFM 6.4 7.5 6.8
Proposed 4.9 4.6 5.4
5.2. Quantitative analysis Sensitivity Baseline 83.5 70.2 60.2
Baseline+EFM 86.7 79.6 64.8
We first conduct ablation experiments to assess our approach and Baseline+DFFM 90.2 79.9 65.6
show the efficacy of the suggested components on complete modalities, Proposed 95.5 81.4 71.3

and then we compare our model with the state-of-the-art approaches on


full modalities to illustrate the significance of our network design.
Finally, we examine the robustness of our methodology when modalities Table 4
are absent. Evaluation of our proposed method on BRATS 2020 training set.
Metric Method WT TC ET
i. A Full Modalities Evaluation of Our Method: To showcase the per­ DSC Baseline 87.6 78.2 71.5
formance of our approach and underline the significance of the Baseline+EFM 89.0 83.1 72.6
incorporated components within our network, we conducted abla­ Baseline+DFFM 89.3 83.9 72.7
tion studies on both the BraTS 2019 and BraTS 2020 datasets. The Proposed 93.7 85.5 80.6
HD Baseline 7.8 10.2 8.3
results are presented in Table 3 and Table 4, respectively. We
Baseline+EFM 6.0 7.2 6.3
established a baseline for the original network, which excludes the Baseline+DFFM 6.1 6.7 5.9
EFM and DFFM. From the results in Table 3, it is evident that the Proposed 3.7 3.2 2.5
baseline method achieves an average Dice Score of 78.03, an average Sensitivity Baseline 85.7 74.8 64.6
Hausdorff Distance of 9.4, and an average Sensitivity of 71.3. Simi­ Baseline+EFM 87.1 81.1 67.4
Baseline+DFFM 92.2 83.6 70.6
larly, in Table 4, the baseline method achieves an average Dice Score Proposed 94.3 92.8 90.7
of 79.1, an average Hausdorff Distance of 8.8, and an average
Sensitivity of 75.03. Upon applying the EFM to the network, we
observe improvements in the average Dice Score, average Hausdorff
Table 5
Distance, and average Sensitivity. This enhancement is attributed to
Comparison of different techniques with the proposed method on BRATS2019
the EFM’s ability to learn edge representations while preserving validation set in terms of sensitivity, DSC, HD, and specificity.
regional edge properties from various modalities, thereby bolstering
Metric Author/year Methods WT TC ET
segmentation outcomes. Furthermore, integration of the DFFM leads
to further improvements in the average Dice Score and average DSC Zhou et al. [21], 2021 LCR 87.1 71.8 72.7
Yan et al. [24], 2022 SEResU-Net 93.7 91.8 87.5
Sensitivity. Our proposed method with all the improvements signif­
Syazwany et al. [31], 2021 MM-BiFPN 90.7 84.8 78.2
icantly contributes to enhanced segmentation. Table 4 demonstrates Luo et al. [32], 2020 HDC-Net 92.7 95.8 84.2
that the proposed method enhances the baseline by 4.8% in terms of Micallef et al. [33], 2021 U-Net++ 87.0 78.2 69.2
average Dice Score, 47.3% in terms of average Hausdorff Distance, - Proposed 95.6 96.5 92.6
and 16.0% in terms of average Sensitivity on the BraTS 2019 dataset. Sensitivity Zhou et al. [21], 2021 LCR 85.6 67.7 75.2
Yan et al. [24], 2022 SEResU-Net 93.7 94.1 87.3
Table 4 mirrors similar enhancements achieved by the proposed
Syazwany et al. [31], 2021 MM-BiFPN 76.1 80.4 71.7
method on the BraTS 2020 dataset, elevating the baseline by 9.1% in Luo et al. [32], 2020 HDC-Net 81.1 94.1 78.5
terms of average Dice Score, 64.2% in terms of average Hausdorff Micallef et al. [33], 2021 U-Net++ 86.54 76.5 72.0
Distance, and 23.4% in terms of average Sensitivity. - Proposed 94.3 95.8 90.7
ii. Full Modalities Comparison With the State-of-the-Art Methods: We Specificity Zhou et al. [21], 2021 LCR
Yan et al. [24], 2022 SEResU-Net 94.7 93.5 89.6
evaluate our proposed approach by contrasting it with state-of-the- Syazwany et al. [31], 2021 MM-BiFPN 87.6 89.5 85.5
art methods on the BraTS 2019 and BraTS 2020 online validation Luo et al. [32], 2020 HDC-Net 90.5 96.5 92.4
sets. Our process involves initially generating segmentation pre­ Micallef et al. [33], 2021 U-Net++ 99.4 99.6 99.7
dictions on a local machine, followed by their submission to an on­ - Proposed 97 98.7 99.3
HD Zhou et al. [21], 2021 LCR 6.7 9.3 6.3
line evaluation platform to acquire the evaluation outcomes. The
Yan et al. [24], 2022 SEResU-Net 2.1 1.3 2.3
comparative results are displayed in Table 5 and Table 6. Notably, Syazwany et al. [31], 2021 MM-BiFPN 5.6 7.4 8.3
for a fair comparison, we exclude methods employing post- Luo et al. [32], 2020 HDC-Net 1.3 0.7 1.4
processing techniques. Micallef et al. [33], 2021 U-Net++ 8.3 9.4 6.8
- Proposed 3.7 3.2 2.1

In both Tables 5 and 6, illustrating the performance on the


BRATS2019 and BRATS2020 validation sets, respectively, our proposed
Table 2 EADFFTU-Net demonstrates substantial improvements over existing
Simulation parameters. brain tumor segmentation methods, incorporating innovative modules
Parameter Value into its architecture. For Table 5, emphasizing the BRATS2019 valida­
Epoch 100
tion set, the results underscore the efficacy of EADFFTU-Net in accu­
Weight decay 0.0004 rately delineating brain tumors. The EFM integrated into the feature
Optimizer Adam extraction phase significantly contributes to this success, capturing
Learning rate 0.0001 intricate edge characteristics and enhancing sensitivity to fine details.
Batch size 32
The DFFM in the decoder further refines segmentation by selectively

9
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

Table 6 ensures that the model captures global contexts and fuses multiscale
Comparison of different techniques with the proposed method on BRATS2020 features effectively, contributing to improved segmentation accuracy.
validation set in terms of sensitivity, DSC, HD and specificity. According to the specified accuracy intervals, the proposed values for
Metric Author/year Methods WT TC ET EADFFTU-Net (DSC: WT: 93.7, TC: 85.5, ET: 80.6; sensitivity: WT: 94.3,
DSC Liew et al. [28], CASPIANET++ 81.08 83.70 78.50
TC: 92.8, ET: 90.7; specificity: WT: 97, TC: 97.1, ET: 97.9; HD: WT: 3.7,
2021 TC: 3.2, ET: 2.5) categorize it as a very good model, meeting or
Ullah et al. [40], MRA-UNet 87.22 86.74 90.18 exceeding the established thresholds for this classification.
2022
Ma et al. [41], Deep Supervision 78.79 75.93 87.26
i. Analysis of Missing Modalities Compared to Modern Techniques:
2020 CNN
Raza et al. [42], DresUNet 83.57 80.04 86.6 Table 7 and Table 8 provide a robust comparison of different brain
2023 tumor segmentation methods, across various combinations of avail­
- Proposed 93.7 85.5 80.6 able modalities on both the BRATS 2019 and BRATS 2020 datasets.
Sensitivity Liew et al. [28], CASPIANET++ 89.19 79.53 77.11 The objective is to evaluate the robustness of each method in infer­
2021
Ullah et al. [40], MRA-UNet 92.71 91.88 93.54
ring missing modalities, a common scenario in clinical practice.
2022
Ma et al. [41], Deep Supervision - - - The table illustrates the DSC percentages for different combinations
2020 CNN of available modalities (T1, T1c, T2, Flair) and tumor regions (WT, TC,
Raza et al. [42], DresUNet 97.86 96.95 96.48
ET). Our proposed method, denoted as "Ours," is compared against the
2023
Proposed 94.3 92.8 90.7 LCR method [21]. Across all 15 possible combination situations of un­
Specificity Liew et al. [28], CASPIANET++ 99.82 99.97 99.92 available modalities, our method consistently outperforms both [21]
2021 and [43] as indicated by higher mean DSC values. This highlights the
Ullah et al. [40], MRA-UNet 93.19 95.64 92.61 exceptional robustness of our multimodal segmentation method,
2022
showcasing superior performance not only on complete modalities
Ma et al. [41], Deep Supervision - - -
2020 CNN datasets but also when dealing with missing modalities.
Raza et al. [42], DresUNet 99.36 98.62 97.91 Similar to the BRATS 2019 analysis, Table 8 presents the robust
2023 comparison on the BRATS 2020 dataset. The same combinations of
Proposed 97 97.1 97.9
available modalities are considered, and the performance of our pro­
HD Liew et al. [28], CASPIANET++ - - -
2021 posed method is evaluated against Once again, our method consistently
Ullah et al. [40], MRA-UNet - - - outperforms the existing method, demonstrating strong robustness in
2022 inferring missing modalities. The mean DSC values further emphasize
Ma et al. [41], Deep Supervision - - - the reliability of our approach across diverse scenarios.
2020 CNN
Fig. 6 illustrates the performance of the brain tumor segmentation
Raza et al. [42], DresUNet - - -
2023 model in terms of loss over the course of training and validation epochs.
Proposed 3.7 3.2 2.5 Each epoch corresponds to one complete pass through the entire training
dataset. The loss value quantifies the difference between the predicted
segmentation and the ground truth segmentation. Lower loss values
fusing informative features, leading to a nuanced outcome. In accor­
indicate better performance. Decreasing trend in both training and
dance with the suggested accuracy intervals, the proposed values for validation loss indicates that the model is learning and improving its
EADFFTU-Net (DSC: WT: 95.6, TC: 96.5, ET: 92.6; sensitivity: WT: 94.3, performance. Fig. 7 illustrates how well the proposed model classifies
TC: 95.8, ET: 90.7; specificity: WT: 97, TC: 98.7, ET: 99.3; HD: WT: 3.7, the input data in terms of accuracy over the course of training and
TC: 3.2, ET: 2.1) position it as an excellent model, surpassing the validation epochs. Accuracy indicates the proportion of correctly clas­
thresholds indicative of very good/excellent models (>=90%). sified brain tumor segments compared to the total number of segments.
Turning to Table 6, evaluating performance on the BRATS2020 The training accuracy shows how the accuracy improves as the model
validation set, EADFFTU-Net continues to outperform existing methods. learns from the training data over epochs. The validation accuracy
The EFM maintains its effectiveness in learning and preserving edge represents measuring the model’s performance on unseen validation
characteristics, enhancing sensitivity to tumor boundaries. The DFFM [Link] increasing trend in both training and validation accuracy

Table 7
Robust comparison of different methods (DICE %) for different combinations of available modalities on BRATS 2019 dataset.
Modalities WT TC ET

T1 T1c T2 Flair [21] Ours [21] Ours [21] Ours

• ◦ ◦ ◦
10.0 30.2 2.2 20.6 62.7 63.8

• ◦ ◦
20.0 40.4 41.7 80.7 46.8 78.4
◦ ◦
• ◦
70 53.6 32.6 63.1 3.8 41.6
◦ ◦ ◦
• 74.3 82.4 52.0 62.5 17.0 40.8
• ◦
• ◦
70 75.8 42.1 71.6 1.2 49.9
• • ◦ ◦
20.7 80.6 39.3 82.8 42.3 73.2
◦ ◦
• • 83.9 89.7 41.8 67.9 11.3 46.5

• ◦
• 81.7 87.9 74.7 82.7 63.5 77.8
• ◦ ◦
• 81.6 88.4 55.6 67.4 3.0 40.5

• • ◦
73.5 85.9 66.8 82.7 60.4 75.7
• • • ◦
74.2 85.3 66.6 83.8 64.0 74.7

• • • 88.9 93.2 76.6 80.7 64.4 75.6
• ◦
• • 88.2 91.1 52.1 63.4 2.9 50.5
• • ◦
• 83.9 89.8 75.7 83.2 67.1 79.6
• • • • 89.7 91.7 77.5 83.9 70.6 83.6
mean 67.37 77.73 53.15 71.8 38.7 63.48

10
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

Table 8
Robust comparison of different methods (DICE %) for different combinations of available modalities on BRATS 2020 dataset.
Modalities WT TC ET

T1 T1c T2 Flair [41] Ours [41] Ours [41] Ours

• ◦ ◦ ◦
77.16 90.6 66.02 75.3 37.30 67.4

• ◦ ◦
76.77 79.9 81.51 84.3 74.85 81.9
◦ ◦
• ◦
86.05 85.4 71.02 69.9 46.29 55.8
◦ ◦ ◦
• 87.32 92.5 69.19 80.8 38.15 43.1
• ◦
• ◦
87.73 90.7 73.13 89.7 45.65 52.6
• • ◦ ◦
81.12 87.4 83.40 91.1 78.01 84.5
◦ ◦
• • 89.87 90.2 74.14 87.6 49.32 74.2

• ◦
• 89.89 93.2 84.65 84.5 76.67 83.1
• ◦ ◦
• 89.73 91.5 73.07 81.5 40.98 52.7

• • ◦
87.74 92.6 83.45 92.8 75.93 81.3
• • • ◦
88.25 94.3 83.47 90.0 76.99 82.6

• • • 90.68 92.4 84.97 82.4 77.12 85.2
• ◦
• • 90.60 92.7 75.19 91.1 49.92 57.5
• • ◦
• 90.69 91.8 85.07 88.5 76.81 82.5
• • • • 91.11 93.7 85.21 92.7 78.00 85.6
mean 86.98 89.98 78.23 83.46 61.47 72.15

indicates that the model is learning to classify brain tumor segments


more accurately.

5.3. Qualitative analysis

The segmentation result on full modalities is provided in Fig. 8. The


results of our analysis reveal that, despite an increasing number of
missing modalities, our robust model’s segmentation results experience
only a slight degradation rather than a sudden, sharp decline. These
findings demonstrate the remarkable robustness of our multimodal
segmentation approach, which can achieve accurate segmentation of the
entire tumor using just the FLAIR modality. Furthermore, when both the
FLAIR and T1c modalities are utilized, the method yields a segmentation
result that is competitive with the ground truth. the segmentation results
are gradually improved when the proposed strategies are integrated,
these comparisons indicate the effectiveness of the proposed strategies.
In addition, with all the proposed strategies, our proposed method can
achieve almost the same results with ground truth.
Upon examination of Fig. 9 it becomes evident that the integration of
our proposed strategies leads to gradual improvements in segmentation
Fig. 6. Training and validation loss.
results. This comparison underscores the effectiveness of our strategies.
Furthermore, when all proposed strategies are employed, our method
achieves segmentation results that closely resemble the ground truth. It
also reveals that as the number of missing modalities increases, the
segmentation results produced by our robust model exhibit only a minor
degradation, rather than a sudden, sharp decline. Notably, using only
the FLAIR modality, our method performs well in segmenting the entire
tumor. When both FLAIR and T1c modalities are utilized, the segmen­
tation results remain competitive with the ground truth. Additionally,
the T1 and T2 modalities contribute to refining the boundary areas of the
tumor regions, leading to the best segmentation outcomes.

5.4. Limitations

The EA-DFFTU-Net represents a significant step forward in brain


tumor segmentation; however, its primary limitation arises from its
reliance on a two-dimensional (2D) processing paradigm. By analyzing
MRI data slice by slice, the model may overlook crucial three-
dimensional (3D) spatial relationships and contextual information pre­
sent in volumetric scans. This limitation can lead to a loss of valuable
details and hinder the model’s ability to accurately delineate tumor
Fig. 7. Training and validation accuracy. boundaries, especially in cases where tumors span multiple slices or
exhibit irregular shapes. Consequently, while the EA-DFFTU-Net ach­
ieves notable segmentation results in certain scenarios, its efficacy may
be compromised when faced with complex tumor geometries or subtle
variations across the 3D volume. Furthermore, the inherent nature of 2D

11
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

Fig. 8. Examples of the segmentation results on full modalities on BraTS 2020 dataset.

Fig. 9. The results of the segmentation process on modalities that are not available are displayed as follows: The necrotic and non-enhancing tumor core are
indicated by Red; the enhancing tumor is indicated by Blue; and edema is indicated by Green.

processing poses challenges in capturing fine-grained local details efficiency of the segmentation model. Moreover, the reliance on slice-
within tumors. While individual slices may contain informative features, based features may hinder the model’s ability to capture global char­
they lack the depth and continuity present in volumetric data. Conse­ acteristics and inter-slice correlations, further impacting its segmenta­
quently, the model may struggle to discern subtle variations in tumor tion accuracy. Overcoming these limitations may require exploring
characteristics or accurately distinguish between tumor and healthy alternative strategies, such as incorporating 3D convolutional neural
tissue, particularly in regions where the distinction is less clear. This networks or developing techniques to better integrate inter-slice
limitation becomes particularly pronounced in cases where tumors context, to enhance the model’s performance in analyzing volumetric
exhibit heterogeneous textures or irregular boundaries, as the model MRI data for brain tumor segmentation.
may fail to integrate information across slices effectively, resulting in
segmentation inaccuracies and potential misclassifications. 6. Conclusion
Additionally, the slice-based approach introduces computational
complexities, especially when dealing with large tumors or high- In this study, we introduced a novel approach named EA-DFFTU-Net,
resolution MRI scans. Slicing volumetric data into individual 2D im­ specifically designed for the precise segmentation of brain tumors, even
ages increases the number of input slices, requiring additional compu­ in scenarios involving missing modalities within multi-modal MRI data.
tational resources and memory. Consequently, processing large datasets Our approach seamlessly integrates various components to enhance
becomes more challenging, potentially limiting the scalability and segmentation accuracy and reliability. The incorporation of a ResNet-50

12
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

encoder enables robust multi-level feature extraction, capturing both and/or animals performed by any of the authors.
local spatial features and global contextual information. The innovative
EFM adds another dimension by preserving crucial edge properties Informed consent
within regions. The use of different encoders for different modalities and
the inclusion of the EFM after each encoder significantly enhance the There is no informed consent for this study.
robustness of our model in cases with missing modalities. The
Discriminative Feature Fusion-based Transformer U-Net (DFFTU-Net) Funding
further refines the segmentation by capturing global contexts and mul­
tiscale features, while the DFFM effectively combines these features for There is no funding for this study.
optimal segmentation outcomes. Our comprehensive evaluation, con­
ducted on benchmark datasets, including BraTS2019 and BraTS2020, CRediT authorship contribution statement
demonstrated that the proposed EA-DFFTU-Net consistently out­
performed state-of-the-art models. This achievement underscores the B. Jagadeesh: Conceptualization, Formal analysis, Writing – orig­
model’s potential to significantly improve the accuracy of brain tumor inal draft, Writing – review & editing. G. Anand Kumar: Conceptuali­
segmentation, which holds promise for enhancing medical diagnosis and zation, Formal analysis, Methodology, Writing – original draft.
treatment planning. Nonetheless, it is essential to acknowledge the
study’s limitations. One notable constraint is that our model operates Declaration of Competing Interest
primarily in a 2D context, and when handling 3D MRI data, slicing
processes may inadvertently result in the loss of some context infor­ Authors declares that they have no conflict of interest.
mation and local details. To overcome this limitation, our future work
will explore the extension of EA-DFFTU-Net into a 3D network archi­ Data availability
tecture, thus allowing us to harness the full potential of 3D information
in MRI data for improved segmentation accuracy. Data will be made available on request.

Ethical approval

This article does not contain any studies with human participants

Appendix

A. Overview of the reviewed works

Fig. A1 Internal connections of Swin Block.

The Table 1 offers a comprehensive overview of various deep learning methods employed for brain tumor segmentation, highlighting key details
such as the year of development, pre-processing techniques, the segmentation model used, the dataset involved, and the metrics utilized for evalu­
ation. Each row corresponds to a distinct method, facilitating a quick and systematic comparison of their characteristics. Metrics range from traditional
ones like Dice score and Sensitivity to more specialized measures such as Hausdorff Distance and Specificity. Figure A1 likely visualizes the internal
connections or architecture of a Swin Block, an essential component in the reviewed deep learning models. This visual aid can provide additional
insights into the intricate details of the Swin Block structure, aiding readers in understanding the internal mechanisms of the associated deep learning
models.

13
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709

References [22] B. Yu, L. Zhou, L. Wang, W. Yang, M. Yang, P. Bourgeat, J. Fripp, SA-LuT-Nets:
learning sample-adaptive intensity lookup tables for brain tumor segmentation,
IEEE Trans. Med. Imaging 40 (5) (2021) 1417–1427.
[1] T. Saba, A.S. Mohamed, M. El-Affendi, J. Amin, M. Sharif, Brain tumor detection
[23] X. Wu, L. Bi, M. Fulham, D.D. Feng, L. Zhou, J. Kim, Unsupervised brain tumor
using fusion of hand crafted and deep learning features, Cogn. Syst. Res. 59 (2020)
segmentation using a symmetric-driven adversarial network, Neurocomputing 455
221–230.
(2021) 242–254.
[2] M.K. Abd-Ellah, A.I. Awad, A.A. Khalaf, H.F. Hamed, Two-phase multi-model
[24] C. Yan, J. Ding, H. Zhang, K. Tong, B. Hua, S. Shi, SEResU-Net for Multimodal
automatic brain tumour diagnosis system from magnetic resonance images using
Brain Tumor Segmentation, IEEE Access 10 (2022) 117033–117044.
convolutional neural networks, EURASIP J. Image Video Process. 2018 (1) (2018),
[25] Z. Huang, Y. Zhao, Y. Liu, G. Song, GCAUNet: A group cross-channel attention
1-0.
residual UNet for slice based brain tumor segmentation, Biomed. Signal Process.
[3] P. Wang, Y.H. Shi, J.Y. Li, C.Z. Zhang, Differentiating Glioblastoma from Primary
Control 70 (2021) 102958.
Central Nervous System Lymphoma: The Value of Shaping and Non enhancing
[26] Z. Xiao, K. He, J. Liu, W. Zhang, Multi-view hierarchical split network for brain
Peritumoral Hyper intense Gyral Lesion on FLAIR Imaging, World Neurosurg. 149
tumor segmentation, Biomed. Signal Process. Control 69 (2021) 102897.
(2021) e696–e704.
[27] Y. Zhang, Y. Lu, W. Chen, Y. Chang, H. Gu, B. Yu, MSMANet: A multi-scale mesh
[4] G. Anand Kumar, P.V. Sridevi, 3D Deep Learning for Automatic Brain MR Tumor
aggregation network for brain tumor segmentation, Appl. Soft Comput. 110 (2021)
Segmentation with T-Spline Intensity Inhomogeneity Correction, Autom. Control
107733.
Comput. Sci. (ACCS) 52 (5) (2018) 439–450.
[28] A. Liew, C.C. Lee, B.L. Lan, M. Tan, CASPIANET++: a multidimensional channel-
[5] D. Zhang, G. Huang, Q. Zhang, J. Han, J. Han, Y. Wang, Y. Yu, Exploring task
spatial asymmetric attention network with noisy student curriculum learning
structure for brain tumor segmentation from multi-modality MR images, IEEE
paradigm for brain tumor segmentation, Comput. Biol. Med. 136 (2021) 104690.
Trans. Image Process. 29 (2020) 9032–9043.
[29] H.M. Rai, K. Chatterjee, S. Dashkevich, Automatic and accurate abnormality
[6] H. Zhou, K. Chang, H.X. Bai, B. Xiao, C. Su, W.L. Bi, P.J. Zhang, J.T. Senders,
detection from brain MR images using a novel hybrid UnetResNext-50 deep CNN
M. Vallières, V.K. Kavouridis, A. Boaro, Machine learning reveals multimodal MRI
model, Biomed. Signal Process. Control 66 (2021) 102477.
patterns predictive of isocitrate dehydrogenase and 1p/19q status in diffuse low-
[30] L. Tan, W. Ma, J. Xia, S. Sarker, Multimodal magnetic resonance image brain tumor
and high-grade gliomas, J. neuro-Oncol. 142 (2) (2019) 299–307.
segmentation based on ACU-net network, IEEE Access 9 (2021) 14608–14618.
[7] G. Anand Kumar, P.V. Sridevi, Deep learning network with Euclidean factor for
[31] N.S. Syazwany, J.H. Nam, S.C. Lee, MM-BiFPN: Multi-Modality Fusion Network
Brain MR Tumor segmentation and volume estimation, Int. J. Model., Simul., Sci.
With Bi-FPN for MRI Brain Tumor Segmentation, IEEE Access 9 (2021)
Comput. (IJMSSC) 10 (17) (2019) 1950039.
160708–160720.
[8] R. Wang, T. Lei, R. Cui, B. Zhang, H. Meng, A.K. Nandi, Medical image
[32] Z. Luo, Z. Jia, Z. Yuan, J. Peng, HDC-Net: Hierarchical decoupled convolution
segmentation using deep learning: A survey, IET Image Process. 2016 (5) (2022)
network for brain tumor segmentation, IEEE J. Biomed. Health Inform. 25 (3)
1243–1267.
(2020) 737–745.
[9] Z. Stosic, P. Rutesic, An improved canny edge detection algorithm for detecting
[33] N. Micallef, D. Seychell, C.J. Bajada, Exploring the U-Net++ Model for Automatic
brain tumors in MRI images, Int. J. Signal Process. 3 (2018).
Brain Tumor Segmentation. IEEE Access 9 (2021) 125523–125539.
[10] M.H. Hesamian, W. Jia, X. He, P. Kennedy, Deep learning techniques for medical
[34] M. Aghalari, A. Aghagolzadeh, M. Ezoji, Brain tumor image segmentation via
image segmentation: achievements and challenges, J. Digit. Imaging 32 (4) (2019)
asymmetric/symmetric UNet based on two-pathway-residual blocks, Biomed.
582–596.
Signal Process. Control 69 (2021) 102841.
[11] N. Yamanakkanavar, J.Y. Choi, B. Lee, MRI segmentation and classification of
[35] K. Hao, S. Lin, J. Qiao, Y. Tu, A generalized pooling for brain tumor segmentation,
human brain using deep learning for diagnosis of Alzheimer’s disease: a survey,
IEEE Access 9 (2021) 159283–159290.
Sensors 20 (11) (2020) 3243.
[36] X. Li, Y. Xu, N. Li, B. Yang, Y. Lei, Remaining useful life prediction with partial
[12] G. Du, X. Cao, J. Liang, X. Chen, Y. Zhan, Medical image segmentation based on u-
sensor malfunctions using deep adversarial networks, IEEE/CAA J. Autom. Sin. 10
net: A review, J. Imaging Sci. Technol. 64 (2020) 1–2.
(1) (2022) 121–134.
[13] O. Çiçek, A. Abdulkadir, S.S. Lienkamp, T. Brox, O. Ronneberger, 3D U-Net:
[37] X. Li, S. Yu, Y. Lei, N. Li, B. Yang, Intelligent machinery fault diagnosis with event-
learning dense volumetric segmentation from sparse annotation. In International
based camera, IEEE Trans. Ind. Inform. (2023).
conference on medical image computing and computer-assisted intervention.
[38] N.J. Tustison, B.B. Avants, P.A. Cook, Y. Zheng, A. Egan, P.A. Yushkevich, J.C. Gee,
Springer Cham, (2016) pp. 424-432.
N4ITK: improved N3 bias correction, IEEE Trans. Med. Imaging 29 (6) (2010)
[14] X. Xiao, S. Lian, Z. Luo, S. Li, Weighted res-unet for high-quality retina vessel
1310–1320.
segmentation. In2018 9th international conference on information technology in
[39] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, (2021) Swin
medicine and education (ITME) IEEE, (2018) pp. 327-331.
transformer: Hierarchical vision transformer using shifted windows. In Proceedings
[15] Z. Zhou, M.M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, Unet++: A nested u-net
of the IEEE/CVF international conference on computer vision (2021) pp. 10012-
architecture for medical image segmentation. Deep learning in medical image
10022.
analysis and multimodal learning for clinical decision support, Springer, Cham,
[40] Z. Ullah, M. Usman, M. Jeon, J. Gwak, Cascade multiscale residual attention cnns
2018, pp. 3–11.
with adaptive roi for automatic brain tumor segmentation, Inf. Sci. 608 (2022)
[16] X. Liu, S. Hou, S. Liu, W. Ding, Y. Zhang, Attention-based multimodal glioma
1541–1556.
segmentation with multi-attention layers for small-intensity dissimilarity, J. King
[41] S. Ma, Z. Zhang, J. Ding, X. Li, J. Tang, F. Guo, A deep supervision CNN network for
Saud. Univ. -Comput. Inf. Sci. 35 (4) (2023) 183–195.
brain tumor segmentation. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and
[17] X. Liu, Y. Liu, W. Fu, S. Liu, SCTV-UNet: a COVID-19 CT segmentation network
Traumatic Brain Injuries: 6th International Workshop, BrainLes 2020, Held in
based on attention mechanism, Soft Comput. 21 (2023), 1-1.
Conjunction with MICCAI 2020, Lima, Peru, Revised Selected Papers, Part II 6
[18] S. Asgari Taghanaki, K. Abhishek, J.P. Cohen, J. Cohen-Adad, G. Hamarneh, Deep
2021 Springer International Publishing. (2020) pp. 158-167.
semantic segmentation of natural and medical images: a review, Artif. Intell. Rev.
[42] R. Raza, U.I. Bajwa, Y. Mehmood, M.W. Anwar, M.H. Jamal, dResU-Net: 3D deep
54 (2021) 137–178.
residual U-Net based brain tumor segmentation from multimodal MRI, Biomed.
[19] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser,
Signal Process. Control 79 (2023) 103861.
I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. (2017) 30.
[43] Y. Ding, X. Yu, Y. Yang, RFNet: Region-aware fusion network for incomplete multi-
[20] P. Shaw, J. Uszkoreit, A. Vaswani, Self-attention with relative position
modal brain tumor segmentation. In Proceedings of the IEEE/CVF international
representations. arXiv preprint arXiv (2018) 1803.02155.
conference on computer vision (2021) pp. 3975-3984.
[21] T. Zhou, S. Canu, P. Vera, S. Ruan, Latent correlation representation learning for
brain tumor segmentation with missing MRI modalities, IEEE Trans. Image Process.
30 (2021) 4263–4274.

14

You might also like