Missing Modality Paper
Missing Modality Paper
Brain tumor segmentation with missing MRI modalities using edge aware
discriminative feature fusion based transformer U-net
B. Jagadeesh a, G. Anand Kumar b, *
a
Professor, Department of Electronics and Communication Engineering, Anil Neerukonda Institute of Technology and Sciences (Autonomous), Visakhapatnam, India
b
Associate Professor, Department of Cyber Security & IoT, School of Engineering, Malla Reddy University, Hyderabad, India
H I G H L I G H T S
A R T I C L E I N F O A B S T R A C T
Keywords: Brain tumor segmentation is an essential task for medical diagnosis and treatment planning. Multi-modal MRI
Brain tumor provides complementary information that is essential for accurate segmentation of brain tumors, but missing
Transformer network modality images are a common problem in clinical practice. Existing segmentation methods often fail to generate
Missing modalities
accurate object boundaries and selectively fuse the tumor region, resulting in unreliable segmentation masks. In
Feature extraction
this work, we propose an Edge-aware Discriminative Feature Fusion Based Transformer U-Net (EA-DFFTU-Net)
Segmentation
to segment the brain tumor effectively even in the absence of modalities. First, the MRI input data is pre-
processed, and features are then extracted using a ResNet-50 encoder, which learns local spatial. We employ
an Edge Feature Module (EFM) to acquire edge attention representations. These features are then transferred to
the Discriminative Feature Fusion based Transformer U-Net (DFFTU-Net) that learns global contexts and captures
multiscale features. We use a Discriminative Feature Fusion Module (DFFM) in the decoder of the DFFTU-Net to
effectively fuse the multiscale features in order to obtain accurate segmentation. The performance of the pro
posed EA-DFFTU-Net segmentation method was determined by evaluating it and comparing the obtained results
with those of existing brain tumor segmentation techniques.
1. Introduction various glioma subregions [3,4]. The tumour core, also known as the
tumour without peritumoral edoema, is highlighted on T1 and T1c. T2
The brain tumor is recognized as one of the most dangerous and and FLAIR emphasise the complete tumour, which is the tumour with
deadly diseases globally, and it is crucial to detect brain tumor earlier for peritumoral edoema. T1c also shows an enhancing tumour core, which
clinical evaluation and treatment planning [1]. A common imaging is an area of the tumour core with hyperintensity. The time, availability
method for evaluating brain tumors is magnetic resonance imaging of the scanners, cost, and patient comfort prevent most patients from
(MRI) because without using radiation, it provides a strong soft tissue having access to the entire variety of multi-modality MRI scans [5].
contrast [2]. There are multiple MRI modalities, namely T2-weighted, Despite numerous techniques being developed for brain tumor seg
T1-weighted, contrast-enhanced T1-weighted (T1c), and FLAIR (Flu mentation, the process is still difficult and challenging especially when
id-Attenuation Inversion Recovery), based on various contrasts and certain modalities are absent. This is because gliomas vary for patients in
functional aspects. Different imaging modalities reveal various tissue size, shape, and texture, variable intensity range, and poor contrast in
architectures and provide additional information for the analysis of MRI [6]. The process of manually segmenting brain tumors by experts is
* Corresponding author.
E-mail addresses: bjagadeesh76@[Link] (B. Jagadeesh), anandlife@[Link] (G. Anand Kumar).
[Link]
Received 27 April 2023; Received in revised form 13 March 2024; Accepted 26 April 2024
Available online 10 May 2024
1568-4946/© 2024 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
both expensive and time-consuming [7]. One of the most essential • Discriminative feature fusing Transformer U-Net (DFFTU-Net) is
components for attaining accurate segmentation is edge information. introduced to capture the global contexts and segment the tumor
Existing medical image segmentation approaches use local gradient region in which we employ a DFFM to find the most informative
representations to identify the edge regions and the obtained closed loop features and effectively fuse them.
area is detected as objects [8,9]. These methods are accurate only for • The EADFFTU-Net approach is shown to have better performance
simple structural images. The lack of object-level information and reli than existing approaches.
ance on local edge representations in various edge detection algorithms
results in trivial segmentation areas and discontinuous borders [10]. The paper’s remaining sections are outlined as follows:
In medical image, segmentation CNNs have administered a variety of Section 2 offers a literature review, Section 3 describes the proposed
applications since the emergence of deep learning [11]. The conven methodology, Section 4 provides the simulation setup and the compar
tional encoder-decoder-based network U-Net has shown excellent seg ative analysis and Section 5 presents the conclusion.
mentation potential among the different CNN variants [12]. Improved
performance in medical image segmentation has been observed with the 2. Related work
introduction of models such as 3D UNet [13], Res-UNet [14], and
UNet++ [15]. CNNs typically increase the receptive field size as the Below are a few recent publications that focus on brain tumor
network depth increases. Attention mechanisms were also introduced segmentation.
[16,17]. However, this strategy has limited efficacy in capturing global A strategy for segmenting brain tumours with missing modalities was
relationships. This causes problems such as loss of global context and not put forth by Zhou et al. [21]. To overcome the problem due to missing
developing long-range dependencies. It also makes the model more modalities, they used a multi-modal segmentation network on the basis
complicated and causes overfitting [18]. of U-Net architecture which integrates fusion block and correlation
Transformer, [19] which successfully enhances network perfor model. The latent correlation representations (LCR) obtained from the
mance in the Natural Language Processing (NLP) space, has recently correlation model are fused by the fusion block that the correlation
proved valuable in computer vision. The Vision Transformer (ViT) al model discovers between modalities through the attention mechanism
gorithm segments an image into multiple patches, making them com which is then decoded to get the segmentation result. However, this
parable. and then the Transformer’s self-attention mechanism is used on approach did not take into account the edge information that is crucial
each patch to capture the long-range dependencies. Limitations that are for precise segmentation.
caused by the networks using convolutional layers are overcome by ViT Yu et al. [22] proposed SA-LuT-Net which is built around backbone
[20]. There are few studies using Transformers for medical image seg networks such as the 3D Unet model and DMFNet. A LuT module and a
mentation, but they neglect the local representations which are essential segmentation module make up the network. Propagation of the LuT
for refining the tumor details as they use a transformer encoder to module with the segmentation module learns an intensity mapping
extract features. Also, the existing methods are not robust to missing function. This method gives better segmentation results but it is
modalities and they mainly focus on the primary region and don’t time-consuming.
consider the edge information, which leads to inaccurate segmentation Wu et al. [23] introduced a generative adversarial network that is
of tumor regions. Additionally, the most essential features of several symmetric driven (SD-GAN). To represent the variation (symmetry) of
modalities that are very pertinent to brain tumors are not fused. So, the normal brains the network was trained to learn a non-linear mapping
features of the tumor region may not get fused properly and produce of brain images. Then the trained model is used in normal brain
inaccurate results. Moreover, it cannot capture long-range correlations. reconstruction and brain tumor segmentation based on increased
Some drawbacks of the existing segmentation approaches include in reconstruction errors caused by their asymmetry. This method is hard to
accuracy, overfitting, and low significant performance. To overcome train.
this issue Edge Aware Discriminative Feature Fusion-based Transformer Yan et al. [24] proposed SEResU-Net, an improved U-Net model that
U-Net (EADFFTU-Net) is proposed. links deep ResNet and the SENet. The network extracts additional
Three phases, namely pre-processing, feature extraction, and seg feature information and overcomes network degradation. It also pre
mentation, make up the proposed model. In the phase of pre-processing, vents information loss. It also resolves the challenge of tiny brain tu
the input data is pre-processed using intensity normalization and N4ITK. mour’s low segmentation accuracy but failed to consider edge
The intensity variations in images are normalized using intensity information and not robust to missing modalities.
normalization and the distortion in MR data is corrected using N4ITK. Huang et al. [25] put forward GCAUNet, which may effectively
After pre-processing, the multi-level features are extracted from the pre- utilise the low-level fine information of tumour areas to enhance the
processed data using the ResNet50 encoder. Here, each input modality is segmentation of brain tumors. They introduced a detail recovering path
processed by an individual encoder to extract features. We develop an (DR) and multiscale input path (MI) to retrieve the precise features of
EFM, which is placed after each encoder block to learn and preserve the brain tumours and to obtain multiscale context data. Moreover, GCA is
edge characteristics of each modality. The obtained multi-level features used to draw attention to the important feature groups and channels.
are fed into the dual path encoder of the proposed DFFTU-Net to capture This method solves the problem of feature loss but is not robust to
global contexts and multiscale features. Effective fusion of multi-scale missing modalities.
features from different modalities is achieved and unwanted features The network’s performance is constrained by employing just 3D
are discarded using the proposed DFFM. Finally, decoding the fused convolution as a module of feature extraction. By considering this issue
features yields the final segmentation result. Xiao et al. [26] proposed a MVHS-Net. Modified 3D UNet introduces the
The proposed EADFFTU-Net framework is characterized by the MVHS block and multi-view fusion convolution block as its structural
following key contributions: components. To gather multi-scale as well as multi-view information,
these modules are introduced. This method decreases the amount of
• To achieve the best segmentation result even in the case of missing irrelevant character information but is not robust to missing modalities.
modalities EADFFTU-Net model is proposed that learns both local Zhang et al. [27] proposed an MSMANet that developed an enhanced
features and global contexts. reception module to gather and extract useful information from multiple
• A new module called the Edge Feature Module (EFM) is proposed to receptive fields. MA technique is utilized in place of the U-Net’s skip
learn edge representations while preserving regional edge connection. This gradually reduces the semantic gap and enhances the
properties. shallow features. Various scale multi-level feature aggregation is
increased in this approach. In order to enhance network identification
2
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
and convergence capabilities, MSMANet additionally uses an attention useful life (RUL) prediction in machinery health assessment. The
technique. This method aims only at the primary areas and discards the method extracts global features from multiple sensors and utilizes
edge information. adversarial learning for generalized feature extraction. Li et al. [37]
Liew et al. [28] introduced the CASPIAN network that uses a noisy introduced the first attempt to use event-based vision data for machine
student curriculum learning paradigm. A multidimensional approach is fault diagnosis. They proposed a vibration event representation and
adapted that integrates multiplanar and multiscale information to employed a deep convolutional neural network model for processing.
improve spatial information inside the network. The proposed Additionally, they introduced an event data augmentation method and a
CASPIANET++ combines channel attention with multiplanar and mul deep representation clustering method. Table 1 presents an overview of
tiscale spatial attention for accurate segmentation but it is the discussed works.
time-consuming and not robust to missing modalities. The majority of currently used techniques rely on a single modality
Rai et al. [29] proposed UnetResNext-50 (hybrid deep CNN model) and are therefore not effective when a modality is absent. Also, it mainly
with a lot of parameters and layers for accurate brain tumor segmen focuses on the primary region and did not consider the edge informa
tation. Unet and ResNext are combined to form UnetResNext-50 in tion, which leads to inaccurate segmentation of tumor regions. It also
which U-net serves as the backbone and it uses the ResNext-50 encoder. fails to fuse the most essential features of several modalities that are
This technique solves the gradient degradation problem by employing a highly pertinent to brain tumors. So, features other than tumor region
cardinality skip connection but it did not consider edge information. may get fused and produce inaccurate results. Moreover, the use of CNN
Tan et al. [30] proposed ACU-Net, a technique for multimodal brain for segmentation cannot capture long-range correlations. Some draw
tumor segmentation. Here, the deep separable convolutional layers are backs of the existing segmentation approaches include inaccuracy,
used instead of U-Net in which the mapped convolutional channel’s overfitting, and low significant performance. To overcome these prob
appearance and spatial correlation are differentiated. To increase the lems EADFFTU-Net is proposed.
capacity of feature propagation and convergence speed residual skip
connections have been included. To improve the segmentation accuracy, 3. Proposed EADFFTU-Net methodology
they insert the active contour model, that ensures proper fitting between
boundary’s inside and outside. This method is computationally efficient, This paper proposes Edge Aware Discriminative Feature Fusion
but it doesn’t give accurate segmentation results. based Transformer U-Net (EADFFTU-Net) for brain tumor segmentation
Syazwany et al. [31] proposed MM-BiFPN to overcome the problems in the absence of MRI modalities. Fig. 1 shows the architectural layout of
in existing methods that ignore the interdependence of the various the proposed method. The framework of the proposed model carries
modalities. The features of each modality are extracted by specific en three phases namely pre-processing, feature extraction, and segmenta
coders to learn the complex relationship among them. The segmentation tion. At the phase of pre-processing, intensity normalization and N4ITK
accuracy is enhanced by the use of the BiFPN layer. This is achieved by are used to pre-process the input data. The intensity variations in images
learning the multiscale features and cross-modality relationships. This are normalized using intensity normalization and the distortion in MR
method works better with missing modalities but discards the edge data is corrected using N4ITK. After pre-processing, multi-level features
features and has high computational overhead. that includes local features are extracted from the pre-processed data
Luo et al. [32] proposed a light-weight HDC-Net, a pseudo-3D using the ResNet50 encoder. Each input modality undergoes feature
network with the architecture of 3D U-Net in which HDC modules are extraction using the individual encoder here. We develop an EFM, which
used in place of 3D convolution modules to extract multiscale and is placed after each encoder block to learn and preserve the edge char
Multiview spatial features more efficiently. In the first stage, periodic acteristics of each modality. The obtained multi-level features are fed
down-shuffling is used as downsampling because it maintains the entire into the dual path encoder of the proposed DFFTU-Net to capture global
information of the image and is parameter-free. View decoupled contexts and multiscale features With the proposed DFFM, multi-scale
convolution in the spatial domain reduce computational complexity but features from various modalities are fused effectively while discarding
it discards the edge information. unwanted features. Finally, decoding the fused features yields the final
Micallef et al. [33] proposed a modified U-Net++ model. In terms of segmentation result.
loss function, deep supervision employing method and number of con
volutional blocks it varies from the U-Net++ model. This network model 3.1. Pre-processing
has five depth levels the last of which connects the network’s decoder.
Instance normalization and a filter map resolution of 16 are also Pre-processing entails adding relevant and helpful information to the
included. This method is not robust to missing modalities. dataset and removing irrelevant data. In the realm of computer vision
Aghalari et al. [34] proposed an improved U-Net based architecture. pre-processing is essential, particularly when analysing medical images,
Three TPR-based UNet models such as TPR-Encoder-Unet, TPR-Deco since erroneous data or undesired information might cause a highly
der-UNet and TPR-Encoder-Decoder-UNet were introduced to enhance effective algorithm to function poorly. Initially, we used N4ITK, and
the effectiveness of segmentation. Initially, TPR blocks are added to the then intensity normalization was performed to pre-process the input
pipeline for down sampling, which is called TPRE-UNet. Then, the initial data. Correcting for bias-field inhomogeneities is essential for improving
method is repeated for up-sampling to create the TPRD-UNet model. the quality of MRI images, as the smooth, low-frequency signal can
Finally, PR blocks are added on both UNet paths, and a model called degrade them. To achieve this, the N4ITK [38] algorithm is used. The
TPRED-UNet (TPR-Encoder-Decoder-UNet) is created. This model process is detailed as follows:
exploit both local and global features simultaneously but it is not robust The initial step involves utilizing the image formation model, which
to missing modalities and not considered edge information. is provided in Eq. (1)
Hao et al. [35] introduced a generalized pooling method (GP) by
B(x) = m(x)e(x) + n(x) (1)
fusing average pooling and maximum pooling with adaptive weights for
the segmentation of brain tumors. This study aims at making the pooling where B corresponds to the observed image, m represents the uncor
operations better. This eliminates the requirement of choosing average rupted image, e stands for the bias field, and n signifies noise, assumed to
or maximum pooling for CNN model down-sampling. GP creates pooling conform to a Gaussian distribution and exhibit independence. Upon
using a variety of expressions on the basis of input images or feature
introducing the notation m ̂ = log m and working within a noise-free
maps. This method solves the problem in conventional pooling methods
setting, the model adopts the subsequent expression as in Eq. (2):
but it does not preserve all the spatial information.
Li et al. [36] proposed a deep learning-based method for remaining
3
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
Table 1
Overview of the reviewed works.
Method Year Pre-processing Deep learning model Dataset Metrics
for segmentation
Modelling the latent multi source 2021 N41TK and intensity U-Net [19] BraTS 2018 and 2019 Dice score, Sensitivity, Hausdroff
correlation and fusion [21] normalization dataset Distance (HD),
SA-LuT-Net [22] 2021 Zero-mean, Unit Standard SA-LuT-Net BraTS 2018 and 2019 Dice score, Hausdorff95
Deviation (STD)normalization dataset
SD-GAN [23] 2021 - SD-GAN BRATS 2012 and 2018 precision, sensitivity, Dice score
SEResU-Net [24] 2022 Data standardization SEResU-Net BraTS 2018 and 2019 Dice similarity coefficient (DSC),
dataset specificity, sensitivity, HD
GCAUNet [25] 2021 Iterative thresholding, zero- GCAUNet BraTS 2017 and BraTS DSC, sensitivity, HD
score normalization 2018 dataset
MVHS-Net [26] 2021 - MVHS-Net BraTS 2018 dataset DSC, Hausdorff95
MSMANET [27] 2021 standardization MSMANET BraTS 2018 dataset DSC, sensitivity, specificity,
Hausdorff95
CASPIANET++ [28] 2021 - CASPIANET++ BraTS 2019 and BraTS Dice score, HD
2020 dataset
UnetResext-50 [29] 2021 Resizing, global pixel UnetResext-50 The Cancer Genome Dice score, Jaccard index, F1-score,
normalization, augmentation Atlas (TCGA) dataset precision, recall and accuracy
ACU-Net [30] 2021 - ACU-Net BraTS 2015, BraTS Dice coefficient, Intersection of Union
2018, BraTS 2019 (IoU), precision, recall, Hausdorff95
dataset
MM-BiFPN [31] 2021 - MM-BiFPN MICCAI BraTS2018 and Dice coefficient, sensitivity, positive
MICCAI BraTS2020. predicted value (PPV), Jaccard
similarity index
HDC-Net [32] 2020 - HDC-Net BraTS 2018 and BraTS Dice sore, Hausdorff95
2017 dataset
U-Net++ [33] 2021 Intensity normalization U-Net++ BraTS 2019 dataset Dice coefficient, HD
Modified UNet [34] 2021 Image normalization TPRED-UNet BraTS 2018 dataset Accuracy, dice similarity coefficient,
sensitivity, specificity, PPV
Generalized Pooling [35] 2021 - Unet, FCN8, UNet++ GP- BraTS 2018 and BraTS Dice similarity coefficient, PPV,
CNN 2019 dataset sensitivity
4
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
b(x) = m(x)
̂ ̂ + e(x) (2)
In accordance with Tustison, the iterative methodology of N4ITK is
expressed in Eq. (3):
n− 1 n− 1
H ∗ {m
̂ − E[ m|
̂ m̂ ]}
̂n = m
m ̂ n− 1 − (3)
̂e 2r
Large intensity variations that occur due to the use of various scan
ners and examination parameters at various times lead to performance
degradation. Hence, to normalize the intensity variations intensity
normalization is used. This is performed by subtracting the gray value of Fig. 2. Feature extraction using Resnet50 with EFM.
the highest frequency from each voxel’s intensity and then dividing the
resulting deviation. This newly adjusted deviation, denoted as ̃ σ , applies (7)
to our MR image denoted as D, which consists of a set of voxels {d1 ,d2 ,d3 ,
…, dN }. The intensity value of each voxel, dk , is represented as Tk . The Fmʹ = F(Fm ; θm ), m ∈ {1, 2, 3, 4} (7)
calculation of ̃ σ is achieved using the Eq. (6)
where, Fmʹ represents the reduced feature maps of each encoder
√̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅
/ ̅
√ N block, F denotes 1 × 1 convolution function, and θm is the respective
√∑
σ=
̃ √ (Tk − T) ̂ 2
N (6) parameter. Edge features of varying granularities can be obtained by
k=1 subtracting average pooling values of various sizes from its local con
volutional feature maps. We specify P pooling operations in order to
where T̂ signifies the gray value of the highest frequency. Additionally, maintain generality as given in equation (8).
to enable the MR images to undergo processing similar to regular im ( )
Fm,l = Fʹm − avg p Fʹm , p ∈ {1, …, P} (8)
(p)
ages, we linearly transform their intensity range to span from 0 to 255.
th th
where F(p)
m signifies the edge features of the current m stage with the p
3.2. Feature extraction pooling operation, andavg p is the relevant average pooling operation.
The edge features are integrated by concatenating them with the fea
Local features are robust to image variations and also helps in dis tures of the current stage, followed by merging them with a convolution
tinguishing the tumor and non tumor region. With the ability to extract operation. This is expressed in Eq. (9).
both high and low-level features, it can learn more complex image ( ( ) )
Fm,l = F C F(1)
m,l , …, F m,l , F m ; θm,l
(P) ʹ
(9)
representations, the Resnet50 encoder is employed for the feature
extraction process.
where, Fm,l represents the feature maps which is the output of the
Fig. 2 shows the feature extraction process using Resnet50 with EFM.
EFM at the encoder network’s present stage, C denotes the concatena
The multi-level features from the pre-processed data are extracted using
tion process, θm,l refers to the corresponding parameter. The design of
ResNet50 encoder. We adopt individual encoders for each modality
extracting edge features offers an effective means of improving the
rather than fusing it to be robust in case of missing modalities. This
representation capability of the relevant level. By integrating edge in
design explicitly takes into account the relationships between the
formation at multiple granularities, the edge features are improved. The
different modalities and specifically extracts features from each mo
output maps are subsequently processed by the multi-task module to
dality. We employ EFM after each encoder block to learn and preserve
enhance the extraction of finer features.
the edge characteristics of each modality.
5
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
each stage includes two Swin Transformer blocks followed by a patch- window-based multi-head self-attention (SW-MSA). Non-overlapping
expanding module to acquire a large feature map. blocks of a specified size are used to divide the input volume in
The encoder has a Swin Transformer block and a patch merging W-MSA, while SW-MSA connects unrelated adjacent blocks using a
module. We use Swin Transformer block [39] because it incorporates a shifted window mechanism.
cleverly designed shifted window operation that sets it apart from other The feature map generated from the resnet50 is converted into non-
transformer-based models. Fig. 4 shows the internal connections of swin overlapping patches. To be more specific, we use patch size ps = 4 to
transformer block. Within the Swin Transformer Block, there are several obtain small-scale feature (fss ) and use ps = 8 those results in a high-
components, including a LayerNorm layer for normalization, residual resolution feature which is large-scale (fls ). Next, two separate encoder
connection, a MSA (multi-head self-attention module) and MLPs of two branches with several down sampling layers are fed with fss and fls . The
layer with activation functions. The MSA module in the swin trans feature dimension is increased at each step while the feature map’s
former Block contains two types of attention mechanisms: resolution is decreased through the use of a Swin Transformer layer and
window-based multi-head self-attention (W-MSA) and shifted a patch merging module [16]. The input patch is partitioned into four
6
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
segments, which are then combined using the patch merging module. leading to more accurate segmentation results, especially in complex
Thus, the multi-scale features are able to be represented hierarchically. medical imaging scenarios. The proposed EADFFTU-Net offers a holistic
During the first three stages of downsampling, the large-scale branch approach to brain tumor segmentation by combining advanced feature
produces outputs with resolutions that are twice as large as those of the extraction techniques, discriminative feature fusion, and state-of-the-art
small-scale branch. In comparison to the small-scale branch, the Transformer architecture. Compared to traditional methods that may
large-scale branch has an additional downsampling stage. rely solely on handcrafted features or shallow learning architectures, our
The decoder of DFFTU-Net comprises three upsampling stages. method achieves superior performance in terms of segmentation accu
Moreover, each stage includes Swin Transformer Block and a Patch racy and robustness.
Expanding module to acquire a large feature map. The output from the The multi-scale features captured by the dual encoders are selec
last Swin Transformer Block of the encoder is fed into the patch tively fused by DFM and then the fused features are transmitted into the
expanding module. The patch expanding module [37] has a rearrange decoder via skip connections.
operation and a linear layer in which the feature resolution is expanded The proposed DFFM consists of spatial as well as channel attention
by a rearrange operation, while the feature dimension is doubled by a followed by feature fusion as depicted in Fig. 5. The two features from
linear layer. the double encoder branch are denoted as Fl ∈ EC×(H×W) , and
Fs ∈ EC×(H/2×W/2) . The unsampled decoder feature is represented as
3.3.1. Discriminative feature fusion module Fde ∈ EC×(H×W) . We initially perform a transposed convolution layer on
We propose a Discriminative Feature Fusion Module, that integrate
Fs ∈ EC×(H×W) due to its poor resolution in order to produce the upsam
the multi-scale features effectively. The proposed method leverages
pled feature FUs ∈ EC×(H×W) . We concatenate Fde with FUs and Fl separately
discriminative feature fusion to effectively integrate multi-scale features
during the spatial attention phase. This allows us to create two spatial
from different modalities. Unlike traditional methods that may rely on
attention maps A1 ∈ EH×W and A2 ∈ EH×W by feeding the concatenated
simple fusion techniques like concatenation or summation, our method
features into point-wise convolution method. Then, the spatial attention
selectively fuses features using spatial and channel attention mecha
maps A1 and A2 are scaled into [0,1] using the sigmoid function. Spatial-
nisms. This allows the model to focus on relevant features while dis
wise multiplication is performed between spatial attention maps and the
carding irrelevant ones, leading to more accurate segmentation results.
input features for minimizing the noisy features while escalating rele
One of the major benefits of our method is its ability to preserve edge
filt
characteristics during feature fusion. The Edge Feature Module (EFM) is vant ones. The filtered features Fl and Ffilt s are subjected to global
specifically designed to capture edge features at different granularities average pooling during the channel attention phase to extract channel
and integrate them into the segmentation process. This ensures that the context descriptors Dl ∈ EC×1 and Ds ∈ EC×1 . Following that, we merge
model can accurately delineate tumor boundaries, which is crucial for them as Fd = [Dl , Ds ] ∈ E2C×1 and apply a Softmax function to the logits
precise segmentation. Traditional methods may struggle when faced across each channel as given in Eq. (10).
with missing modalities or incomplete data. Our method addresses this j j
challenge by employing individual encoders for each modality, allowing eDl eDs
ajm = j , bjm = j (10)
the model to extract features independently from each input. This design
j j
Dl
e +e Ds Dl
e + eDs
choice ensures robustness to missing modalities and enables the model
where ajm and bjm correspond to the jth element of their respective
to effectively utilize whatever information is available.
By incorporating Transformer architecture, particularly the Swin channel attention maps. Here we have ajm + bjm = [Link] jth element of Dl
Transformer blocks, our method benefits from their ability to capture and Ds are denoted by Djl and Djs . By combining the features from both
long-range dependencies and global context information. This enhances branches, we obtain an informative feature which is expressed in Eq.
the model’s understanding of spatial relationships within the input data, (11).
7
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
8
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
relied on the Adam optimizer. A learning rate decay scheme was inte Table 3
grated, reducing the rate by 0.5 when no improvement was observed Evaluation of our proposed method on BRATS 2019 training set.
within a 10-epoch window. The simulation parameters are provided in Metric Method WT TC ET
Table 2.
DSC Baseline 87.2 77.3 69.6
Additionally, early stopping measures were introduced to prevent Baseline+EFM 88.6 78.0 71.3
over-fitting, terminating training in cases where the validation loss Baseline+DFFM 88.9 77.9 72.4
failed to improve over 50 epochs. A random split is employed to divide Proposed 90.2 81.5 73.7
the dataset into 80% for training and 20% for testing. HD Baseline 8.0 11.5 8.8
Baseline+EFM 6.5 7.7 7.2
Baseline+DFFM 6.4 7.5 6.8
Proposed 4.9 4.6 5.4
5.2. Quantitative analysis Sensitivity Baseline 83.5 70.2 60.2
Baseline+EFM 86.7 79.6 64.8
We first conduct ablation experiments to assess our approach and Baseline+DFFM 90.2 79.9 65.6
show the efficacy of the suggested components on complete modalities, Proposed 95.5 81.4 71.3
9
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
Table 6 ensures that the model captures global contexts and fuses multiscale
Comparison of different techniques with the proposed method on BRATS2020 features effectively, contributing to improved segmentation accuracy.
validation set in terms of sensitivity, DSC, HD and specificity. According to the specified accuracy intervals, the proposed values for
Metric Author/year Methods WT TC ET EADFFTU-Net (DSC: WT: 93.7, TC: 85.5, ET: 80.6; sensitivity: WT: 94.3,
DSC Liew et al. [28], CASPIANET++ 81.08 83.70 78.50
TC: 92.8, ET: 90.7; specificity: WT: 97, TC: 97.1, ET: 97.9; HD: WT: 3.7,
2021 TC: 3.2, ET: 2.5) categorize it as a very good model, meeting or
Ullah et al. [40], MRA-UNet 87.22 86.74 90.18 exceeding the established thresholds for this classification.
2022
Ma et al. [41], Deep Supervision 78.79 75.93 87.26
i. Analysis of Missing Modalities Compared to Modern Techniques:
2020 CNN
Raza et al. [42], DresUNet 83.57 80.04 86.6 Table 7 and Table 8 provide a robust comparison of different brain
2023 tumor segmentation methods, across various combinations of avail
- Proposed 93.7 85.5 80.6 able modalities on both the BRATS 2019 and BRATS 2020 datasets.
Sensitivity Liew et al. [28], CASPIANET++ 89.19 79.53 77.11 The objective is to evaluate the robustness of each method in infer
2021
Ullah et al. [40], MRA-UNet 92.71 91.88 93.54
ring missing modalities, a common scenario in clinical practice.
2022
Ma et al. [41], Deep Supervision - - - The table illustrates the DSC percentages for different combinations
2020 CNN of available modalities (T1, T1c, T2, Flair) and tumor regions (WT, TC,
Raza et al. [42], DresUNet 97.86 96.95 96.48
ET). Our proposed method, denoted as "Ours," is compared against the
2023
Proposed 94.3 92.8 90.7 LCR method [21]. Across all 15 possible combination situations of un
Specificity Liew et al. [28], CASPIANET++ 99.82 99.97 99.92 available modalities, our method consistently outperforms both [21]
2021 and [43] as indicated by higher mean DSC values. This highlights the
Ullah et al. [40], MRA-UNet 93.19 95.64 92.61 exceptional robustness of our multimodal segmentation method,
2022
showcasing superior performance not only on complete modalities
Ma et al. [41], Deep Supervision - - -
2020 CNN datasets but also when dealing with missing modalities.
Raza et al. [42], DresUNet 99.36 98.62 97.91 Similar to the BRATS 2019 analysis, Table 8 presents the robust
2023 comparison on the BRATS 2020 dataset. The same combinations of
Proposed 97 97.1 97.9
available modalities are considered, and the performance of our pro
HD Liew et al. [28], CASPIANET++ - - -
2021 posed method is evaluated against Once again, our method consistently
Ullah et al. [40], MRA-UNet - - - outperforms the existing method, demonstrating strong robustness in
2022 inferring missing modalities. The mean DSC values further emphasize
Ma et al. [41], Deep Supervision - - - the reliability of our approach across diverse scenarios.
2020 CNN
Fig. 6 illustrates the performance of the brain tumor segmentation
Raza et al. [42], DresUNet - - -
2023 model in terms of loss over the course of training and validation epochs.
Proposed 3.7 3.2 2.5 Each epoch corresponds to one complete pass through the entire training
dataset. The loss value quantifies the difference between the predicted
segmentation and the ground truth segmentation. Lower loss values
fusing informative features, leading to a nuanced outcome. In accor
indicate better performance. Decreasing trend in both training and
dance with the suggested accuracy intervals, the proposed values for validation loss indicates that the model is learning and improving its
EADFFTU-Net (DSC: WT: 95.6, TC: 96.5, ET: 92.6; sensitivity: WT: 94.3, performance. Fig. 7 illustrates how well the proposed model classifies
TC: 95.8, ET: 90.7; specificity: WT: 97, TC: 98.7, ET: 99.3; HD: WT: 3.7, the input data in terms of accuracy over the course of training and
TC: 3.2, ET: 2.1) position it as an excellent model, surpassing the validation epochs. Accuracy indicates the proportion of correctly clas
thresholds indicative of very good/excellent models (>=90%). sified brain tumor segments compared to the total number of segments.
Turning to Table 6, evaluating performance on the BRATS2020 The training accuracy shows how the accuracy improves as the model
validation set, EADFFTU-Net continues to outperform existing methods. learns from the training data over epochs. The validation accuracy
The EFM maintains its effectiveness in learning and preserving edge represents measuring the model’s performance on unseen validation
characteristics, enhancing sensitivity to tumor boundaries. The DFFM [Link] increasing trend in both training and validation accuracy
Table 7
Robust comparison of different methods (DICE %) for different combinations of available modalities on BRATS 2019 dataset.
Modalities WT TC ET
• ◦ ◦ ◦
10.0 30.2 2.2 20.6 62.7 63.8
◦
• ◦ ◦
20.0 40.4 41.7 80.7 46.8 78.4
◦ ◦
• ◦
70 53.6 32.6 63.1 3.8 41.6
◦ ◦ ◦
• 74.3 82.4 52.0 62.5 17.0 40.8
• ◦
• ◦
70 75.8 42.1 71.6 1.2 49.9
• • ◦ ◦
20.7 80.6 39.3 82.8 42.3 73.2
◦ ◦
• • 83.9 89.7 41.8 67.9 11.3 46.5
◦
• ◦
• 81.7 87.9 74.7 82.7 63.5 77.8
• ◦ ◦
• 81.6 88.4 55.6 67.4 3.0 40.5
◦
• • ◦
73.5 85.9 66.8 82.7 60.4 75.7
• • • ◦
74.2 85.3 66.6 83.8 64.0 74.7
◦
• • • 88.9 93.2 76.6 80.7 64.4 75.6
• ◦
• • 88.2 91.1 52.1 63.4 2.9 50.5
• • ◦
• 83.9 89.8 75.7 83.2 67.1 79.6
• • • • 89.7 91.7 77.5 83.9 70.6 83.6
mean 67.37 77.73 53.15 71.8 38.7 63.48
10
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
Table 8
Robust comparison of different methods (DICE %) for different combinations of available modalities on BRATS 2020 dataset.
Modalities WT TC ET
• ◦ ◦ ◦
77.16 90.6 66.02 75.3 37.30 67.4
◦
• ◦ ◦
76.77 79.9 81.51 84.3 74.85 81.9
◦ ◦
• ◦
86.05 85.4 71.02 69.9 46.29 55.8
◦ ◦ ◦
• 87.32 92.5 69.19 80.8 38.15 43.1
• ◦
• ◦
87.73 90.7 73.13 89.7 45.65 52.6
• • ◦ ◦
81.12 87.4 83.40 91.1 78.01 84.5
◦ ◦
• • 89.87 90.2 74.14 87.6 49.32 74.2
◦
• ◦
• 89.89 93.2 84.65 84.5 76.67 83.1
• ◦ ◦
• 89.73 91.5 73.07 81.5 40.98 52.7
◦
• • ◦
87.74 92.6 83.45 92.8 75.93 81.3
• • • ◦
88.25 94.3 83.47 90.0 76.99 82.6
◦
• • • 90.68 92.4 84.97 82.4 77.12 85.2
• ◦
• • 90.60 92.7 75.19 91.1 49.92 57.5
• • ◦
• 90.69 91.8 85.07 88.5 76.81 82.5
• • • • 91.11 93.7 85.21 92.7 78.00 85.6
mean 86.98 89.98 78.23 83.46 61.47 72.15
5.4. Limitations
11
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
Fig. 8. Examples of the segmentation results on full modalities on BraTS 2020 dataset.
Fig. 9. The results of the segmentation process on modalities that are not available are displayed as follows: The necrotic and non-enhancing tumor core are
indicated by Red; the enhancing tumor is indicated by Blue; and edema is indicated by Green.
processing poses challenges in capturing fine-grained local details efficiency of the segmentation model. Moreover, the reliance on slice-
within tumors. While individual slices may contain informative features, based features may hinder the model’s ability to capture global char
they lack the depth and continuity present in volumetric data. Conse acteristics and inter-slice correlations, further impacting its segmenta
quently, the model may struggle to discern subtle variations in tumor tion accuracy. Overcoming these limitations may require exploring
characteristics or accurately distinguish between tumor and healthy alternative strategies, such as incorporating 3D convolutional neural
tissue, particularly in regions where the distinction is less clear. This networks or developing techniques to better integrate inter-slice
limitation becomes particularly pronounced in cases where tumors context, to enhance the model’s performance in analyzing volumetric
exhibit heterogeneous textures or irregular boundaries, as the model MRI data for brain tumor segmentation.
may fail to integrate information across slices effectively, resulting in
segmentation inaccuracies and potential misclassifications. 6. Conclusion
Additionally, the slice-based approach introduces computational
complexities, especially when dealing with large tumors or high- In this study, we introduced a novel approach named EA-DFFTU-Net,
resolution MRI scans. Slicing volumetric data into individual 2D im specifically designed for the precise segmentation of brain tumors, even
ages increases the number of input slices, requiring additional compu in scenarios involving missing modalities within multi-modal MRI data.
tational resources and memory. Consequently, processing large datasets Our approach seamlessly integrates various components to enhance
becomes more challenging, potentially limiting the scalability and segmentation accuracy and reliability. The incorporation of a ResNet-50
12
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
encoder enables robust multi-level feature extraction, capturing both and/or animals performed by any of the authors.
local spatial features and global contextual information. The innovative
EFM adds another dimension by preserving crucial edge properties Informed consent
within regions. The use of different encoders for different modalities and
the inclusion of the EFM after each encoder significantly enhance the There is no informed consent for this study.
robustness of our model in cases with missing modalities. The
Discriminative Feature Fusion-based Transformer U-Net (DFFTU-Net) Funding
further refines the segmentation by capturing global contexts and mul
tiscale features, while the DFFM effectively combines these features for There is no funding for this study.
optimal segmentation outcomes. Our comprehensive evaluation, con
ducted on benchmark datasets, including BraTS2019 and BraTS2020, CRediT authorship contribution statement
demonstrated that the proposed EA-DFFTU-Net consistently out
performed state-of-the-art models. This achievement underscores the B. Jagadeesh: Conceptualization, Formal analysis, Writing – orig
model’s potential to significantly improve the accuracy of brain tumor inal draft, Writing – review & editing. G. Anand Kumar: Conceptuali
segmentation, which holds promise for enhancing medical diagnosis and zation, Formal analysis, Methodology, Writing – original draft.
treatment planning. Nonetheless, it is essential to acknowledge the
study’s limitations. One notable constraint is that our model operates Declaration of Competing Interest
primarily in a 2D context, and when handling 3D MRI data, slicing
processes may inadvertently result in the loss of some context infor Authors declares that they have no conflict of interest.
mation and local details. To overcome this limitation, our future work
will explore the extension of EA-DFFTU-Net into a 3D network archi Data availability
tecture, thus allowing us to harness the full potential of 3D information
in MRI data for improved segmentation accuracy. Data will be made available on request.
Ethical approval
This article does not contain any studies with human participants
Appendix
The Table 1 offers a comprehensive overview of various deep learning methods employed for brain tumor segmentation, highlighting key details
such as the year of development, pre-processing techniques, the segmentation model used, the dataset involved, and the metrics utilized for evalu
ation. Each row corresponds to a distinct method, facilitating a quick and systematic comparison of their characteristics. Metrics range from traditional
ones like Dice score and Sensitivity to more specialized measures such as Hausdorff Distance and Specificity. Figure A1 likely visualizes the internal
connections or architecture of a Swin Block, an essential component in the reviewed deep learning models. This visual aid can provide additional
insights into the intricate details of the Swin Block structure, aiding readers in understanding the internal mechanisms of the associated deep learning
models.
13
B. Jagadeesh and G. Anand Kumar Applied Soft Computing 161 (2024) 111709
References [22] B. Yu, L. Zhou, L. Wang, W. Yang, M. Yang, P. Bourgeat, J. Fripp, SA-LuT-Nets:
learning sample-adaptive intensity lookup tables for brain tumor segmentation,
IEEE Trans. Med. Imaging 40 (5) (2021) 1417–1427.
[1] T. Saba, A.S. Mohamed, M. El-Affendi, J. Amin, M. Sharif, Brain tumor detection
[23] X. Wu, L. Bi, M. Fulham, D.D. Feng, L. Zhou, J. Kim, Unsupervised brain tumor
using fusion of hand crafted and deep learning features, Cogn. Syst. Res. 59 (2020)
segmentation using a symmetric-driven adversarial network, Neurocomputing 455
221–230.
(2021) 242–254.
[2] M.K. Abd-Ellah, A.I. Awad, A.A. Khalaf, H.F. Hamed, Two-phase multi-model
[24] C. Yan, J. Ding, H. Zhang, K. Tong, B. Hua, S. Shi, SEResU-Net for Multimodal
automatic brain tumour diagnosis system from magnetic resonance images using
Brain Tumor Segmentation, IEEE Access 10 (2022) 117033–117044.
convolutional neural networks, EURASIP J. Image Video Process. 2018 (1) (2018),
[25] Z. Huang, Y. Zhao, Y. Liu, G. Song, GCAUNet: A group cross-channel attention
1-0.
residual UNet for slice based brain tumor segmentation, Biomed. Signal Process.
[3] P. Wang, Y.H. Shi, J.Y. Li, C.Z. Zhang, Differentiating Glioblastoma from Primary
Control 70 (2021) 102958.
Central Nervous System Lymphoma: The Value of Shaping and Non enhancing
[26] Z. Xiao, K. He, J. Liu, W. Zhang, Multi-view hierarchical split network for brain
Peritumoral Hyper intense Gyral Lesion on FLAIR Imaging, World Neurosurg. 149
tumor segmentation, Biomed. Signal Process. Control 69 (2021) 102897.
(2021) e696–e704.
[27] Y. Zhang, Y. Lu, W. Chen, Y. Chang, H. Gu, B. Yu, MSMANet: A multi-scale mesh
[4] G. Anand Kumar, P.V. Sridevi, 3D Deep Learning for Automatic Brain MR Tumor
aggregation network for brain tumor segmentation, Appl. Soft Comput. 110 (2021)
Segmentation with T-Spline Intensity Inhomogeneity Correction, Autom. Control
107733.
Comput. Sci. (ACCS) 52 (5) (2018) 439–450.
[28] A. Liew, C.C. Lee, B.L. Lan, M. Tan, CASPIANET++: a multidimensional channel-
[5] D. Zhang, G. Huang, Q. Zhang, J. Han, J. Han, Y. Wang, Y. Yu, Exploring task
spatial asymmetric attention network with noisy student curriculum learning
structure for brain tumor segmentation from multi-modality MR images, IEEE
paradigm for brain tumor segmentation, Comput. Biol. Med. 136 (2021) 104690.
Trans. Image Process. 29 (2020) 9032–9043.
[29] H.M. Rai, K. Chatterjee, S. Dashkevich, Automatic and accurate abnormality
[6] H. Zhou, K. Chang, H.X. Bai, B. Xiao, C. Su, W.L. Bi, P.J. Zhang, J.T. Senders,
detection from brain MR images using a novel hybrid UnetResNext-50 deep CNN
M. Vallières, V.K. Kavouridis, A. Boaro, Machine learning reveals multimodal MRI
model, Biomed. Signal Process. Control 66 (2021) 102477.
patterns predictive of isocitrate dehydrogenase and 1p/19q status in diffuse low-
[30] L. Tan, W. Ma, J. Xia, S. Sarker, Multimodal magnetic resonance image brain tumor
and high-grade gliomas, J. neuro-Oncol. 142 (2) (2019) 299–307.
segmentation based on ACU-net network, IEEE Access 9 (2021) 14608–14618.
[7] G. Anand Kumar, P.V. Sridevi, Deep learning network with Euclidean factor for
[31] N.S. Syazwany, J.H. Nam, S.C. Lee, MM-BiFPN: Multi-Modality Fusion Network
Brain MR Tumor segmentation and volume estimation, Int. J. Model., Simul., Sci.
With Bi-FPN for MRI Brain Tumor Segmentation, IEEE Access 9 (2021)
Comput. (IJMSSC) 10 (17) (2019) 1950039.
160708–160720.
[8] R. Wang, T. Lei, R. Cui, B. Zhang, H. Meng, A.K. Nandi, Medical image
[32] Z. Luo, Z. Jia, Z. Yuan, J. Peng, HDC-Net: Hierarchical decoupled convolution
segmentation using deep learning: A survey, IET Image Process. 2016 (5) (2022)
network for brain tumor segmentation, IEEE J. Biomed. Health Inform. 25 (3)
1243–1267.
(2020) 737–745.
[9] Z. Stosic, P. Rutesic, An improved canny edge detection algorithm for detecting
[33] N. Micallef, D. Seychell, C.J. Bajada, Exploring the U-Net++ Model for Automatic
brain tumors in MRI images, Int. J. Signal Process. 3 (2018).
Brain Tumor Segmentation. IEEE Access 9 (2021) 125523–125539.
[10] M.H. Hesamian, W. Jia, X. He, P. Kennedy, Deep learning techniques for medical
[34] M. Aghalari, A. Aghagolzadeh, M. Ezoji, Brain tumor image segmentation via
image segmentation: achievements and challenges, J. Digit. Imaging 32 (4) (2019)
asymmetric/symmetric UNet based on two-pathway-residual blocks, Biomed.
582–596.
Signal Process. Control 69 (2021) 102841.
[11] N. Yamanakkanavar, J.Y. Choi, B. Lee, MRI segmentation and classification of
[35] K. Hao, S. Lin, J. Qiao, Y. Tu, A generalized pooling for brain tumor segmentation,
human brain using deep learning for diagnosis of Alzheimer’s disease: a survey,
IEEE Access 9 (2021) 159283–159290.
Sensors 20 (11) (2020) 3243.
[36] X. Li, Y. Xu, N. Li, B. Yang, Y. Lei, Remaining useful life prediction with partial
[12] G. Du, X. Cao, J. Liang, X. Chen, Y. Zhan, Medical image segmentation based on u-
sensor malfunctions using deep adversarial networks, IEEE/CAA J. Autom. Sin. 10
net: A review, J. Imaging Sci. Technol. 64 (2020) 1–2.
(1) (2022) 121–134.
[13] O. Çiçek, A. Abdulkadir, S.S. Lienkamp, T. Brox, O. Ronneberger, 3D U-Net:
[37] X. Li, S. Yu, Y. Lei, N. Li, B. Yang, Intelligent machinery fault diagnosis with event-
learning dense volumetric segmentation from sparse annotation. In International
based camera, IEEE Trans. Ind. Inform. (2023).
conference on medical image computing and computer-assisted intervention.
[38] N.J. Tustison, B.B. Avants, P.A. Cook, Y. Zheng, A. Egan, P.A. Yushkevich, J.C. Gee,
Springer Cham, (2016) pp. 424-432.
N4ITK: improved N3 bias correction, IEEE Trans. Med. Imaging 29 (6) (2010)
[14] X. Xiao, S. Lian, Z. Luo, S. Li, Weighted res-unet for high-quality retina vessel
1310–1320.
segmentation. In2018 9th international conference on information technology in
[39] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, (2021) Swin
medicine and education (ITME) IEEE, (2018) pp. 327-331.
transformer: Hierarchical vision transformer using shifted windows. In Proceedings
[15] Z. Zhou, M.M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, Unet++: A nested u-net
of the IEEE/CVF international conference on computer vision (2021) pp. 10012-
architecture for medical image segmentation. Deep learning in medical image
10022.
analysis and multimodal learning for clinical decision support, Springer, Cham,
[40] Z. Ullah, M. Usman, M. Jeon, J. Gwak, Cascade multiscale residual attention cnns
2018, pp. 3–11.
with adaptive roi for automatic brain tumor segmentation, Inf. Sci. 608 (2022)
[16] X. Liu, S. Hou, S. Liu, W. Ding, Y. Zhang, Attention-based multimodal glioma
1541–1556.
segmentation with multi-attention layers for small-intensity dissimilarity, J. King
[41] S. Ma, Z. Zhang, J. Ding, X. Li, J. Tang, F. Guo, A deep supervision CNN network for
Saud. Univ. -Comput. Inf. Sci. 35 (4) (2023) 183–195.
brain tumor segmentation. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and
[17] X. Liu, Y. Liu, W. Fu, S. Liu, SCTV-UNet: a COVID-19 CT segmentation network
Traumatic Brain Injuries: 6th International Workshop, BrainLes 2020, Held in
based on attention mechanism, Soft Comput. 21 (2023), 1-1.
Conjunction with MICCAI 2020, Lima, Peru, Revised Selected Papers, Part II 6
[18] S. Asgari Taghanaki, K. Abhishek, J.P. Cohen, J. Cohen-Adad, G. Hamarneh, Deep
2021 Springer International Publishing. (2020) pp. 158-167.
semantic segmentation of natural and medical images: a review, Artif. Intell. Rev.
[42] R. Raza, U.I. Bajwa, Y. Mehmood, M.W. Anwar, M.H. Jamal, dResU-Net: 3D deep
54 (2021) 137–178.
residual U-Net based brain tumor segmentation from multimodal MRI, Biomed.
[19] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser,
Signal Process. Control 79 (2023) 103861.
I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. (2017) 30.
[43] Y. Ding, X. Yu, Y. Yang, RFNet: Region-aware fusion network for incomplete multi-
[20] P. Shaw, J. Uszkoreit, A. Vaswani, Self-attention with relative position
modal brain tumor segmentation. In Proceedings of the IEEE/CVF international
representations. arXiv preprint arXiv (2018) 1803.02155.
conference on computer vision (2021) pp. 3975-3984.
[21] T. Zhou, S. Canu, P. Vera, S. Ruan, Latent correlation representation learning for
brain tumor segmentation with missing MRI modalities, IEEE Trans. Image Process.
30 (2021) 4263–4274.
14