Multi-task Learning for Breast Tumor Analysis
Multi-task Learning for Breast Tumor Analysis
A R T I C L E I N F O A B S T R A C T
Keywords: Segmentation and classification of breast tumors are critical components of breast ultrasound (BUS) computer-
Multi-task learning aided diagnosis (CAD), which significantly improves the diagnostic accuracy of breast cancer. However, the
Breast tumor characteristics of tumor regions in BUS images, such as non-uniform intensity distributions, ambiguous or
Ultrasound images
missing boundaries, and varying tumor shapes and sizes, pose significant challenges to automated segmentation
Segmentation
Classification
and classification solutions. Many previous studies have proposed multi-task learning methods to jointly tackle
tumor segmentation and classification by sharing the features extracted by the encoder. Unfortunately, this often
introduces redundant or misleading information, which hinders effective feature exploitation and adversely
affects performance. To address this issue, we present ACSNet, a novel multi-task learning network designed to
optimize tumor segmentation and classification in BUS images. The segmentation network incorporates a novel
gate unit to allow optimal transfer of valuable contextual information from the encoder to the decoder. In
addition, we develop the Deformable Spatial Attention Module (DSAModule) to improve segmentation accuracy
by overcoming the limitations of conventional convolution in dealing with morphological variations of tumors.
In the classification branch, multi-scale feature extraction and channel attention mechanisms are integrated to
discriminate between benign and malignant breast tumors. Experiments on two publicly available BUS datasets
demonstrate that ACSNet not only outperforms mainstream multi-task learning methods for both breast tumor
segmentation and classification tasks, but also achieves state-of-the-art results for BUS tumor segmentation. Code
and models are available at [Link]
1. Introduction and shape, irregular and blurred tumor boundaries, and low
signal-to-noise ratios in ultrasound images, as shown in Fig. 1.
Female breast cancer has surpassed lung cancer as the most Convolutional Neural Networks (CNNs) have achieved remarkable
commonly diagnosed cancer worldwide, and ranked fifth as the leading success in medical image analysis due to their powerful capability for
cause of cancer deaths worldwide [1]. Ultrasound imaging has become automatic feature extraction [4–10]. Specifically, in medical image
an important tool for breast cancer diagnosis due to its versatility, safety, segmentation tasks, UNet [11] is preferred for its ability to reconstruct
and sensitivity. However, evaluating a whole breast ultrasound (BUS) high-resolution segmentation images by integrating multi-level infor
examination slice by slice is highly time-consuming, even for experi mation. In BUS image segmentation, Almajalid et al. [12] enhanced the
enced radiologists [2], and there is a high inter-observer variation rate. performance of UNet in the precision of breast lesion segmentation by
Therefore, an effective computer-aided diagnosis (CAD) system is implementing strategies such as contrast enhancement and speckle noise
essential in assisting physicians with the diagnosis and treatment of reduction. Ning et al. [13] pointed out that a simple modeling frame
breast cancer. Segmentation and classification of BUS tumors are highly work is difficult to obtain ideal segmentation results on BUS images due
relevant tasks and are two fundamental objectives of CAD systems, to the complexity of the echo patterns of ultrasound images and the
because both share generic image features, such as shape and boundary interference from surrounding tissues. Additionally, Zhao et al. [14]
features [3]. Nevertheless, designing this system for BUS images is pointed out issues with the encoder-decoder model architecture,
challenging due to the problems such as large variations in tumor size emphasizing that while high-quality segmentation relies on useful
* Corresponding author.
E-mail address: yangqiuju@[Link] (Q. Yang).
[Link]
Received 11 July 2023; Received in revised form 3 March 2024; Accepted 12 March 2024
Available online 18 March 2024
0010-4825/© 2024 Elsevier Ltd. All rights reserved.
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
encoder features, direct transfer through skip connections could intro analysis within a unified network framework. This approach promotes
duce misleading contextual information, resulting in underutilization of the sharing of breast tumor-specific features, effectively improving the
valuable features and potential introduction of irrelevant information performance of both tasks and emerging as a promising avenue for
leading to pixel classification errors. Moreover, the feature extraction further investigation. However, most existing multi-task learning ar
capacity at a single level is limited, causing the global contextual in chitectures rely on an encoder-decoder approach and often overlook the
formation captured by the encoder at deeper levels to be progressively impact of redundant or misleading information within the encoder on
lost during the up-sampling process in the decoder [15]. the results. Such oversight can hinder the selective transfer of valuable
Similarly, the progress in deep learning has significantly advanced contextual information from the encoder to the decoder [14]. Addi
tumor classification in BUS imaging. CNNs adopt a hierarchical struc tionally, tumor boundaries in BUS images are often poorly defined and
ture of multiple convolutional and down-sampling layers to capture there is significant variability in tumor size and shape. This variability
deep contextual features, an architecture capable of recognizing pat requires a network that can adapt to tumors of different sizes.
terns ranging from simple edges to complex high-order arrangements. In this paper, we present ACSNet, an end-to-end multi-task learning
This allows the model to learn advanced semantic features in images and network tailored for the simultaneous learning of BUS tumor segmen
effectively identify diverse visual content. Notable studies in BUS tumor tation and classification tasks. The architecture consists of an encoder-
classification include Daoud et al.’s [16] investigation of deep feature decoder network dedicated to segmentation and a robust multi-scale
extraction and transfer learning techniques to enhance the classification feature aggregation network for benign and malignant classification.
accuracy. In addition, Saad et al. [17] introduced the BreastUS Trans To refine the information flow during skip connections, gate unit mod
former, a Transformer model that incorporates self-attention mecha ules for segmentation are designed to control the feature extraction
nisms specifically designed for BUS image classification. process of the encoder. Furthermore, a deformable spatial attention
Both the breast tumor segmentation and classification tasks rely module (DSAModule) is introduced to capture the complex morpho
heavily on the accurate identification and detailed interpretation of logical variations of breast tumors. In the classification task, a channel
common morphological features of tumors, including their shape, size, attention mechanism within the multi-scale feature extraction network
texture, and edge characteristics [3]. Proper delineation of tumor highlights salient features while mitigating the influence of extraneous
boundaries for segmentation and determination of their benign or ma information, thereby improving the accuracy of the network in dis
lignant nature for classification are closely related goals [18]. Accurate tinguishing between benign and malignant breast tumors. In addition, to
tumor segmentation can provide key morphological features that are suppress noise and further improve the multi-task learning performance,
essential for accurate classification, while a sound understanding of the LossNet, a generic no-training loss network proposed by Zhao et al. [23],
classification will aid in more accurate tumor segmentation. This is used to supervise the feature mapping extracted from the bottom to
interdependence and information sharing between tasks lead to higher the top layer, providing crucial structural detail supervision to the
accuracy and efficiency in breast cancer diagnosis. feature layers. We evaluate the effectiveness of the proposed method
However, single-task CNNs designed for BUS imaging may face through extensive experiments on two publicly available datasets. Our
limitations in scalability and dependence on substantial amounts of main contributions can be summarized as follows:
annotated data when tackling specific segmentation or classification
tasks. In contrast, multi-task learning (MTL), a well-established learning (1) We present ACSNet, a novel multi-task learning model for
paradigm in machine learning, has demonstrated its ability to improve simultaneous segmentation and classification of BUS tumors,
the performance of individual tasks by sharing information between which effectively addresses issues related to tumor size and shape
related tasks [19]. The inductive bias introduced by MTL serves as a variability and irregular boundaries in BUS images.
regularization mechanism, effectively preventing overfitting and pro (2) For tumor segmentation, we are developing the DSAModule and
moting the acquisition of more robust feature representations. In addi gate units. The DSAModule effectively adapts to morphological
tion, MTL proves beneficial in mitigating the challenges associated with changes, while the gate unit optimizes the information flow be
small datasets by facilitating knowledge transfer between different tasks tween the encoder and decoder, thereby enhancing the accuracy
and improving the overall generalizability of the model [20]. of segmentation.
For example, Shareef et al. [20] showed that multi-task learning (3) For classification, ACSNet integrates a channel attention mecha
outperformed single-task learning methods in the context of breast ul nism and a multi-scale feature fusion method, achieving excellent
trasound classification. Recent research on MTL [2,21,22] advocates the performance in breast cancer classification.
integration of segmentation and classification tasks in BUS image
2
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
(4) Our experiments on two BUS tumor datasets show that ACSNet 2.3. Multi-task learning for breast ultrasound image segmentation and
outperforms current mainstream multi-task learning methods for classification
both segmentation and classification tasks. Furthermore, ACSNet
achieves state-of-the-art results for BUS tumor segmentation. Many multi-task learning (MTL) methods have been used to improve
the results of both tumor classification and segmentation in BUS images.
The rest of the paper is organized as follows. Section 2 provides an Zhou et al. [2] developed a multi-task learning network for 3D auto
overview of breast ultrasound image segmentation, classification, and matic BUS images, utilizing a multi-scale feature connectivity network
multi-task learning for both tasks. Section 3 describes the proposed for tumor classification and an iterative feature refinement training
ACSNet in detail. Section 4 introduces the datasets used in our research strategy to focus on tumor regions. Xu et al. [46] proposed a
and provides experimental details. The experimental results are pre region-focused MTL framework (RMTL-Net) for simultaneous BUS
sented in Section 5. Further discussions are given in Section 6, and image segmentation and classification, using a region-attention (RA)
Section 7 concludes the study. module to automatically learn weighted category-sensitive information
in tumor, peri-tumor, and background regions. Wang et al. [47] pro
2. Related work posed a multi-feature-guided CNN architecture for enhancement, seg
mentation, and classification of bone surfaces in ultrasound images.
2.1. Breast ultrasound image segmentation Zhang et al. [21] proposed the use of an Attention Gate (AG) module to
integrate segmentation and classification tasks and improve the accu
BUS tumor segmentation methods can be broadly categorized into racy by effectively using lesion region information. Multi-task learning
traditional methods and deep learning-based approaches. Traditional has also shown advantages in other medical image analysis tasks. For
methods include region-based methods [24], deformable models [25, example, Chen et al. [48] proposed a fully automated framework for
26], graph-based methods [27,28], and learning-based methods [29, gadolinium-enhanced magnetic resonance image (GE-MRI) left atrial
30]. These methods rely heavily on hand-designed features and texture segmentation based on deep learning, Qu et al. [49] proposed a new
analysis to detect and segment breast lesion regions in ultrasound im multi-task learning method for cell nucleus segmentation and classifi
ages, which could lead to misidentification of breast lesions in complex cation within a unified framework.
environments [31].
In recent years, deep learning-based approaches have shown 3. Methodology
considerable progress in BUS medical image processing. Shareef et al.
[32] introduced row-column convolutional kernels that adapt to the Fig. 2 shows the proposed multi-task learning network, ACSNet,
anatomical structure of the breast and fused contextual information at which integrates the BUS image segmentation and classification into a
different scales to segment small breast tumors. Hu et al. [33] proposed unified end-to-end model. ACSNet takes a BUS image as input and
to combine an expanded convolutional network with a phase-based produces two outputs: a tumor segmentation map and an image-level
active contour model for automatic breast lesion region segmentation. classification probability.
Yap et al. [34] comparatively studied patched-based LeNet [35], UNet More specifically, the architecture of ACSNet consists of two main
[11], and migration learning methods for breast ultrasound lesion components: a segmentation network and a classification branch. The
detection. Zhu et al. [36] developed a second-order sub-region pooling core of the segmentation network is a U-shaped encoder-decoder
network for breast lesion segmentation by using second-order statistics structure that includes stages 1–5 for the encoder and stages 6–9 for
for multiple feature sub-regions. Vakanski et al. [37] proposed to inte the decoder. The final output stage of the decoder, stage 9, effectively
grate visual saliency into a deep learning model for breast tumor seg achieves the segmentation of tumor lesions in ultrasound images. At the
mentation in ultrasound images. He et al. [38] designed a cross same time, the classification branch utilizes features from stages 4, 5,
CNN-Transformer network, named HCTNet, for breast ultrasound lesion and 6 to classify breast tumors in these images as either benign or
segmentation. malignant.
To optimize the flow of information, gate units have been introduced
2.2. Breast ultrasound image classification at each skip connection between the encoder and decoder for the seg
mentation task. These gate units are designed to extract more robust
For BUS image classification, deep learning-based methods have features from the lesion areas during segmentation, thereby making
shown superior feature extraction capabilities when compared to more efficient use of the encoder features and generating more accurate
traditional machine learning methods. Qi et al. [39] proposed a deep predictive segmentation maps. In addition, the DSAModule is developed
CNN with multi-scale kernels and skip connections for diagnosing breast to enhance the spatial feature information of the lesions at different
ultrasonography images. Byra et al. [40] developed a Selective Kernel stages of the decoder, taking into account the variation in tumor size and
(SK) UNet, which utilized an attention mechanism to adjust the recep the interference caused by background noise.
tive field of the network and fused the feature maps extracted by dilation For the classification task, multi-scale features are extracted by
and traditional convolution. Qian et al. [41] proposed a multi-pathway merging feature maps from different scales, and a channel attention
deep learning architecture for automated breast cancer risk prediction, mechanism [50] is employed to enhance the learning capability of the
using multimodal and multi-view BUS images to mimic routine clinical network for feature representation. By jointly handling the segmenta
working processes and improve clinical applicability. Cui et al. [42] tion and classification tasks and utilizing common features, MTL en
constructed a CNN that adaptively fuses information from multiple hances the model’s ability to analyze and process breast tumors.
tumor regions for breast tumor classification in the testing process using Furthermore, this approach improves data efficiency and reduces the
images without segmentation masks. Zhuang et al. [43] utilized transfer dependence on large, task-specific datasets.
learning [44] methods with the original BUS images, as well as
enhanced and bilaterally filtered images, to improve breast lesion clas 3.1. Multi-task learning network
sification. Mo et al. [45] proposed a HoVer-Trans model that associates
Transformers with CNNs for breast cancer diagnosis in BUS images. The proposed ACSNet utilizes UNet [11] as the backbone architec
ture for multi-task learning due to its excellent performance in medical
image segmentation. UNet consists of encoding paths, decoding paths,
and skip connections between them. Skipping connections are essential
mechanisms for disseminating spatial information, bridging semantic
3
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
Fig. 2. Overview of the proposed multi-task learning network. The architecture consists of a segmentation network and a classification branch. The segmentation
network includes gate units to regulate the features extracted from different encoder levels, while the DSAModules enhance the network’s adaptability to tumor
deformations. The input to the classification branch consists of tumor features at different scales, which are aggregated using the channel attention mechanism.
gaps, and refining segmentation results. These connections greatly connection in ACSNet, allowing for precise control and optimization of
enhance the model’s ability to capture and utilize global and local in the feature information flowing from the encoder. In contrast to the
formation, making UNet a powerful architecture for medical image approach of Zhao et al. [14], which evaluates the importance of different
segmentation tasks. Given UNet’s excellent performance in segmenta stages by combining encoder and decoder features, our gate unit em
tion tasks, strong feature fusion capability, adaptability to different ploys a distinctive combination of Global Average Pooling (GAP) and
shapes, effectiveness on limited data, and previous success in related Global Max Pooling (GMP). This approach enhances the selectivity of
works [22,40], we make it the backbone of ACSNet. As shown in the the feature control process, facilitating effective identification and
segmentation network in Fig. 2, UNet contains convolutional layers at retention of key features. Consequently, our method significantly im
each stage. The encoding path employs four down-sampling operations proves the overall multi-task learning performance.
to extract high-level semantic features, while the decoding path uses Fig. 3 shows the designed gate unit. The gate unit allows the infor
four upsampling operations to restore the feature maps to the original mation selection of the input features and improves the effectiveness of
input size; in addition, the feature maps obtained from the encoder are the output features. The gate unit takes the convolutional feature map
connected to the decoder using skip connections to propagate spatial X ∈ RC×H×W as input, where H and W are the spatial height and width,
information and refine the segmentation results. Each convolution and C is the number of channels. With the convolution operation, the
operation is followed by group normalization (GN) and ReLU. input X is first divided into two feature maps F1 and F2 , (F1 ,
Image classification networks mostly use the high-level feature maps F2 ∈ R1×H×W ), which allows the gate unit to pay attention to effective
of CNNs such as VGG [51] and ResNet [52]. Inspired by this, we add a information from different representation subspaces, and then F1 and F2
classification branch at the bottom of UNet. First, the feature maps from are normalized to the interval [0,1] using the Sigmoid function to obtain
stages 4, 5, and 6 of UNet are passed into the classification network. F1′ and F2′. A pair of gate values G1 and G2 are computed by using the
Then, the shared feature maps from these stages are fused using the
global average pooling and global maximum pooling operations on F1′
channel attention (CA) mechanism [50] to explicitly model in
terdependencies between channels and adaptively recalibrate the and F2′. Finally, G1 and G2 are used to select the contextual information
feature responses of the channels. Finally, the fused features are input from the input features respectively, and to sum and fuse the informa
into two fully connected (FC) layers and a Softmax layer to predict the tion to obtain the output feature Y ∈ RC×H×W . In this way, the gate unit
benign and malignant categories of the input images. can selectively send the features from the encoder to the decoder during
the skip connection process. The whole calculation process for the gate
unit can be expressed as follows:
3.2. Gate control mechanism
⎛ ⎞
To overcome the problem of redundant or misleading information in ⎜
G1 = GAP⎝Sigmoid(Conv(X))⎠,
⎟
(1)
the encoder during multi-task learning, several solutions have been ⏟̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏞⏞̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏟
F′
proposed by researchers. Among them, the application of gating mech 1
anisms plays a crucial role, not only by improving the quality of feature ⎛ ⎞
extraction, but also by enhancing the adaptability of the models. For ⎜ ⎟
example, the Gated Dual-branch Network (GateNet) [14] was designed G2 = GMP⎝Sigmoid(Conv(X))⎠, (2)
⏟̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏞⏞̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏟
to optimize the transfer of contextual information from the encoder to F′2
the decoder. Inspired by this, we have integrated gate units at each skip
4
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
Fig. 3. Overview of the proposed gate unit. The input feature map is directed to two branches with spatial attention mechanisms. These branches utilize GAP and
GMP, respectively, to extract relevant feature details. After summing the feature maps from both branches, a refined output is generated via a 3 × 3 Conv.
Fig. 4. Overview of the proposed DSAModule. The input feature maps first pass through a deformable convolution to preliminarily extract the features of the tumor
regions. These features are then fed into a spatial attention module to refine relevant feature details, and finally processed through another 3 × 3 deformable
convolution to obtain enhanced output features.
5
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
6
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
7
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
Table 2
Quantitative results on BUSI for all networks constructed in ablation study. The optimal results are shown in bold.
Model Segmentation (Mean ± Std) Classification (Mean ± Std)
Dice (%) ↑ Jaccard (%) ↑ 95HD ↓ ASD ↓ ACC (%) ↑ PRE (%) ↑ REC (%) ↑
observed in CUANet and CAUNet, as compared to CUNet, are due to their 5.2. Comparison with the state-of-the-arts
respective attentional mechanisms, which contribute to performance
enhancements by focusing on critical features, thereby refining the Table 3 compares the proposed ACSNet with ten advanced methods,
overall model accuracy and effectiveness. Furthermore, CMSUNet, unlike including six recent single-task segmentation methods (FCN [61], UNet
CUNet, employs multi-scale feature extraction for classification, result [11], DeeplabV3+ [62], FPN [63], Unet++ [64], and TransUnet [65])
ing in superior performance over both ClsNet and CUNet in classification and four state-of-the-art MTL methods (CUNet, Wang et al. [47], Chen
tasks. This result highlights the effectiveness of multi-scale feature et al. [48], and Qu et al. [49]) on the BUSI dataset. CUNet is the model
extraction and provides a compelling argument for its superiority over discussed in ablation study. For a fair comparison, we used the same
single-scale approaches. CMSUNet’s ability to capture a wider range of backbone (UNet) as the proposed method in all compared methods, and
feature details allows for more accurate classification, further validating applied five-fold cross-validation to all models.
the model’s robustness and versatility. Table 4 reports the mean and standard deviation values of the
Fig. 6 visually compares the segmentation results of various methods
in the ablation experiment. This paper designs a gate unit and a DSA
Table 3
Module for tumor segmentation. From Fig. 6, we can conclude that the
Briefly comparing the proposed ACSNet with 10 advanced methods.
segmentation performance of CUGNet has greatly improved compared to
Methods Backbone Tasks Feature
CUNet, because its gate unit can select and control the feature infor
Enhancement
mation in the encoder; Also, the segmentation accuracy of CUANet is
superior to CUNet, and over and under-segmentation problems are Single Multi- Attention
Segmentation Task Mechanism
significantly reduced, indicating that the developed DSAModule can
learn the location information of lesions in the feature map. CUGANet UNet [11] ResNet-18 ✓
FCN [61] ResNet-18
combines the strengths of the gate unit and DSAModule, further ✓
DeepLabV3+ ResNet-18 ✓
improving the segmentation performance of CUGNet and CUANet, [62]
resulting in more accurate lesion location and improved boundaries. In FPN [63] ResNet-18 ✓
addition, CUGANet achieves significantly better segmentation results Unet++ [64] ResNet-18 ✓
than CAUNet, CMSUNet, and CAMSUNet, which do not contain the gate TransUnet [65] R50+ViT- ✓
B_16
unit and DSAModule. All the results demonstrate that the effectiveness CUNet ResNet-18 ✓
of the proposed gate unit and DSAModule in breast ultrasound image Wang et al. ResNet-18 ✓ ✓
segmentation. [47]
Qu et al. [49] ResNet-18 ✓ ✓
Chen et al. [48] ResNet-18 ✓ ✓
ACSNet ResNet-18 ✓ ✓
8
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
Table 4 encoding the tokenized image patches from the feature maps of CNNs as
Classification performance (Mean ± SD) of all compared methods on BUSI. the input sequence for extracting the global semantics, while the
Methods ACC (%) ↑ PRE (%) ↑ REC (%) ↑ decoder upsamples the encoded features and then combines them with
the high-resolution CNN feature maps for precise localization. Deep
CUNet 92.54 ± 2.87 92.92 ± 2.39 92.20 ± 2.04
Wang et al. [47] 92.86 ± 2.84 92.93 ± 2.93 92.13 ± 3.61 labV3+ obtained the second-best overall segmentation performance,
Qu et al. [49] 94.13 ± 1.98 94.51 ± 1.48 93.58 ± 1.36 followed by FPN. Among the five MTL methods, the proposed ACSNet
Chen et al. [48] 92.22 ± 1.91 92.32 ± 1.88 92.25 ± 1.89 achieved the best segmentation performance on all four metrics, with
ACSNet 94.44 ± 2.07 94.61 ± 1.56 93.86 ± 2.33 the lowest 95HD and ASD of 1.65 mm and 0.66 mm, respectively, while
Dice and Jaccard were 1.08% and 1.50% higher than the sub-optimal
classification metrics (ACC, PRE, and REC) among ACSNet and all the method (i.e., TransUnet), respectively.
competitors on BUSI. Among the five MTL methods, the proposed In addition, Table 5 also provides a clear comparison of floating-
ACSNet achieves the best performance in all metrics. Low-level features point operations (FLOPs) and the number of parameters (in millions)
are usually used for classification because they mainly capture shape for each model. ACSNet has a computational complexity of 30.96 G
and boundary information, while high-level features extract semantic FLOPs and 42.99 M parameters. Although ACSNet’s parameters are the
features of the targets. Capturing the features of small objects is chal largest among these models, its model complexity is low, and in terms of
lenging when the network depth increases with more convolution and results, ACSNet achieves optimal results on the task of breast tumor
down-sampling operations. To address this issue, ACSNet combines and segmentation and classification from ultrasound images, which we
fuses the multi-scale features for the classification task. Specifically, believe strikes a favorable balance between performance and efficiency.
ACSNet inputs the feature maps from stage 4 to stage 6 of the segmen The segmentation results of ACSNet and all competitors on BUSI are
tation network to the classification branch, which enables the multi- given in Fig. 7. Among the single-task segmentation methods (Fig. 7(d)–
scale features to contain more useful information through channel (i)), TransUnet has the best overall segmentation performance due to the
attention CA, thus achieving the best classification performance on the advantage of using the Transformer to obtain global information, while
BUSI. the CNN models (Fig. 7(d)–(h)) tend to ignore breast lesion details or
Table 5 summarizes the segmentation results of ACSNet and ten misclassify lesion regions as non-breast lesions in the predicted seg
compared methods on Dice, Jaccard, 95HD, and ASD on the BUSI, mentation maps. The multi-task models shown in Fig. 7(j)–(m) have
including six single-task segmentation methods (i.e., UNet, FCN, FPN, poor results for tumor boundary segmentation and suffer from more
DeepLabV3+, UNet++ and TransUnet) and five MTL methods (i.e., serious over- or under-segmentation problems, which are due to the fact
CUNet, Wang et al. [47], Chen et al. [48], Qu et al. [49], and ACSNet). that the tumor morphology is highly variable and the features extracted
Among the six single-task segmentation methods, TransUnet obtained by the encoder are not fully suitable for the segmentation task. The
the highest Jaccard and Dice values, the lowest 95HD and ASD values, proposed ACSNet achieved the best segmentation performance among
and the best overall segmentation performance on BUSI, because it the five MTL methods and provides more accurate segmentation of the
combines the advantages of Transformer and UNet, with Transformer breast lesion region, with fewer over- or under-segmentation errors.
Table 6 summarizes the classification results of ACSNet and four
compared methods on BUS. It is clear that ACSNet has the highest ac
Table 5 curacy, precision, and recall and improves over the second-best method
Segmentation performance of all compared methods on BUSI. (i.e., Qu et al. [49]) by 0.63%, 0.13%, and 3.4%, respectively.
Methods Dice Jaccard 95HD ↓ ASD ↓ FLOPs Param Table 7 reports the mean and standard deviation values of the four
(%) ↑ (%) ↑ (G) (M) segmentation metrics among our ACSNet and ten compared methods on
UNet [11] 83.04 76.34 ± 15.95 4.14 25.69 29.38 the dataset BUS. Among the single-task segmentation methods (UNet,
± 1.64 2.04 ± 2.29 ± FCN, FPN, DeepLabV3+, UNet++, and TransUnet), TransUnet has the
0.67
best overall segmentation performance, with the highest values on Dice
FCN [61] 81.46 73.90 ± 16.14 4.50 9.54 22.35
± 1.50 1.66 ± 1.69 ± and Jaccard and the lowest values on 95HD. Among the five MTL
0.36 methods (CUNet, Wang et al. [47], Chen et al. [48], Qu et al. [49], and
FPN [63] 82.68 75.88 ± 15.98 4.26 14.41 28.91 ACSNet), ACSNet achieved the best segmentation performance on all
± 0.69 0.72 ± 1.27 ± four metrics, with increased Dice and Jaccard values over the
0.68
DeepLabV3+ 83.34 76.30 ± 14.87 4.26 15.08 27.78
second-best method (TransUnet) by 1.26% and 1.66%, respectively,
[62] ± 0.96 0.88 ± 1.18 ± while reducing 95HD and ASD by 0.63 mm and 0.69 mm, respectively.
0.64 Fig. 8 shows the visual segmentation results of different segmenta
UNet++ [64] 82.80 76.04 ± 14.83 4.18 35.20 16.80 tion methods on the dataset BUS. The proposed ACSNet effectively
± 0.87 1.03 ± 0.67 ±
mitigated the influence of tumor size, surrounding tissue, and shadow
1.26
TransUnet 83.82 77.12 ± 14.69 4.11 32.23 93.23 regions, providing segmentation results closer to the ground truth and
[65] ± 0.89 0.84 ± 1.68 ± fewer missed and false detections. The comprehensive evaluation results
1.09 and visual effects demonstrate the effectiveness of the proposed ACSNet
CUNet 83.56 76.36 ± 15.22 4.01 25.69 29.12 for breast ultrasound image segmentation and classification.
2.19 ±
± 1.72 ± 1.59
1.13
Wang et al. 82.76 76.30 ± 14.39 4.45 37.86 29.38 6. Discussion
[47] 2.03 ± 2.79
± 1.62 ±
1.76 BUS tumor segmentation and classification face significant chal
Chen et al. 82.38 75.48 ± 16.51 5.20 27.90 34.72 lenges, such as variations in tumor size and shape, irregular and blurred
[48]
± 0.82
0.47 ± 0.67 ± tumor boundaries, and low signal-to-noise ratio of ultrasound images
1.52
[2]. To address these challenges, we propose ACSNet, a novel multi-task
Qu et al. [49] 82.28 75.48 ± 17.71 5.41 5.82 14.55
± 0.99 1.03 ± 2.12 ± learning model specifically designed for tumor segmentation and clas
1.45 sification in BUS images. ACSNet starts by extracting basic patterns and
ACSNet 84.90 78.62 ± 13.04 3.45 30.96 42.99 features at shallow levels and gradually progresses to complex,
± 1.69 1.75 ± 2.58 ± high-level semantic feature extraction. As the model’s receptive field
1.59
expands, it not only performs more accurate tumor feature analysis, but
9
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
Fig. 7. Visualization of segmentation results of different methods on the dataset BUSI, (a)–(m) from left to right are (a) input image, (b) ground truth, and (c) the
segmentation results of ACSNet (Ours), (d) UNet, (e) Unet++, (f) FPN, (g) FCN, (h) DeeplabV3+, (i) TransUnet, (j) CUNet, (k) Wang et al. [47], (l) Chen et al. [48],
and (m) Qu et al. [49] respectively.
Fig. 8. Visualization of segmentation results of different methods on BUS, (a)–(m) from left to right are (a) input image, (b) ground truth, and (c) the segmentation
results of ACSNet (Ours), (d) UNet, (e) Unet++, (f) FPN, (g) FCN, (h) DeeplabV3+, (i) TransUnet, (j) CUNet, (k) Qu et al. [49], (l) Chen et al. [48], and (m) Wang
et al. [47] respectively.
10
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
11
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319
[26] A. Madabhushi, D.N. Metaxas, Combining low-, high-level and empirical domain aware HoVer-transformer for ROI-free breast cancer diagnosis in ultrasound
knowledge for automated segmentation of ultrasonic breast lesions, IEEE Trans. images, IEEE Trans. Med. Imag. (2023) 1, 1.
Med. Imag. 22 (2003) 155–169. [46] M. Xu, K. Huang, X. Qi, A regional-attentive multi-task learning framework for
[27] C.-M. Chen, H.H.-S. Lu, Y.-S. Huang, Cell-based dual snake model: a new approach breast ultrasound image segmentation and classification, IEEE Access 11 (2023)
to extracting highly winding boundaries in the ultrasound images, Ultrasound Med. 5377–5392.
Biol. 28 (2002) 1061–1073. [47] P. Wang, V.M. Patel, I. Hacihaliloglu, Simultaneous Segmentation and
[28] M. Xian, Y. Zhang, H.D. Cheng, Fully automatic segmentation of breast ultrasound Classification of Bone Surfaces from Ultrasound Using a Multi-Feature Guided
images based on breast characteristics in space and frequency domains, Pattern CNN, Springer International Publishing, 2018, pp. 134–142.
Recogn. 48 (2015) 485–497. [48] C. Chen, W. Bai, D. Rueckert, Multi-task Learning for Left Atrial Segmentation on
[29] X. Guofang, M. Brady, J.A. Noble, Z. Yongyue, Segmentation of ultrasound B-mode GE-MRI, Springer International Publishing, 2019, pp. 292–301.
images with intensity inhomogeneity correction, IEEE Trans. Med. Imag. 21 (2002) [49] H. Qu, G. Riedlinger, P. Wu, Q. Huang, J. Yi, S. De, D. Metaxas, Joint Segmentation
48–57. and Fine-Grained Classification of Nuclei in Histopathology Images, IEEE.
[30] B. Liu, H.D. Cheng, J. Huang, J. Tian, X. Tang, J. Liu, Fully automatic and [50] J. Hu, L. Shen, G. Sun, Squeeze-and-Excitation networks, in: 2018 IEEE/CVF
segmentation-robust classification of breast tumors based on local texture analysis Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141.
of ultrasound images, Pattern Recogn. 43 (2010) 280–298. [51] K. Simonyan, A. Zisserman, Very Deep Convolutional Networks for Large-Scale
[31] C. Xue, L. Zhu, H. Fu, X. Hu, X. Li, H. Zhang, P.-A. Heng, Global guidance network Image Recognition, 2014 arXiv preprint arXiv:1409.1556.
for breast lesion segmentation in ultrasound images, Med. Image Anal. 70 (2021) [52] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in:
101989. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016,
[32] B. Shareef, A. Vakanski, P.E. Freer, M. Xian, ESTAN: enhanced small tumor-aware pp. 770–778.
network for breast ultrasound image segmentation, Healthcare 10 (2022) 2262. [53] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, Y. Wei, Deformable convolutional
[33] Y. Hu, Y. Guo, Y. Wang, J. Yu, J. Li, S. Zhou, C. Chang, Automatic tumor networks, IEEE.
segmentation in breast ultrasound images using a dilated fully convolutional [54] F. Yuan, L. Zhang, X. Xia, Q. Huang, X. Li, A gated recurrent network with dual
network combined with an active contour model, Med. Phys. 46 (2019) 215–228. classification assistance for smoke semantic segmentation, IEEE Trans. Image
[34] M.H. Yap, G. Pons, J. Marti, S. Ganau, M. Sentis, R. Zwiggelaar, A.K. Davison, Process. 30 (2021) 4409–4422.
R. Marti, Automated breast ultrasound lesions detection using convolutional neural [55] P.-T. De Boer, D.P. Kroese, S. Mannor, R.Y. Rubinstein, A tutorial on the cross-
networks, IEEE Journal of Biomedical and Health Informatics 22 (2018) entropy method, Ann. Oper. Res. 134 (2005) 19–67.
1218–1226. [56] W. Al-Dhabyani, M. Gomaa, H. Khaled, A. Fahmy, Dataset of breast ultrasound
[35] Y. LeCun, B.E. Boser, J.S. Denker, D. Henderson, R.E. Howard, W.E. Hubbard, L. images, Data Brief 28 (2020) 104863.
D. Jackel, Handwritten Digit Recognition with a Back-Propagation Network, NIPS, [57] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin,
1989. N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison,
[36] L. Zhu, R. Chen, H. Fu, C. Xie, L. Wang, L. Wan, P.-A. Heng, A Second-Order A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala, PyTorch: an
Subregion Pooling Network for Breast Lesion Segmentation in Ultrasound, Springer imperative style, high-performance deep learning library, in: Proceedings of the
International Publishing, 2020, pp. 160–170. 33rd International Conference on Neural Information Processing Systems, Curran
[37] A. Vakanski, M. Xian, P.E. Freer, Attention-enriched deep learning model for breast Associates Inc., 2019. Article 721.
tumor segmentation in ultrasound images, Ultrasound Med. Biol. 46 (2020) [58] B. Chen, Y. Liu, Z. Zhang, G. Lu, A.W.K. Kong, TransAttUnet: multi-level attention-
2819–2833. guided U-net with transformer for medical image segmentation, IEEE Transactions
[38] Q. He, Q. Yang, M. Xie, HCTNet: a hybrid CNN-transformer network for breast on Emerging Topics in Computational Intelligence (2023) 1–14.
ultrasound image segmentation, Comput. Biol. Med. 155 (2023) 106629. [59] H.Y. Zhou, J. Guo, Y. Zhang, X. Han, L. Yu, L. Wang, Y. Yu, nnFormer: volumetric
[39] X. Qi, L. Zhang, Y. Chen, Y. Pi, Y. Chen, Q. Lv, Z. Yi, Automated diagnosis of breast medical image segmentation via a 3D transformer, IEEE Trans. Image Process. 32
ultrasonography images using deep neural networks, Med. Image Anal. 52 (2019) (2023) 4036–4045.
185–198. [60] B.N. Li, C.K. Chui, S. Chang, S.H. Ong, A new unified level set method for semi-
[40] M. Byra, P. Jarosik, A. Szubert, M. Galperin, H. Ojeda-Fournier, L. Olson, automatic liver tumor segmentation on contrast-enhanced CT images, Expert Syst.
M. O’Boyle, C. Comstock, M. Andre, Breast mass segmentation in ultrasound with Appl. 39 (2012) 9661–9668.
selective kernel U-Net convolutional neural network, Biomed. Signal Process [61] E. Shelhamer, J. Long, T. Darrell, Fully convolutional networks for semantic
Control 61 (2020) 102027. segmentation, IEEE Trans. Pattern Anal. Mach. Intell. 39 (2017) 640–651.
[41] X. Qian, J. Pei, H. Zheng, X. Xie, L. Yan, H. Zhang, C. Han, X. Gao, H. Zhang, [62] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder-Decoder with
W. Zheng, Q. Sun, L. Lu, K.K. Shung, Prospective assessment of breast cancer risk Atrous Separable Convolution for Semantic Image Segmentation, Springer
from multimodal multiview ultrasound images via clinically applicable deep International Publishing, 2018, pp. 833–851.
learning, Nat. Biomed. Eng. 5 (2021) 522–532. [63] T.Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, S. Belongie, Feature pyramid
[42] W. Cui, Y. Peng, G. Yuan, W. Cao, Y. Cao, Z. Lu, X. Ni, Z. Yan, J. Zheng, FMRNet: a networks for object detection, in: 2017 IEEE Conference on Computer Vision and
fused network of multiple tumoral regions for breast tumor classification with Pattern Recognition (CVPR), 2017, pp. 936–944.
ultrasound images, Med. Phys. 49 (2022) 144–157. [64] Z. Zhou, M.M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, UNet++: A Nested U-Net
[43] Z. Zhuang, Y. Kang, A.N. Joseph Raj, Y. Yuan, W. Ding, S. Qiu, Breast ultrasound Architecture for Medical Image Segmentation, Springer International Publishing,
lesion classification based on image decomposition and transfer learning, Med. 2018, pp. 3–11.
Phys. 47 (2020) 6257–6269. [65] J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, Alan, Y. Zhou, arXiv Pre-
[44] S.J. Pan, Q. Yang, A survey on transfer learning, IEEE Trans. Knowl. Data Eng. 22 print Server, in: TransUNet: Transformers Make Strong Encoders for Medical Image
(2010) 1345–1359. Segmentation, 2021.
[45] Y. Mo, C. Han, Y. Liu, M. Liu, Z. Shi, J. Lin, B. Zhao, C. Huang, B. Qiu, Y. Cui, L. Wu,
X. Pan, Z. Xu, X. Huang, Z. Li, Z. Liu, Y. Wang, C. Liang, HoVer-trans: anatomy-
12
ACSNet outperforms other multi-task learning methods such as CUNet, Wang et al., Chen et al., and Qu et al. in both segmentation and classification tasks. It achieves the highest values in Dice and Jaccard coefficients, with improvements of 1.26% and 1.66% over TransUnet respectively, along with reduced 95HD and ASD values. For classification, ACSNet improves mean accuracy to 94.44% ± 2.07, surpassing other methods .
The multi-scale feature fusion method in ACSNet allows the integration of features at various scales, providing comprehensive context and detailed information for both segmentation and classification tasks. This fusion enables the model to capture essential information from different resolution levels, enhancing both spatial resolution for segmentation and feature depth for classification, thereby significantly improving overall model performance .
Breast ultrasound image segmentation and classification face challenges such as variations in tumor size and shape, irregular and blurred tumor boundaries, and low signal-to-noise ratio. ACSNet addresses these challenges by employing a multi-task learning model that effectively extracts and utilizes both low-level and high-level features. It integrates a DSAModule to adapt to morphological changes and a gate unit to optimize feature information flow. Additionally, it employs channel attention mechanisms and multi-scale feature fusion to enhance classification accuracy .
Transfer learning techniques in models for breast ultrasound image segmentation, such as those compared to ACSNet, allow these models to leverage pre-trained knowledge from large datasets to improve segmentation performance on smaller datasets. This approach helps in better feature representation and reduces the need for extensive training data, leading to improvements in segmentation accuracy as observed in methods like UNet++ and TransUnet .
The channel attention mechanism in ACSNet enhances the model's classification performance by focusing on the most informative features across channels. This mechanism helps to suppress less important features, leading to more accurate feature extraction and improved classification results. It plays a critical role in achieving better accuracy and precision in identifying breast cancer from ultrasound images .
The DSAModule in ACSNet adapts to the morphological variability of breast tumors, which is crucial for accurate segmentation. It helps the model to dynamically adjust its segmentation strategies according to the shape and size changes of tumors, improving overall segmentation accuracy by handling the elusive nature of tumor boundaries effectively .
ACSNet exploits contextual information effectively by integrating both low-level and high-level features through its multi-scale feature fusion and channel attention mechanisms. It uses the context surrounding tumors to better differentiate between actual tumor boundaries and similar background patterns, enhancing accuracy in detection and classification .
The primary contributions of ACSNet include the development of a novel multi-task learning model that simultaneously performs segmentation and classification more effectively than previous methods. It introduces DSAModule for morphological adaptation, gate units for optimized information flow, and integrates channel attention and multi-scale feature fusion to enhance classification. These innovations address issues like tumor size, shape variability, and irregular tumor boundaries more efficiently than traditional methods .
Experimental evidence for ACSNet's state-of-the-art results in BUS tumor segmentation includes achieving the highest Dice and Jaccard values among tested models, with values increased by 1.26% and 1.66% over TransUnet respectively. It also demonstrated reductions in 95HD and ASD by 0.63 mm and 0.69 mm. Visual results showed fewer false detections and better alignment with the ground truth .
Using LossNet in multi-task learning frameworks like ACSNet provides structural detail supervision without requiring additional training, which helps suppress noise and improves feature mapping. This enhances overall learning performance by preserving important structural features, leading to better segmentation and classification outcomes by maintaining the balance between different task demands .