0% found this document useful (0 votes)
7 views12 pages

Multi-task Learning for Breast Tumor Analysis

The document discusses ACSNet, a novel multi-task learning network designed for the segmentation and classification of breast tumors from ultrasound images. It addresses challenges such as non-uniform intensity distributions and varying tumor shapes by incorporating a gate unit and a Deformable Spatial Attention Module to enhance feature extraction and improve accuracy. Experimental results demonstrate that ACSNet outperforms existing methods, achieving state-of-the-art results in breast tumor segmentation and classification tasks.

Uploaded by

Santiago Moreno
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views12 pages

Multi-task Learning for Breast Tumor Analysis

The document discusses ACSNet, a novel multi-task learning network designed for the segmentation and classification of breast tumors from ultrasound images. It addresses challenges such as non-uniform intensity distributions and varying tumor shapes by incorporating a gate unit and a Deformable Spatial Attention Module to enhance feature extraction and improve accuracy. Experimental results demonstrate that ACSNet outperforms existing methods, achieving state-of-the-art results in breast tumor segmentation and classification tasks.

Uploaded by

Santiago Moreno
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Computers in Biology and Medicine 173 (2024) 108319

Contents lists available at ScienceDirect

Computers in Biology and Medicine


journal homepage: [Link]/locate/compbiomed

Multi-task learning for segmentation and classification of breast tumors


from ultrasound images
Qiqi He a, b, Qiuju Yang a, *, Hang Su a, Yixuan Wang a
a
School of Physics and Information Technology, Shaanxi Normal University, Xi’an, China
b
School of Life Science and Technology, Xidian University, Xi’an, China

A R T I C L E I N F O A B S T R A C T

Keywords: Segmentation and classification of breast tumors are critical components of breast ultrasound (BUS) computer-
Multi-task learning aided diagnosis (CAD), which significantly improves the diagnostic accuracy of breast cancer. However, the
Breast tumor characteristics of tumor regions in BUS images, such as non-uniform intensity distributions, ambiguous or
Ultrasound images
missing boundaries, and varying tumor shapes and sizes, pose significant challenges to automated segmentation
Segmentation
Classification
and classification solutions. Many previous studies have proposed multi-task learning methods to jointly tackle
tumor segmentation and classification by sharing the features extracted by the encoder. Unfortunately, this often
introduces redundant or misleading information, which hinders effective feature exploitation and adversely
affects performance. To address this issue, we present ACSNet, a novel multi-task learning network designed to
optimize tumor segmentation and classification in BUS images. The segmentation network incorporates a novel
gate unit to allow optimal transfer of valuable contextual information from the encoder to the decoder. In
addition, we develop the Deformable Spatial Attention Module (DSAModule) to improve segmentation accuracy
by overcoming the limitations of conventional convolution in dealing with morphological variations of tumors.
In the classification branch, multi-scale feature extraction and channel attention mechanisms are integrated to
discriminate between benign and malignant breast tumors. Experiments on two publicly available BUS datasets
demonstrate that ACSNet not only outperforms mainstream multi-task learning methods for both breast tumor
segmentation and classification tasks, but also achieves state-of-the-art results for BUS tumor segmentation. Code
and models are available at [Link]

1. Introduction and shape, irregular and blurred tumor boundaries, and low
signal-to-noise ratios in ultrasound images, as shown in Fig. 1.
Female breast cancer has surpassed lung cancer as the most Convolutional Neural Networks (CNNs) have achieved remarkable
commonly diagnosed cancer worldwide, and ranked fifth as the leading success in medical image analysis due to their powerful capability for
cause of cancer deaths worldwide [1]. Ultrasound imaging has become automatic feature extraction [4–10]. Specifically, in medical image
an important tool for breast cancer diagnosis due to its versatility, safety, segmentation tasks, UNet [11] is preferred for its ability to reconstruct
and sensitivity. However, evaluating a whole breast ultrasound (BUS) high-resolution segmentation images by integrating multi-level infor­
examination slice by slice is highly time-consuming, even for experi­ mation. In BUS image segmentation, Almajalid et al. [12] enhanced the
enced radiologists [2], and there is a high inter-observer variation rate. performance of UNet in the precision of breast lesion segmentation by
Therefore, an effective computer-aided diagnosis (CAD) system is implementing strategies such as contrast enhancement and speckle noise
essential in assisting physicians with the diagnosis and treatment of reduction. Ning et al. [13] pointed out that a simple modeling frame­
breast cancer. Segmentation and classification of BUS tumors are highly work is difficult to obtain ideal segmentation results on BUS images due
relevant tasks and are two fundamental objectives of CAD systems, to the complexity of the echo patterns of ultrasound images and the
because both share generic image features, such as shape and boundary interference from surrounding tissues. Additionally, Zhao et al. [14]
features [3]. Nevertheless, designing this system for BUS images is pointed out issues with the encoder-decoder model architecture,
challenging due to the problems such as large variations in tumor size emphasizing that while high-quality segmentation relies on useful

* Corresponding author.
E-mail address: yangqiuju@[Link] (Q. Yang).

[Link]
Received 11 July 2023; Received in revised form 3 March 2024; Accepted 12 March 2024
Available online 18 March 2024
0010-4825/© 2024 Elsevier Ltd. All rights reserved.
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

encoder features, direct transfer through skip connections could intro­ analysis within a unified network framework. This approach promotes
duce misleading contextual information, resulting in underutilization of the sharing of breast tumor-specific features, effectively improving the
valuable features and potential introduction of irrelevant information performance of both tasks and emerging as a promising avenue for
leading to pixel classification errors. Moreover, the feature extraction further investigation. However, most existing multi-task learning ar­
capacity at a single level is limited, causing the global contextual in­ chitectures rely on an encoder-decoder approach and often overlook the
formation captured by the encoder at deeper levels to be progressively impact of redundant or misleading information within the encoder on
lost during the up-sampling process in the decoder [15]. the results. Such oversight can hinder the selective transfer of valuable
Similarly, the progress in deep learning has significantly advanced contextual information from the encoder to the decoder [14]. Addi­
tumor classification in BUS imaging. CNNs adopt a hierarchical struc­ tionally, tumor boundaries in BUS images are often poorly defined and
ture of multiple convolutional and down-sampling layers to capture there is significant variability in tumor size and shape. This variability
deep contextual features, an architecture capable of recognizing pat­ requires a network that can adapt to tumors of different sizes.
terns ranging from simple edges to complex high-order arrangements. In this paper, we present ACSNet, an end-to-end multi-task learning
This allows the model to learn advanced semantic features in images and network tailored for the simultaneous learning of BUS tumor segmen­
effectively identify diverse visual content. Notable studies in BUS tumor tation and classification tasks. The architecture consists of an encoder-
classification include Daoud et al.’s [16] investigation of deep feature decoder network dedicated to segmentation and a robust multi-scale
extraction and transfer learning techniques to enhance the classification feature aggregation network for benign and malignant classification.
accuracy. In addition, Saad et al. [17] introduced the BreastUS Trans­ To refine the information flow during skip connections, gate unit mod­
former, a Transformer model that incorporates self-attention mecha­ ules for segmentation are designed to control the feature extraction
nisms specifically designed for BUS image classification. process of the encoder. Furthermore, a deformable spatial attention
Both the breast tumor segmentation and classification tasks rely module (DSAModule) is introduced to capture the complex morpho­
heavily on the accurate identification and detailed interpretation of logical variations of breast tumors. In the classification task, a channel
common morphological features of tumors, including their shape, size, attention mechanism within the multi-scale feature extraction network
texture, and edge characteristics [3]. Proper delineation of tumor highlights salient features while mitigating the influence of extraneous
boundaries for segmentation and determination of their benign or ma­ information, thereby improving the accuracy of the network in dis­
lignant nature for classification are closely related goals [18]. Accurate tinguishing between benign and malignant breast tumors. In addition, to
tumor segmentation can provide key morphological features that are suppress noise and further improve the multi-task learning performance,
essential for accurate classification, while a sound understanding of the LossNet, a generic no-training loss network proposed by Zhao et al. [23],
classification will aid in more accurate tumor segmentation. This is used to supervise the feature mapping extracted from the bottom to
interdependence and information sharing between tasks lead to higher the top layer, providing crucial structural detail supervision to the
accuracy and efficiency in breast cancer diagnosis. feature layers. We evaluate the effectiveness of the proposed method
However, single-task CNNs designed for BUS imaging may face through extensive experiments on two publicly available datasets. Our
limitations in scalability and dependence on substantial amounts of main contributions can be summarized as follows:
annotated data when tackling specific segmentation or classification
tasks. In contrast, multi-task learning (MTL), a well-established learning (1) We present ACSNet, a novel multi-task learning model for
paradigm in machine learning, has demonstrated its ability to improve simultaneous segmentation and classification of BUS tumors,
the performance of individual tasks by sharing information between which effectively addresses issues related to tumor size and shape
related tasks [19]. The inductive bias introduced by MTL serves as a variability and irregular boundaries in BUS images.
regularization mechanism, effectively preventing overfitting and pro­ (2) For tumor segmentation, we are developing the DSAModule and
moting the acquisition of more robust feature representations. In addi­ gate units. The DSAModule effectively adapts to morphological
tion, MTL proves beneficial in mitigating the challenges associated with changes, while the gate unit optimizes the information flow be­
small datasets by facilitating knowledge transfer between different tasks tween the encoder and decoder, thereby enhancing the accuracy
and improving the overall generalizability of the model [20]. of segmentation.
For example, Shareef et al. [20] showed that multi-task learning (3) For classification, ACSNet integrates a channel attention mecha­
outperformed single-task learning methods in the context of breast ul­ nism and a multi-scale feature fusion method, achieving excellent
trasound classification. Recent research on MTL [2,21,22] advocates the performance in breast cancer classification.
integration of segmentation and classification tasks in BUS image

Fig. 1. Heavily shadowed areas resembling lesions in ultrasound images.

2
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

(4) Our experiments on two BUS tumor datasets show that ACSNet 2.3. Multi-task learning for breast ultrasound image segmentation and
outperforms current mainstream multi-task learning methods for classification
both segmentation and classification tasks. Furthermore, ACSNet
achieves state-of-the-art results for BUS tumor segmentation. Many multi-task learning (MTL) methods have been used to improve
the results of both tumor classification and segmentation in BUS images.
The rest of the paper is organized as follows. Section 2 provides an Zhou et al. [2] developed a multi-task learning network for 3D auto­
overview of breast ultrasound image segmentation, classification, and matic BUS images, utilizing a multi-scale feature connectivity network
multi-task learning for both tasks. Section 3 describes the proposed for tumor classification and an iterative feature refinement training
ACSNet in detail. Section 4 introduces the datasets used in our research strategy to focus on tumor regions. Xu et al. [46] proposed a
and provides experimental details. The experimental results are pre­ region-focused MTL framework (RMTL-Net) for simultaneous BUS
sented in Section 5. Further discussions are given in Section 6, and image segmentation and classification, using a region-attention (RA)
Section 7 concludes the study. module to automatically learn weighted category-sensitive information
in tumor, peri-tumor, and background regions. Wang et al. [47] pro­
2. Related work posed a multi-feature-guided CNN architecture for enhancement, seg­
mentation, and classification of bone surfaces in ultrasound images.
2.1. Breast ultrasound image segmentation Zhang et al. [21] proposed the use of an Attention Gate (AG) module to
integrate segmentation and classification tasks and improve the accu­
BUS tumor segmentation methods can be broadly categorized into racy by effectively using lesion region information. Multi-task learning
traditional methods and deep learning-based approaches. Traditional has also shown advantages in other medical image analysis tasks. For
methods include region-based methods [24], deformable models [25, example, Chen et al. [48] proposed a fully automated framework for
26], graph-based methods [27,28], and learning-based methods [29, gadolinium-enhanced magnetic resonance image (GE-MRI) left atrial
30]. These methods rely heavily on hand-designed features and texture segmentation based on deep learning, Qu et al. [49] proposed a new
analysis to detect and segment breast lesion regions in ultrasound im­ multi-task learning method for cell nucleus segmentation and classifi­
ages, which could lead to misidentification of breast lesions in complex cation within a unified framework.
environments [31].
In recent years, deep learning-based approaches have shown 3. Methodology
considerable progress in BUS medical image processing. Shareef et al.
[32] introduced row-column convolutional kernels that adapt to the Fig. 2 shows the proposed multi-task learning network, ACSNet,
anatomical structure of the breast and fused contextual information at which integrates the BUS image segmentation and classification into a
different scales to segment small breast tumors. Hu et al. [33] proposed unified end-to-end model. ACSNet takes a BUS image as input and
to combine an expanded convolutional network with a phase-based produces two outputs: a tumor segmentation map and an image-level
active contour model for automatic breast lesion region segmentation. classification probability.
Yap et al. [34] comparatively studied patched-based LeNet [35], UNet More specifically, the architecture of ACSNet consists of two main
[11], and migration learning methods for breast ultrasound lesion components: a segmentation network and a classification branch. The
detection. Zhu et al. [36] developed a second-order sub-region pooling core of the segmentation network is a U-shaped encoder-decoder
network for breast lesion segmentation by using second-order statistics structure that includes stages 1–5 for the encoder and stages 6–9 for
for multiple feature sub-regions. Vakanski et al. [37] proposed to inte­ the decoder. The final output stage of the decoder, stage 9, effectively
grate visual saliency into a deep learning model for breast tumor seg­ achieves the segmentation of tumor lesions in ultrasound images. At the
mentation in ultrasound images. He et al. [38] designed a cross same time, the classification branch utilizes features from stages 4, 5,
CNN-Transformer network, named HCTNet, for breast ultrasound lesion and 6 to classify breast tumors in these images as either benign or
segmentation. malignant.
To optimize the flow of information, gate units have been introduced
2.2. Breast ultrasound image classification at each skip connection between the encoder and decoder for the seg­
mentation task. These gate units are designed to extract more robust
For BUS image classification, deep learning-based methods have features from the lesion areas during segmentation, thereby making
shown superior feature extraction capabilities when compared to more efficient use of the encoder features and generating more accurate
traditional machine learning methods. Qi et al. [39] proposed a deep predictive segmentation maps. In addition, the DSAModule is developed
CNN with multi-scale kernels and skip connections for diagnosing breast to enhance the spatial feature information of the lesions at different
ultrasonography images. Byra et al. [40] developed a Selective Kernel stages of the decoder, taking into account the variation in tumor size and
(SK) UNet, which utilized an attention mechanism to adjust the recep­ the interference caused by background noise.
tive field of the network and fused the feature maps extracted by dilation For the classification task, multi-scale features are extracted by
and traditional convolution. Qian et al. [41] proposed a multi-pathway merging feature maps from different scales, and a channel attention
deep learning architecture for automated breast cancer risk prediction, mechanism [50] is employed to enhance the learning capability of the
using multimodal and multi-view BUS images to mimic routine clinical network for feature representation. By jointly handling the segmenta­
working processes and improve clinical applicability. Cui et al. [42] tion and classification tasks and utilizing common features, MTL en­
constructed a CNN that adaptively fuses information from multiple hances the model’s ability to analyze and process breast tumors.
tumor regions for breast tumor classification in the testing process using Furthermore, this approach improves data efficiency and reduces the
images without segmentation masks. Zhuang et al. [43] utilized transfer dependence on large, task-specific datasets.
learning [44] methods with the original BUS images, as well as
enhanced and bilaterally filtered images, to improve breast lesion clas­ 3.1. Multi-task learning network
sification. Mo et al. [45] proposed a HoVer-Trans model that associates
Transformers with CNNs for breast cancer diagnosis in BUS images. The proposed ACSNet utilizes UNet [11] as the backbone architec­
ture for multi-task learning due to its excellent performance in medical
image segmentation. UNet consists of encoding paths, decoding paths,
and skip connections between them. Skipping connections are essential
mechanisms for disseminating spatial information, bridging semantic

3
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

Fig. 2. Overview of the proposed multi-task learning network. The architecture consists of a segmentation network and a classification branch. The segmentation
network includes gate units to regulate the features extracted from different encoder levels, while the DSAModules enhance the network’s adaptability to tumor
deformations. The input to the classification branch consists of tumor features at different scales, which are aggregated using the channel attention mechanism.

gaps, and refining segmentation results. These connections greatly connection in ACSNet, allowing for precise control and optimization of
enhance the model’s ability to capture and utilize global and local in­ the feature information flowing from the encoder. In contrast to the
formation, making UNet a powerful architecture for medical image approach of Zhao et al. [14], which evaluates the importance of different
segmentation tasks. Given UNet’s excellent performance in segmenta­ stages by combining encoder and decoder features, our gate unit em­
tion tasks, strong feature fusion capability, adaptability to different ploys a distinctive combination of Global Average Pooling (GAP) and
shapes, effectiveness on limited data, and previous success in related Global Max Pooling (GMP). This approach enhances the selectivity of
works [22,40], we make it the backbone of ACSNet. As shown in the the feature control process, facilitating effective identification and
segmentation network in Fig. 2, UNet contains convolutional layers at retention of key features. Consequently, our method significantly im­
each stage. The encoding path employs four down-sampling operations proves the overall multi-task learning performance.
to extract high-level semantic features, while the decoding path uses Fig. 3 shows the designed gate unit. The gate unit allows the infor­
four upsampling operations to restore the feature maps to the original mation selection of the input features and improves the effectiveness of
input size; in addition, the feature maps obtained from the encoder are the output features. The gate unit takes the convolutional feature map
connected to the decoder using skip connections to propagate spatial X ∈ RC×H×W as input, where H and W are the spatial height and width,
information and refine the segmentation results. Each convolution and C is the number of channels. With the convolution operation, the
operation is followed by group normalization (GN) and ReLU. input X is first divided into two feature maps F1 and F2 , (F1 ,
Image classification networks mostly use the high-level feature maps F2 ∈ R1×H×W ), which allows the gate unit to pay attention to effective
of CNNs such as VGG [51] and ResNet [52]. Inspired by this, we add a information from different representation subspaces, and then F1 and F2
classification branch at the bottom of UNet. First, the feature maps from are normalized to the interval [0,1] using the Sigmoid function to obtain
stages 4, 5, and 6 of UNet are passed into the classification network. F1′ and F2′. A pair of gate values G1 and G2 are computed by using the
Then, the shared feature maps from these stages are fused using the
global average pooling and global maximum pooling operations on F1′
channel attention (CA) mechanism [50] to explicitly model in­
terdependencies between channels and adaptively recalibrate the and F2′. Finally, G1 and G2 are used to select the contextual information
feature responses of the channels. Finally, the fused features are input from the input features respectively, and to sum and fuse the informa­
into two fully connected (FC) layers and a Softmax layer to predict the tion to obtain the output feature Y ∈ RC×H×W . In this way, the gate unit
benign and malignant categories of the input images. can selectively send the features from the encoder to the decoder during
the skip connection process. The whole calculation process for the gate
unit can be expressed as follows:
3.2. Gate control mechanism
⎛ ⎞
To overcome the problem of redundant or misleading information in ⎜
G1 = GAP⎝Sigmoid(Conv(X))⎠,

(1)
the encoder during multi-task learning, several solutions have been ⏟̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏞⏞̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏟
F′
proposed by researchers. Among them, the application of gating mech­ 1

anisms plays a crucial role, not only by improving the quality of feature ⎛ ⎞
extraction, but also by enhancing the adaptability of the models. For ⎜ ⎟
example, the Gated Dual-branch Network (GateNet) [14] was designed G2 = GMP⎝Sigmoid(Conv(X))⎠, (2)
⏟̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏞⏞̅̅̅̅̅̅̅̅̅̅̅̅̅̅̅⏟
to optimize the transfer of contextual information from the encoder to F′2

the decoder. Inspired by this, we have integrated gate units at each skip

4
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

Fig. 3. Overview of the proposed gate unit. The input feature map is directed to two branches with spatial attention mechanisms. These branches utilize GAP and
GMP, respectively, to extract relevant feature details. After summing the feature maps from both branches, a refined output is generated via a 3 × 3 Conv.

( ) we add a spatial attention module after the deformable convolution. In


Y = Conv X × G1 + X × G2 , (3)
the spatial attention module, we use average pooling and maximum
where GAP denotes global average pooling operation and GMP denotes pooling operations for X′ along the channel dimension C to obtain in­
global maximum pooling operation. formation about the spatial context of the object and the important in­
formation in the features, respectively [54]. Then the 3 × 3
convolutional layers and Sigmoid function are used to generate the
3.3. Deformable spatial attention module corresponding spatial attention map S1,2 ∈ R1×H×W respectively, and the
attention-weighted feature maps A1,2 ∈ RC×H×W are obtained by multi­
In the proposed ACSNet, we further improve the segmentation per­
plying S1,2 ∈ R1×H×W with the deformable convolved features X′,
formance of the model by addressing the problem of under- and over-
respectively. Finally, A1 and A2 are concatenated and the lesion features
segmentation due to the variation in lesion size and different morpho­
are again enhanced by a deformable convolution to obtain the final
logical characteristics of the tumor. We design a deformable spatial
attention module (DSAModule) that overcomes the limitations of con­ output Y ∈ RC×H×W . The whole DSAModule process can be expressed as
ventional convolution by applying deformable convolution [53] to the follows:
significant changes of the lesion region accordingly. Specifically, while X′ = DConv(X), (4)
traditional convolutional operations in image processing rely on a fixed
and regular grid, deformable convolution introduces learnable offsets S1 = Sigmoid(Conv(Avgaxis=channel (X′))), (5)
that allow the convolutional kernel sampling points to be dynamically
adjusted based on image features. This design provides greater flexi­ S2 = Sigmoid(Conv(Maxaxis=channel (X′))), (6)
bility to adapt to local variations in the image. We use deformable
convolution to address the limitations of traditional convolution when ⎛ ⎛ ⎞⎞
dealing with the diverse morphologies of tumors. In addition, we inte­ Y = DConv⎝Cat⎝X′ × S1 , X′ × S2 ⎠⎠, (7)
⏟̅̅̅⏞⏞̅̅̅⏟ ⏟̅̅̅⏞⏞̅̅̅⏟
grate a spatial attention mechanism into the DSAModule that merges A1 A2
max-pooling and average-pooling operations. This integration effec­
tively extracts and consolidates critical information from the feature where DConv is the deformable convolution.
maps, resulting in the generation of the final feature map. This inclusion
of the spatial attention mechanism significantly improves the model’s
ability to identify and focus on critical regions within the images. 3.4. Multi-task loss function
Fig. 4 shows the proposed DSAModule. DSAModule takes the feature
map X ∈ RC×H×W of the current stage of the encoder as input, and first In the classification task, we use the binary cross-entropy [55] loss
obtain X′ ∈ RC×H×W by 3 × 3 deformable convolution. To further function to train the model to classify the whole image instead of each
enhance the spatial details and improve the segmentation performance, pixel, which can be expressed as:

Fig. 4. Overview of the proposed DSAModule. The input feature maps first pass through a deformable convolution to preliminarily extract the features of the tumor
regions. These features are then fed into a spatial attention module to refine relevant feature details, and finally processed through another 3 × 3 deformable
convolution to obtain enhanced output features.

5
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

1 ∑ where Ltotal is denoted as multi-task loss and λ ∈ [0, 1] is the classification


N
( ) ( k ( ) ( ))
Lcls pk , gk = − g log pk + 1 − gk log 1 − pk , (8)
N k=1 task weight. In our experiment, segmentation and classification tasks are
weighted differently, with more weights given to the more difficult
where gk is the ground truth class of sample k and pk is the predicted segmentation tasks.
classification probability from the proposed network for sample k,
respectively. 4. Data and experiments
In the segmentation task, we use a segmentation loss based on the
Dice coefficient to address the issue of class imbalance between the 4.1. Dataset
foreground (object) and background in the image, which can cause
segmentation bias. The Dice coefficient measures the similarity between We evaluated the effectiveness of the proposed ACSNet using two
two sets of data, and is commonly used as an evaluation metric in seg­ publicly available breast ultrasound image datasets. The first dataset,
mentation tasks. The segmentation loss is defined as: BUSI [56], was collected at Baheya Women’s Cancer Early Detection and
( ) 2Pseg Yseg + 1 Treatment Hospital in Cairo, Egypt in 2018, utilizing the LOGIQ E9 ul­
Ldice Pseg , Yseg = 1 − , (9) trasound and LOGIQ E9 Agile ultrasound systems. The BUSI dataset
Pseg + Yseg + 1
consists of 780 breast cancer ultrasound images from 600 female pa­
where Ldice is the dice segmentation loss, Pseg and Yseg denote the pre­ tients aged 25–75 years with an average image size of 500 × 500 pixels.
dicted segmentation map from the proposed network and the ground It includes 437 benign cases, 210 malignant masses, and 133 normal
truth respectively. cases. The selection of the BUSI dataset was motivated by its inherent
Additionally, inspired by the polyp segmentation described in Zhao challenges, including variations in tumor size, irregular boundaries, and
et al. [23], we further employed their LossNet to refine the segmentation significant morphological differences.
results, ensuring comprehensive supervision of both details and struc­ The second dataset, BUS [34], was collected by the UDIAT Diag­
tures at the feature level. To achieve this, we extract the predicted and nostic Centre of Parc Taulí in Sabadell, using a Siemens ACUSON
ground truth multi-scale features using an ImageNet pre-trained classi­ Sequoia C512 system and a 17L5HD linear array transducer (8.5 MHz)
fication network, namely VGG-16 [51]. Then, the feature difference is to acquire ultrasound images. The BUS dataset consists of 163 breast
computed as a loss LLossNet : cancer ultrasound images from different women, with an average image
size of 760 × 570 pixels, including 53 images of malignant cases and 110
LLossNet = l1LossNet + l2LossNet + l3LossNet + l4LossNet , (10) images of benign cases. Since tumors tend to be smaller in the early
⃦ ⃦ stages of clinical detection, accurate segmentation or classification be­
liLossNet = ⃦FPi − FGi ⃦2 , i = 1, 2, 3, 4 (11) comes a challenge. Therefore, the BUS dataset deliberately focuses on
smaller tumors, which pose a significant challenge due to their limited
where FPi and FGi separately represent the i-th level feature maps size and the possibility of superimposed benign or malignant features,
extracted from the prediction and ground truth, and liLossNet is calculated complicating the segmentation and classification process.
as their Euclidean distance (L2-Loss). Fig. 5 illustrates the calculation of As our multi-task learning aims to segment and classify benign and
LossNet. malignant lesions in breast ultrasound images, we removed normal cases
Thus, the total segmentation loss function can be expressed as: without breast lesions from the BUSI dataset. In addition, a five-fold
cross-validation method was used to evaluate the performance of
Lseg (Ldice , LLossNet ) = Ldice + γ × LLossNet , (12)
different methods on the two aforementioned datasets.
where γ ∈ [0, 1] is the LLossNet weight in the total segmentation loss
function. 4.2. Experiment details
In our approach, the classification loss Lcls and the segmentation loss
Lseg are linearly combined into a multi-task loss by a hyperparameter λ. All experiments were conducted on an Ubuntu 16.04 system equip­
The multi-task loss is defined as: ped with an Intel(R) Xeon(R) CPU E5-2680 V3. All models were
implemented using the deep learning toolbox PyTorch [57] and were
Ltotal = λLcls + (1 − λ)Lseg , (13) trained and tested on an NVIDIA GeForce GTX 1080Ti with 11,264 MB
of memory. Using the Adam optimizer with an initialized learning rate

Fig. 5. Illustration of the calculation of LossNet.

6
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

of 1e-4, a momentum β1 of 0.9, a momentum β2 of 0.99, and a weight Table 1


decay of 0.0001. We employed LambdaLR to adjust the learning rate Networks constructed in ablation study. Cls: classification; Seg: segmentation.
during training. The hyperparameter γ was set to 0.2 and λ was set to 0.3. Model Task CA Gate unit DSAModule Multi-scale
For the BUSI dataset, we set the GPU batch size to 16 and trained the
UNet Seg ⨯ ⨯ ⨯ ⨯
model for 100 epochs, whereas for the BUS dataset, we set the GPU ClsNet Cls ⨯ ⨯ ⨯ ⨯
batch size to 4 and trained the model for 90 epochs. All images were CUNet Seg + Cls ⨯ ⨯ ⨯ ⨯
resized to 256 × 256. We employed standard data augmentation tech­ CUGNet Seg + Cls ⨯ ✓ ⨯ ⨯
niques including random rotation, horizontal flip, and vertical flip. CAUNet Seg + Cls ✓ ⨯ ⨯ ⨯
CUANet Seg + Cls
Rotations were applied randomly within a range of 10–30◦ range to
⨯ ⨯ ✓ ⨯
CUGANet Seg + Cls ⨯ ✓ ✓ ⨯
introduce variability in tumor orientation. Horizontal and vertical flips CMSUNet Seg + Cls ⨯ ⨯ ⨯ ✓
were used to simulate different tumor appearances from different ori­ CAMSUNet Seg + Cls ✓ ⨯ ⨯ ✓
entations and sides of the image. These well-established augmentation ACSNet Seg + Cls ✓ ✓ ✓ ✓
methods are commonly used in deep learning-based medical image
analysis. By artificially expanding the variety and size of the dataset,
and UNet are used as baseline models for classification and segmenta­
they play a key role in enhancing the generalization ability of the model,
tion, respectively. ClsNet uses the encoding path of UNet followed by a
a critical aspect for robust performance in various clinical scenarios.
GAP layer and two FC layers for classification. In the multi-task models,
the baseline is denoted as CUNet. CUNet uses UNet as the framework for
4.3. Evaluation metrics the segmentation task and adds a single-scale classification branch to the
UNet model, which uses the features of stage 5 of UNet as the classifi­
To quantitatively evaluate the effectiveness of tumor segmentation, cation features. We then add the CA module to the single-scale classi­
we utilized four commonly used evaluation metrics: Dice similarity co­ fication branch to obtain CAUNet. Based on CUNet, by using the gate
efficient (Dice) [58], Jaccard index (Jaccard) [31], 95% asymmetric unit to learn the skip connections, we obtain CUGNet. Based on CUNet,
Hausdorff distance (95HD) [59], and average surface distance (ASD) we add DSAModule to enhance the spatial location information of le­
[60]. These indices were calculated as follows: sions, resulting in CUANet, and use both gate unit and DSAModule to
2 × |A ∩ B| obtain CUGANet. In addition, we use a multi-scale classification branch
Dice(A, B) = , (14) for multi-task learning instead of a single-scale classification branch
|A| + |B|
(denoted as CMSUNet) and add the proposed CA module to the classifi­
|A ∩ B| cation branch for multi-task learning (denoted as CAMSUNet). Finally, we
Jaccard(A, B) = , (15) integrate all proposed components to obtain the final model ACSNet.
|A ∪ B|
The quantitative results of the ablation experiments are presented in
{ }
Table 2. These results show that the integration of all principal com­
95HD(A, B) = max supinf ‖x − y‖, supinf ‖x − y‖ , (16)
x⊂A y⊂B y⊂B x⊂A ponents in our model leads to the best performance compared to all
other configurations. Specifically, for the segmentation tasks, the pro­

miny∈B d(x, y) posed ACSNet achieved remarkable results, with the highest Dice score
ASD(A, B) = x∈A
, (17)
|A| of 84.90%, the highest Jaccard index of 78.62%, the lowest 95HD of
13.04 mm, and the second-lowest ASD of 3.45 mm among all configu­
where A and B are the segmentation results and the ground truth, and x rations. ACSNet also performed outstandingly in the classification tasks,
and y are the pixels in A and B, respectively. Dice and Jaccard are sen­ achieving the highest ACC of 94.44%, the highest PRE of 94.61% and a
sitive to tumor size and 95HD sensitive to tumor shape. remarkably high REC of 93.86%. These results highlight the critical
For the tumor classification task, we used recall (REC), precision importance of each component, including the channel attention mech­
(PRE), and accuracy (ACC) for the quantitative evaluation: anism, gate unit, DSAModule, and multi-scale feature fusion, in
improving the accuracy of both tumor segmentation and classification in
TP + TN
ACC(%) = × 100, (18) breast ultrasound images.
TP + TN + FN + FP
Furthermore, a comparison between CUGNet and CUNet shows that
TP the gate unit optimizes the features extracted from the encoder to a
REC(%) = × 100, (19) certain extent. This optimization helps to avoid interference from
TP + FN
redundant or misleading information, thereby enhancing the perfor­
PRE(%) =
TP
× 100, (20) mance of the model in breast tumor segmentation and classification.
TP + FP Additionally, when comparing CUGANet with CUGNet, it can be seen that
the DSAModule effectively improves the model’s ability to detect de­
where TP, FP, TN, FN are the number of true positive, false positive, true formations in breast tumors. The arrows accompanying each indicator
negative and false negative, respectively. See Ref. [58] for details of indicate the desired trend: an upward arrow (↑) indicates a preference
these classification metrics. for higher values of the indicator, while a downward arrow (↓) indicates
a desire for lower values to achieve optimal performance.
5. Results When comparing the results of CUNet and UNet, it is clear that
CUNet shows superior performance in the segmentation task. This
5.1. Ablation study improvement clearly confirms the benefit of integrating classification
information into segmentation, and underscores the advantages of
In this section, we conducted ablation experiments to evaluate the multi-task learning over conventional single-task approaches. Further­
effectiveness of the principal components of our network, including CA, more, the inclusion of a gate unit in the skip connections of the CUNet
gate unit, DSAModule, and multi-scale feature fusion. We performed model, as seen in CUGNet, leads to improved results across all metrics for
these experiments on the BUSI dataset since it contains the most chal­ both segmentation and classification. This enhancement signifies the
lenging data with significant lesion morphological variation and the effectiveness of the gate unit in selectively filtering the features
largest number of samples in both datasets. extracted by the encoder, ensuring that only the most relevant features
Table 1 presents the comparison results of our method with different are passed to the decoder. Additionally, the significant improvements
combinations of components. Specifically, the single-task models ClsNet

7
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

Table 2
Quantitative results on BUSI for all networks constructed in ablation study. The optimal results are shown in bold.
Model Segmentation (Mean ± Std) Classification (Mean ± Std)

Dice (%) ↑ Jaccard (%) ↑ 95HD ↓ ASD ↓ ACC (%) ↑ PRE (%) ↑ REC (%) ↑

UNet 83.04 ± 1.64 76.34 ± 2.04 15.95 ± 2.29 5.14 ± 0.67 – – –


ClsNet – – – – 92.38 ± 2.59 92.54 ± 2.29 91.38 ± 2.30
CUNet 83.56 ± 1.72 76.36 ± 2.19 15.07 ± 2.47 4.01 ± 1.13 92.54 ± 2.87 92.92 ± 2.39 92.20 ± 2.04
CUGNet 84.26 ± 1.14 77.58 ± 1.36 14.04 ± 1.15 3.75 ± 0.61 93.02 ± 2.21 93.20 ± 2.13 91.70 ± 1.62
CAUNet 83.44 ± 1.28 76.98 ± 1.17 14.33 ± 1.39 3.95 ± 0.93 93.81 ± 1.46 93.89 ± 1.40 92.89 ± 1.25
CUANet 84.48 ± 1.26 77.78 ± 0.94 14.03 ± 1.75 4.22 ± 1.33 93.33 ± 2.91 93.61 ± 2.43 92.82 ± 2.12
CUGANet 84.76 ± 0.83 78.20 ± 0.94 13.12 ± 1.08 3.33 ± 0.68 93.17 ± 2.91 93.35 ± 2.68 92.03 ± 2.74
CMSUNet 83.76 ± 1.90 77.12 ± 1.99 14.84 ± 2.27 4.26 ± 1.29 94.16 ± 2.77 94.35 ± 2.11 93.76 ± 1.67
CAMSUNet 84.02 ± 1.18 77.12 ± 1.29 14.63 ± 1.68 4.45 ± 0.95 94.29 ± 2.10 94.46 ± 2.07 94.07 ± 1.48
ACSNet 84.90 ± 1.69 78.62 ± 1.75 13.04 ± 2.58 3.45 ± 1.59 94.44 ± 2.07 94.61 ± 1.56 93.86 ± 2.33

observed in CUANet and CAUNet, as compared to CUNet, are due to their 5.2. Comparison with the state-of-the-arts
respective attentional mechanisms, which contribute to performance
enhancements by focusing on critical features, thereby refining the Table 3 compares the proposed ACSNet with ten advanced methods,
overall model accuracy and effectiveness. Furthermore, CMSUNet, unlike including six recent single-task segmentation methods (FCN [61], UNet
CUNet, employs multi-scale feature extraction for classification, result­ [11], DeeplabV3+ [62], FPN [63], Unet++ [64], and TransUnet [65])
ing in superior performance over both ClsNet and CUNet in classification and four state-of-the-art MTL methods (CUNet, Wang et al. [47], Chen
tasks. This result highlights the effectiveness of multi-scale feature et al. [48], and Qu et al. [49]) on the BUSI dataset. CUNet is the model
extraction and provides a compelling argument for its superiority over discussed in ablation study. For a fair comparison, we used the same
single-scale approaches. CMSUNet’s ability to capture a wider range of backbone (UNet) as the proposed method in all compared methods, and
feature details allows for more accurate classification, further validating applied five-fold cross-validation to all models.
the model’s robustness and versatility. Table 4 reports the mean and standard deviation values of the
Fig. 6 visually compares the segmentation results of various methods
in the ablation experiment. This paper designs a gate unit and a DSA­
Table 3
Module for tumor segmentation. From Fig. 6, we can conclude that the
Briefly comparing the proposed ACSNet with 10 advanced methods.
segmentation performance of CUGNet has greatly improved compared to
Methods Backbone Tasks Feature
CUNet, because its gate unit can select and control the feature infor­
Enhancement
mation in the encoder; Also, the segmentation accuracy of CUANet is
superior to CUNet, and over and under-segmentation problems are Single Multi- Attention
Segmentation Task Mechanism
significantly reduced, indicating that the developed DSAModule can
learn the location information of lesions in the feature map. CUGANet UNet [11] ResNet-18 ✓
FCN [61] ResNet-18
combines the strengths of the gate unit and DSAModule, further ✓
DeepLabV3+ ResNet-18 ✓
improving the segmentation performance of CUGNet and CUANet, [62]
resulting in more accurate lesion location and improved boundaries. In FPN [63] ResNet-18 ✓
addition, CUGANet achieves significantly better segmentation results Unet++ [64] ResNet-18 ✓
than CAUNet, CMSUNet, and CAMSUNet, which do not contain the gate TransUnet [65] R50+ViT- ✓
B_16
unit and DSAModule. All the results demonstrate that the effectiveness CUNet ResNet-18 ✓
of the proposed gate unit and DSAModule in breast ultrasound image Wang et al. ResNet-18 ✓ ✓
segmentation. [47]
Qu et al. [49] ResNet-18 ✓ ✓
Chen et al. [48] ResNet-18 ✓ ✓
ACSNet ResNet-18 ✓ ✓

Fig. 6. Visual segmentation results of ablation study.

8
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

Table 4 encoding the tokenized image patches from the feature maps of CNNs as
Classification performance (Mean ± SD) of all compared methods on BUSI. the input sequence for extracting the global semantics, while the
Methods ACC (%) ↑ PRE (%) ↑ REC (%) ↑ decoder upsamples the encoded features and then combines them with
the high-resolution CNN feature maps for precise localization. Deep­
CUNet 92.54 ± 2.87 92.92 ± 2.39 92.20 ± 2.04
Wang et al. [47] 92.86 ± 2.84 92.93 ± 2.93 92.13 ± 3.61 labV3+ obtained the second-best overall segmentation performance,
Qu et al. [49] 94.13 ± 1.98 94.51 ± 1.48 93.58 ± 1.36 followed by FPN. Among the five MTL methods, the proposed ACSNet
Chen et al. [48] 92.22 ± 1.91 92.32 ± 1.88 92.25 ± 1.89 achieved the best segmentation performance on all four metrics, with
ACSNet 94.44 ± 2.07 94.61 ± 1.56 93.86 ± 2.33 the lowest 95HD and ASD of 1.65 mm and 0.66 mm, respectively, while
Dice and Jaccard were 1.08% and 1.50% higher than the sub-optimal
classification metrics (ACC, PRE, and REC) among ACSNet and all the method (i.e., TransUnet), respectively.
competitors on BUSI. Among the five MTL methods, the proposed In addition, Table 5 also provides a clear comparison of floating-
ACSNet achieves the best performance in all metrics. Low-level features point operations (FLOPs) and the number of parameters (in millions)
are usually used for classification because they mainly capture shape for each model. ACSNet has a computational complexity of 30.96 G
and boundary information, while high-level features extract semantic FLOPs and 42.99 M parameters. Although ACSNet’s parameters are the
features of the targets. Capturing the features of small objects is chal­ largest among these models, its model complexity is low, and in terms of
lenging when the network depth increases with more convolution and results, ACSNet achieves optimal results on the task of breast tumor
down-sampling operations. To address this issue, ACSNet combines and segmentation and classification from ultrasound images, which we
fuses the multi-scale features for the classification task. Specifically, believe strikes a favorable balance between performance and efficiency.
ACSNet inputs the feature maps from stage 4 to stage 6 of the segmen­ The segmentation results of ACSNet and all competitors on BUSI are
tation network to the classification branch, which enables the multi- given in Fig. 7. Among the single-task segmentation methods (Fig. 7(d)–
scale features to contain more useful information through channel (i)), TransUnet has the best overall segmentation performance due to the
attention CA, thus achieving the best classification performance on the advantage of using the Transformer to obtain global information, while
BUSI. the CNN models (Fig. 7(d)–(h)) tend to ignore breast lesion details or
Table 5 summarizes the segmentation results of ACSNet and ten misclassify lesion regions as non-breast lesions in the predicted seg­
compared methods on Dice, Jaccard, 95HD, and ASD on the BUSI, mentation maps. The multi-task models shown in Fig. 7(j)–(m) have
including six single-task segmentation methods (i.e., UNet, FCN, FPN, poor results for tumor boundary segmentation and suffer from more
DeepLabV3+, UNet++ and TransUnet) and five MTL methods (i.e., serious over- or under-segmentation problems, which are due to the fact
CUNet, Wang et al. [47], Chen et al. [48], Qu et al. [49], and ACSNet). that the tumor morphology is highly variable and the features extracted
Among the six single-task segmentation methods, TransUnet obtained by the encoder are not fully suitable for the segmentation task. The
the highest Jaccard and Dice values, the lowest 95HD and ASD values, proposed ACSNet achieved the best segmentation performance among
and the best overall segmentation performance on BUSI, because it the five MTL methods and provides more accurate segmentation of the
combines the advantages of Transformer and UNet, with Transformer breast lesion region, with fewer over- or under-segmentation errors.
Table 6 summarizes the classification results of ACSNet and four
compared methods on BUS. It is clear that ACSNet has the highest ac­
Table 5 curacy, precision, and recall and improves over the second-best method
Segmentation performance of all compared methods on BUSI. (i.e., Qu et al. [49]) by 0.63%, 0.13%, and 3.4%, respectively.
Methods Dice Jaccard 95HD ↓ ASD ↓ FLOPs Param Table 7 reports the mean and standard deviation values of the four
(%) ↑ (%) ↑ (G) (M) segmentation metrics among our ACSNet and ten compared methods on
UNet [11] 83.04 76.34 ± 15.95 4.14 25.69 29.38 the dataset BUS. Among the single-task segmentation methods (UNet,
± 1.64 2.04 ± 2.29 ± FCN, FPN, DeepLabV3+, UNet++, and TransUnet), TransUnet has the
0.67
best overall segmentation performance, with the highest values on Dice
FCN [61] 81.46 73.90 ± 16.14 4.50 9.54 22.35
± 1.50 1.66 ± 1.69 ± and Jaccard and the lowest values on 95HD. Among the five MTL
0.36 methods (CUNet, Wang et al. [47], Chen et al. [48], Qu et al. [49], and
FPN [63] 82.68 75.88 ± 15.98 4.26 14.41 28.91 ACSNet), ACSNet achieved the best segmentation performance on all
± 0.69 0.72 ± 1.27 ± four metrics, with increased Dice and Jaccard values over the
0.68
DeepLabV3+ 83.34 76.30 ± 14.87 4.26 15.08 27.78
second-best method (TransUnet) by 1.26% and 1.66%, respectively,
[62] ± 0.96 0.88 ± 1.18 ± while reducing 95HD and ASD by 0.63 mm and 0.69 mm, respectively.
0.64 Fig. 8 shows the visual segmentation results of different segmenta­
UNet++ [64] 82.80 76.04 ± 14.83 4.18 35.20 16.80 tion methods on the dataset BUS. The proposed ACSNet effectively
± 0.87 1.03 ± 0.67 ±
mitigated the influence of tumor size, surrounding tissue, and shadow
1.26
TransUnet 83.82 77.12 ± 14.69 4.11 32.23 93.23 regions, providing segmentation results closer to the ground truth and
[65] ± 0.89 0.84 ± 1.68 ± fewer missed and false detections. The comprehensive evaluation results
1.09 and visual effects demonstrate the effectiveness of the proposed ACSNet
CUNet 83.56 76.36 ± 15.22 4.01 25.69 29.12 for breast ultrasound image segmentation and classification.
2.19 ±
± 1.72 ± 1.59
1.13
Wang et al. 82.76 76.30 ± 14.39 4.45 37.86 29.38 6. Discussion
[47] 2.03 ± 2.79
± 1.62 ±
1.76 BUS tumor segmentation and classification face significant chal­
Chen et al. 82.38 75.48 ± 16.51 5.20 27.90 34.72 lenges, such as variations in tumor size and shape, irregular and blurred
[48]
± 0.82
0.47 ± 0.67 ± tumor boundaries, and low signal-to-noise ratio of ultrasound images
1.52
[2]. To address these challenges, we propose ACSNet, a novel multi-task
Qu et al. [49] 82.28 75.48 ± 17.71 5.41 5.82 14.55
± 0.99 1.03 ± 2.12 ± learning model specifically designed for tumor segmentation and clas­
1.45 sification in BUS images. ACSNet starts by extracting basic patterns and
ACSNet 84.90 78.62 ± 13.04 3.45 30.96 42.99 features at shallow levels and gradually progresses to complex,
± 1.69 1.75 ± 2.58 ± high-level semantic feature extraction. As the model’s receptive field
1.59
expands, it not only performs more accurate tumor feature analysis, but

9
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

Fig. 7. Visualization of segmentation results of different methods on the dataset BUSI, (a)–(m) from left to right are (a) input image, (b) ground truth, and (c) the
segmentation results of ACSNet (Ours), (d) UNet, (e) Unet++, (f) FPN, (g) FCN, (h) DeeplabV3+, (i) TransUnet, (j) CUNet, (k) Wang et al. [47], (l) Chen et al. [48],
and (m) Qu et al. [49] respectively.

also effectively ignores background information in non-tumor regions,


Table 6
utilizing contextual information to enhance accuracy.
Classification performance (Mean ± SD) of all compared methods on BUS.
We have systematically evaluated the effectiveness of ACSNet’s
Methods ACC (%) ↑ PRE (%) ↑ REC (%) ↑ components and parameters through ablation studies. Experimental
CUNet 87.08 ± 3.70 87.00 ± 3.75 83.91 ± 4.54 results show that ACSNet achieves the best performance in breast cancer
Chen et al. [48] 84.02 ± 3.69 84.20 ± 3.56 80.38 ± 4.52 classification experiments, highlighting the critical role of the integra­
Qu et al. [49] 88.33 ± 4.16 89.16 ± 2.89 84.12 ± 6.99
tion of the channel attention mechanism and the multi-scale feature
Wang et al. [47] 85.89 ± 2.42 86.68 ± 1.97 83.22 ± 3.27
ACSNet 88.96 ± 1.49 89.29 ± 1.46 87.52 ± 2.66 fusion network in achieving superior classification results. ACSNet also
outperforms other methods in the segmentation task. The gate units
optimize the encoder’s feature information, enhancing the model’s
ability to accurately identify and exploit the contextual information. At
Table 7
the same time, the DSAModule adapts to the morphological variability
Segmentation performance (Mean ± SD) of all compared methods on BUS.
of tumors, further improving the accuracy of the model in segmenting
Methods Dice (%) ↑ Jaccard (%) 95HD ↓ ASD ↓
breast tumors.

For the multitask learning of breast tumor segmentation and classi­
UNet [11] 87.36 ± 1.99 81.28 ± 2.41 8.77 ± 1.12 2.56 ± 0.90 fication from ultrasound images, the deep learning architecture is of
FCN [61] 84.36 ± 3.31 76.62 ± 4.21 8.57 ± 4.30 2.31 ± 2.21
FPN [63] 83.00 ± 3.73 75.52 ± 3.36 14.28 ± 4.97 ± 1.61
paramount importance. Additionally, we utilize data augmentation
1.87 techniques such as random rotation and flipping to mitigate the limi­
DeepLabV3+ 86.78 ± 3.70 80.34 ± 4.58 9.81 ± 5.06 3.11 ± 2.24 tations of small datasets. By implementing weight regularization, we
[62] prevent the model from overfitting, ultimately improving its general­
Unet++ [64] 87.37 ± 2.29 81.30 ± 2.89 9.05 ± 1.49 2.19 ± 0.72
ization ability and overall performance. Through these strategies, ACS­
TransUnet [65] 87.86 ± 2.51 81.38 ± 3.28 8.39 ± 1.39 2.74 ± 1.21
CUNet 87.26 ± 3.70 80.90 ± 3.74 10.95 ± 3.35 ± 2.36 Net effectively identifies and handles common patterns and artifacts in
4.33 images, ensuring accurate analysis and recognition of tumor features,
Chen et al. [48] 86.32 ± 2.19 79.98 ± 1.87 9.79 ± 2.48 3.11 ± 1.96 thereby significantly enhancing the overall segmentation and classifi­
Qu et al. [49] 83.06 ± 4.05 75.56 ± 3.68 13.51 ± 3.50 ± 1.74 cation performance and effectively exploiting contextual information.
4.29
Wang et al. [47] 85.76 ± 4.54 79.40 ± 4.91 9.98 ± 4.37 3.35 ± 2.46
The analysis of the experimental results on the segmentation task
ACSNet 89.12 ± 83.04 ± 7.76 ± 1.73 2.05 ± leads to several conclusions. U-shaped structures with skip connections
2.31 2.60 1.38 effectively merge low-level features from the encoding stage with high-
level features from the decoding stage, leading to commendable seg­
mentation results. Furthermore, multi-task learning shows improved

Fig. 8. Visualization of segmentation results of different methods on BUS, (a)–(m) from left to right are (a) input image, (b) ground truth, and (c) the segmentation
results of ACSNet (Ours), (d) UNet, (e) Unet++, (f) FPN, (g) FCN, (h) DeeplabV3+, (i) TransUnet, (j) CUNet, (k) Qu et al. [49], (l) Chen et al. [48], and (m) Wang
et al. [47] respectively.

10
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

segmentation performance compared to single-task segmentation References


methods. Visual results from breast lesion segmentation experiments
highlight the effectiveness of our proposed method, showing improved [1] H. Sung, J. Ferlay, R.L. Siegel, M. Laversanne, I. Soerjomataram, A. Jemal, F. Bray,
Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality
performance compared to TransUnet. The gate unit and the spatial worldwide for 36 cancers in 185 countries, Ca - Cancer J. Clin. 71 (2021) 209–249.
attention mechanism play critical roles in overcoming challenges such [2] Y. Zhou, H. Chen, Y. Li, Q. Liu, X. Xu, S. Wang, P.-T. Yap, D. Shen, Multi-task
as irregular boundaries and variable morphologies, contributing to the learning for segmentation and classification of tumors in 3D automated breast
ultrasound images, Med. Image Anal. 70 (2021) 101918.
observed improvement. In the segmentation task, ACSNet excels on the [3] K. Yang, A. Suzuki, J. Ye, H. Nosato, A. Izumori, H. Sakanashi, Multi-task learning
Dice and Jaccard metrics, demonstrating its ability to accurately identify with consistent prediction for efficient breast ultrasound tumor detection, in: 2022
and delineate tumor regions. The robust performance observed high­ IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2022,
pp. 3201–3208.
lights the effectiveness of our advanced feature extraction and seg­ [4] Y. Chen, X.-H. Yang, Z. Wei, A.A. Heidari, N. Zheng, Z. Li, H. Chen, H. Hu, Q. Zhou,
mentation techniques in improving the accuracy of breast cancer Q. Guan, Generative adversarial networks in medical image augmentation: a
diagnosis. review, Comput. Biol. Med. 144 (2022) 105382.
[5] M. Rostami, M. Oussalah, K. Berahmand, V. Farrahi, Community detection
In classification tasks, our model demonstrates high accuracy and
algorithms in healthcare applications: a systematic review, IEEE Access 11 (2023)
precision, both of which are critical for clinical diagnosis. This suggests 30247–30272.
that ACSNet not only accurately segments tumor regions, but also pro­ [6] J. Zhou, Z. Wu, Z. Jiang, K. Huang, K. Guo, S. Zhao, Background selection schema
vides important information about the nature of the tumors, thereby on deep learning-based classification of dermatological disease, Comput. Biol. Med.
149 (2022) 105966.
helping to formulate treatment strategies. Despite its state-of-the-art [7] J. Wang, D. Wu, Y. Gao, X. Wang, X. Li, G. Xu, W. Dong, Integral real-time
performance, ACSNet has some limitations. Firstly, there may be a locomotion mode recognition based on GA-CNN for lower limb exoskeleton, JBE 19
classification bias due to the imbalance between benign and malignant (2022) 1359–1373.
[8] Y. Wang, T. Bai, T. Li, L. Huang, Osteoporotic vertebral fracture classification in X-
tumor samples. Secondly, the manual adjustment of weights in the rays based on a multi-modal semantic consistency network, JBE 19 (2022)
multi-task loss function requires further research into more intelligent 1816–1829.
methods. Thirdly, segmentation challenges include over- and under- [9] C.-f. Chen, Z.-j. Du, L. He, Y.-j. Shi, J.-q. Wang, W. Dong, A novel gait pattern
recognition method based on LSTM-CNN for lower limb exoskeleton, JBE 18
segmentation of complex BUS images, which requires the develop­ (2021) 1059–1072.
ment of an effective boundary detection module. Future research di­ [10] Q. Guan, Y. Chen, Z. Wei, A.A. Heidari, H. Hu, X.-H. Yang, J. Zheng, Q. Zhou,
rections include the development of a unified large-model approach H. Chen, F. Chen, Medical image augmentation for lesion detection using a texture-
constrained multichannel progressive GAN, Comput. Biol. Med. 145 (2022)
integrated with a CAD system to accelerate the diagnosis of tumor dis­ 105444.
eases and to address clinical needs in ultrasound imaging tasks. [11] O. Ronneberger, P. Fischer, T. Brox, U-net: convolutional networks for biomedical
image segmentation, in: N. Navab, J. Hornegger, W.M. Wells, A.F. Frangi (Eds.),
Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015,
7. Conclusion
Springer International Publishing, Cham, 2015, pp. 234–241.
[12] R. Almajalid, J. Shan, Y. Du, M. Zhang, Development of a deep-learning-based
In this paper, we have presented ACSNet, a multi-task learning method for breast ultrasound image segmentation, in: 2018 17th IEEE
network designed for simultaneous segmentation and classification of International Conference on Machine Learning and Applications (ICMLA), 2018,
pp. 1103–1108.
BUS images. Our network incorporates a multi-scale feature connectiv­ [13] Z. Ning, S. Zhong, Q. Feng, W. Chen, Y. Zhang, SMU-net: saliency-guided
ity network for classification, and integrates a channel attention mech­ morphology-aware U-net for breast lesion segmentation in ultrasound image, IEEE
anism to enhance classification features. For segmentation, we designed Trans. Med. Imag. 41 (2022) 476–490.
[14] X. Zhao, Y. Pang, L. Zhang, H. Lu, L. Zhang, Suppress and balance: a simple gated
the DSAModule, which uses deformable convolutions to implement a network for salient object detection, in: A. Vedaldi, H. Bischof, T. Brox, J.-
spatial attention mechanism that captures features representing tumor M. Frahm (Eds.), Computer Vision – ECCV 2020, Springer International Publishing,
morphological changes. In addition, our gate unit module enables the Cham, 2020, pp. 35–51.
[15] S. Feng, H. Zhao, F. Shi, X. Cheng, M. Wang, Y. Ma, D. Xiang, W. Zhu, X. Chen,
selective transfer of valuable contextual information from the encoder to CPFNet: context pyramid fusion network for medical image segmentation, IEEE
the decoder, thereby reducing the impact of redundant or misleading Trans. Med. Imag. 39 (2020) 3008–3018.
information on the results. We evaluated the effectiveness of ACSNet [16] M.I. Daoud, S. Abdel-Rahman, R. Alazrai, Breast ultrasound image classification
using a pre-trained convolutional neural network, in: 2019 15th International
using two publicly available breast ultrasound image datasets. The Conference on Signal-Image Technology & Internet-Based Systems (SITIS), 2019,
experimental results showed that the proposed ACSNet not only out­ pp. 167–171.
performed mainstream multi-task learning methods for both breast [17] M. Saad, M. Ullah, H. Afridi, F.A. Cheikh, M. Sajjad, BreastUS: vision transformer
for breast cancer classification using breast ultrasound images, in: 2022 16th
tumor segmentation and classification tasks, but also achieved state-of-
International Conference on Signal-Image Technology & Internet-Based Systems
the-art results for BUS tumor segmentation. (SITIS), 2022, pp. 246–253.
[18] W. Yang, S. Zhang, Y. Chen, W. Li, Y. Chen, Measuring shape complexity of breast
lesions on ultrasound images, in: Medical Imaging 2008: Ultrasonic Imaging and
CRediT authorship contribution statement
Signal Processing, SPIE, 2008, pp. 169–178.
[19] Y. Zhang, Q. Yang, An overview of multi-task learning, Natl. Sci. Rev. 5 (2018)
Qiqi He: Writing – original draft, Methodology, Conceptualization. 30–43.
Qiuju Yang: Writing – review & editing, Supervision, Resources, [20] B. Shareef, M. Xian, A. Vakanski, H. Wang, Breast ultrasound tumor classification
using a hybrid multitask CNN-transformer network, in: H. Greenspan,
Funding acquisition. Hang Su: Validation, Supervision, Investigation. A. Madabhushi, P. Mousavi, S. Salcudean, J. Duncan, T. Syeda-Mahmood, R. Taylor
Yixuan Wang: Visualization, Validation, Investigation. (Eds.), Medical Image Computing and Computer Assisted Intervention – MICCAI
2023, Springer Nature Switzerland, Cham, 2023, pp. 344–353.
[21] G. Zhang, K. Zhao, Y. Hong, X. Qiu, K. Zhang, B. Wei, SHA-MTL: soft and hard
Declaration of competing interest attention multi-task learning for automated breast cancer ultrasound image
segmentation and classification, Int. J. Comput. Assist. Radiol. Surg. 16 (2021)
All authors declare that they have no known competing financial 1719–1725.
[22] M. Xu, K. Huang, X. Qi, Multi-task learning with context-oriented self-attention for
interests or personal relationships that could have appeared to influence breast ultrasound image classification and segmentation, in: 2022 IEEE 19th
the work reported in this paper. International Symposium on Biomedical Imaging (ISBI), 2022, pp. 1–5.
[23] X. Zhao, L. Zhang, H. Lu, Automatic Polyp Segmentation via Multi-Scale
Subtraction Network, Springer International Publishing, 2021, pp. 120–130.
Acknowledgments [24] J. Shan, Y. Wang, H.D. Cheng, Completely automatic segmentation for breast
ultrasound using multiple-domain features. 2010 IEEE International Conference on
This work was supported by the Natural Science Basic Research Plan Image Processing, 2010, pp. 1713–1716.
[25] J. Shan, H.D. Cheng, W. Yuxuan, A novel automatic seed point selection algorithm
in Shaanxi Province of China under Grant 2023-JC-YB-228, and the
for breast ultrasound images. 2008 19th International Conference on Pattern
Open Fund of State Key Laboratory of Loess and Quaternary Geology Recognition, 2008, pp. 1–4.
under Grant SKLLQGZR2201.

11
Q. He et al. Computers in Biology and Medicine 173 (2024) 108319

[26] A. Madabhushi, D.N. Metaxas, Combining low-, high-level and empirical domain aware HoVer-transformer for ROI-free breast cancer diagnosis in ultrasound
knowledge for automated segmentation of ultrasonic breast lesions, IEEE Trans. images, IEEE Trans. Med. Imag. (2023) 1, 1.
Med. Imag. 22 (2003) 155–169. [46] M. Xu, K. Huang, X. Qi, A regional-attentive multi-task learning framework for
[27] C.-M. Chen, H.H.-S. Lu, Y.-S. Huang, Cell-based dual snake model: a new approach breast ultrasound image segmentation and classification, IEEE Access 11 (2023)
to extracting highly winding boundaries in the ultrasound images, Ultrasound Med. 5377–5392.
Biol. 28 (2002) 1061–1073. [47] P. Wang, V.M. Patel, I. Hacihaliloglu, Simultaneous Segmentation and
[28] M. Xian, Y. Zhang, H.D. Cheng, Fully automatic segmentation of breast ultrasound Classification of Bone Surfaces from Ultrasound Using a Multi-Feature Guided
images based on breast characteristics in space and frequency domains, Pattern CNN, Springer International Publishing, 2018, pp. 134–142.
Recogn. 48 (2015) 485–497. [48] C. Chen, W. Bai, D. Rueckert, Multi-task Learning for Left Atrial Segmentation on
[29] X. Guofang, M. Brady, J.A. Noble, Z. Yongyue, Segmentation of ultrasound B-mode GE-MRI, Springer International Publishing, 2019, pp. 292–301.
images with intensity inhomogeneity correction, IEEE Trans. Med. Imag. 21 (2002) [49] H. Qu, G. Riedlinger, P. Wu, Q. Huang, J. Yi, S. De, D. Metaxas, Joint Segmentation
48–57. and Fine-Grained Classification of Nuclei in Histopathology Images, IEEE.
[30] B. Liu, H.D. Cheng, J. Huang, J. Tian, X. Tang, J. Liu, Fully automatic and [50] J. Hu, L. Shen, G. Sun, Squeeze-and-Excitation networks, in: 2018 IEEE/CVF
segmentation-robust classification of breast tumors based on local texture analysis Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141.
of ultrasound images, Pattern Recogn. 43 (2010) 280–298. [51] K. Simonyan, A. Zisserman, Very Deep Convolutional Networks for Large-Scale
[31] C. Xue, L. Zhu, H. Fu, X. Hu, X. Li, H. Zhang, P.-A. Heng, Global guidance network Image Recognition, 2014 arXiv preprint arXiv:1409.1556.
for breast lesion segmentation in ultrasound images, Med. Image Anal. 70 (2021) [52] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in:
101989. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016,
[32] B. Shareef, A. Vakanski, P.E. Freer, M. Xian, ESTAN: enhanced small tumor-aware pp. 770–778.
network for breast ultrasound image segmentation, Healthcare 10 (2022) 2262. [53] J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, Y. Wei, Deformable convolutional
[33] Y. Hu, Y. Guo, Y. Wang, J. Yu, J. Li, S. Zhou, C. Chang, Automatic tumor networks, IEEE.
segmentation in breast ultrasound images using a dilated fully convolutional [54] F. Yuan, L. Zhang, X. Xia, Q. Huang, X. Li, A gated recurrent network with dual
network combined with an active contour model, Med. Phys. 46 (2019) 215–228. classification assistance for smoke semantic segmentation, IEEE Trans. Image
[34] M.H. Yap, G. Pons, J. Marti, S. Ganau, M. Sentis, R. Zwiggelaar, A.K. Davison, Process. 30 (2021) 4409–4422.
R. Marti, Automated breast ultrasound lesions detection using convolutional neural [55] P.-T. De Boer, D.P. Kroese, S. Mannor, R.Y. Rubinstein, A tutorial on the cross-
networks, IEEE Journal of Biomedical and Health Informatics 22 (2018) entropy method, Ann. Oper. Res. 134 (2005) 19–67.
1218–1226. [56] W. Al-Dhabyani, M. Gomaa, H. Khaled, A. Fahmy, Dataset of breast ultrasound
[35] Y. LeCun, B.E. Boser, J.S. Denker, D. Henderson, R.E. Howard, W.E. Hubbard, L. images, Data Brief 28 (2020) 104863.
D. Jackel, Handwritten Digit Recognition with a Back-Propagation Network, NIPS, [57] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin,
1989. N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison,
[36] L. Zhu, R. Chen, H. Fu, C. Xie, L. Wang, L. Wan, P.-A. Heng, A Second-Order A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, S. Chintala, PyTorch: an
Subregion Pooling Network for Breast Lesion Segmentation in Ultrasound, Springer imperative style, high-performance deep learning library, in: Proceedings of the
International Publishing, 2020, pp. 160–170. 33rd International Conference on Neural Information Processing Systems, Curran
[37] A. Vakanski, M. Xian, P.E. Freer, Attention-enriched deep learning model for breast Associates Inc., 2019. Article 721.
tumor segmentation in ultrasound images, Ultrasound Med. Biol. 46 (2020) [58] B. Chen, Y. Liu, Z. Zhang, G. Lu, A.W.K. Kong, TransAttUnet: multi-level attention-
2819–2833. guided U-net with transformer for medical image segmentation, IEEE Transactions
[38] Q. He, Q. Yang, M. Xie, HCTNet: a hybrid CNN-transformer network for breast on Emerging Topics in Computational Intelligence (2023) 1–14.
ultrasound image segmentation, Comput. Biol. Med. 155 (2023) 106629. [59] H.Y. Zhou, J. Guo, Y. Zhang, X. Han, L. Yu, L. Wang, Y. Yu, nnFormer: volumetric
[39] X. Qi, L. Zhang, Y. Chen, Y. Pi, Y. Chen, Q. Lv, Z. Yi, Automated diagnosis of breast medical image segmentation via a 3D transformer, IEEE Trans. Image Process. 32
ultrasonography images using deep neural networks, Med. Image Anal. 52 (2019) (2023) 4036–4045.
185–198. [60] B.N. Li, C.K. Chui, S. Chang, S.H. Ong, A new unified level set method for semi-
[40] M. Byra, P. Jarosik, A. Szubert, M. Galperin, H. Ojeda-Fournier, L. Olson, automatic liver tumor segmentation on contrast-enhanced CT images, Expert Syst.
M. O’Boyle, C. Comstock, M. Andre, Breast mass segmentation in ultrasound with Appl. 39 (2012) 9661–9668.
selective kernel U-Net convolutional neural network, Biomed. Signal Process [61] E. Shelhamer, J. Long, T. Darrell, Fully convolutional networks for semantic
Control 61 (2020) 102027. segmentation, IEEE Trans. Pattern Anal. Mach. Intell. 39 (2017) 640–651.
[41] X. Qian, J. Pei, H. Zheng, X. Xie, L. Yan, H. Zhang, C. Han, X. Gao, H. Zhang, [62] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, H. Adam, Encoder-Decoder with
W. Zheng, Q. Sun, L. Lu, K.K. Shung, Prospective assessment of breast cancer risk Atrous Separable Convolution for Semantic Image Segmentation, Springer
from multimodal multiview ultrasound images via clinically applicable deep International Publishing, 2018, pp. 833–851.
learning, Nat. Biomed. Eng. 5 (2021) 522–532. [63] T.Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, S. Belongie, Feature pyramid
[42] W. Cui, Y. Peng, G. Yuan, W. Cao, Y. Cao, Z. Lu, X. Ni, Z. Yan, J. Zheng, FMRNet: a networks for object detection, in: 2017 IEEE Conference on Computer Vision and
fused network of multiple tumoral regions for breast tumor classification with Pattern Recognition (CVPR), 2017, pp. 936–944.
ultrasound images, Med. Phys. 49 (2022) 144–157. [64] Z. Zhou, M.M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, UNet++: A Nested U-Net
[43] Z. Zhuang, Y. Kang, A.N. Joseph Raj, Y. Yuan, W. Ding, S. Qiu, Breast ultrasound Architecture for Medical Image Segmentation, Springer International Publishing,
lesion classification based on image decomposition and transfer learning, Med. 2018, pp. 3–11.
Phys. 47 (2020) 6257–6269. [65] J. Chen, Y. Lu, Q. Yu, X. Luo, E. Adeli, Y. Wang, L. Lu, Alan, Y. Zhou, arXiv Pre-
[44] S.J. Pan, Q. Yang, A survey on transfer learning, IEEE Trans. Knowl. Data Eng. 22 print Server, in: TransUNet: Transformers Make Strong Encoders for Medical Image
(2010) 1345–1359. Segmentation, 2021.
[45] Y. Mo, C. Han, Y. Liu, M. Liu, Z. Shi, J. Lin, B. Zhao, C. Huang, B. Qiu, Y. Cui, L. Wu,
X. Pan, Z. Xu, X. Huang, Z. Li, Z. Liu, Y. Wang, C. Liang, HoVer-trans: anatomy-

12

Common questions

Powered by AI

ACSNet outperforms other multi-task learning methods such as CUNet, Wang et al., Chen et al., and Qu et al. in both segmentation and classification tasks. It achieves the highest values in Dice and Jaccard coefficients, with improvements of 1.26% and 1.66% over TransUnet respectively, along with reduced 95HD and ASD values. For classification, ACSNet improves mean accuracy to 94.44% ± 2.07, surpassing other methods .

The multi-scale feature fusion method in ACSNet allows the integration of features at various scales, providing comprehensive context and detailed information for both segmentation and classification tasks. This fusion enables the model to capture essential information from different resolution levels, enhancing both spatial resolution for segmentation and feature depth for classification, thereby significantly improving overall model performance .

Breast ultrasound image segmentation and classification face challenges such as variations in tumor size and shape, irregular and blurred tumor boundaries, and low signal-to-noise ratio. ACSNet addresses these challenges by employing a multi-task learning model that effectively extracts and utilizes both low-level and high-level features. It integrates a DSAModule to adapt to morphological changes and a gate unit to optimize feature information flow. Additionally, it employs channel attention mechanisms and multi-scale feature fusion to enhance classification accuracy .

Transfer learning techniques in models for breast ultrasound image segmentation, such as those compared to ACSNet, allow these models to leverage pre-trained knowledge from large datasets to improve segmentation performance on smaller datasets. This approach helps in better feature representation and reduces the need for extensive training data, leading to improvements in segmentation accuracy as observed in methods like UNet++ and TransUnet .

The channel attention mechanism in ACSNet enhances the model's classification performance by focusing on the most informative features across channels. This mechanism helps to suppress less important features, leading to more accurate feature extraction and improved classification results. It plays a critical role in achieving better accuracy and precision in identifying breast cancer from ultrasound images .

The DSAModule in ACSNet adapts to the morphological variability of breast tumors, which is crucial for accurate segmentation. It helps the model to dynamically adjust its segmentation strategies according to the shape and size changes of tumors, improving overall segmentation accuracy by handling the elusive nature of tumor boundaries effectively .

ACSNet exploits contextual information effectively by integrating both low-level and high-level features through its multi-scale feature fusion and channel attention mechanisms. It uses the context surrounding tumors to better differentiate between actual tumor boundaries and similar background patterns, enhancing accuracy in detection and classification .

The primary contributions of ACSNet include the development of a novel multi-task learning model that simultaneously performs segmentation and classification more effectively than previous methods. It introduces DSAModule for morphological adaptation, gate units for optimized information flow, and integrates channel attention and multi-scale feature fusion to enhance classification. These innovations address issues like tumor size, shape variability, and irregular tumor boundaries more efficiently than traditional methods .

Experimental evidence for ACSNet's state-of-the-art results in BUS tumor segmentation includes achieving the highest Dice and Jaccard values among tested models, with values increased by 1.26% and 1.66% over TransUnet respectively. It also demonstrated reductions in 95HD and ASD by 0.63 mm and 0.69 mm. Visual results showed fewer false detections and better alignment with the ground truth .

Using LossNet in multi-task learning frameworks like ACSNet provides structural detail supervision without requiring additional training, which helps suppress noise and improves feature mapping. This enhances overall learning performance by preserving important structural features, leading to better segmentation and classification outcomes by maintaining the balance between different task demands .

You might also like