BiRefNet for High-Resolution DIS
BiRefNet for High-Resolution DIS
zhengpeng0108@[Link]
GT Ours GT Ours
Figure 1. Visual comparison between the results of our proposed BiRefNet and the latest state-of-the-art methods (e.g., IS-Net [37]
and UDUN [34]) for high-resolution dichotomous image segmentation (DIS). Details of segmentation are zoomed in for better display.
Abstract 1. Introduction
We introduce a novel bilateral reference framework With the advancement in high-resolution image acquisi-
(BiRefNet) for high-resolution dichotomous image segmen- tion, image segmentation technology has evolved from tra-
tation (DIS). It comprises two essential components: the ditional coarse localization to achieving high-precision ob-
localization module (LM) and the reconstruction module ject segmentation. This task, whether it involves salient [12]
(RM) with our proposed bilateral reference (BiRef). LM or concealed object detection [10, 11], is referred to as high-
aids in object localization using global semantic informa- resolution dichotomous image segmentation (DIS) [37] and
tion. Within the RM, we utilize BiRef for the reconstruc- has attracted widespread attention and use in the industry,
tion process, where hierarchical patches of images pro- e.g., by Samsung, Adobe, and Disney.
vide the source reference, and gradient maps serve as the For the new DIS task, recent works have considered
target reference. These components collaborate to gen- strategies such as intermediate supervision [37], frequency
erate the final predicted maps. We also introduce auxil- prior [58], and unite-divide-unite [34], and have achieved
iary gradient supervision to enhance the focus on regions favorable results. Essentially, they either split the supervi-
with finer details. In addition, we outline practical train- sion [34, 37] at the feature-level or introduce an additional
ing strategies tailored for DIS to improve map quality and prior [58] to enhance feature extraction. These strategies
the training process. To validate the general applicability are, however, still insufficient to capture very fine features
of our approach, we conduct extensive experiments on four (see Fig. 1). Based on our observations, we found that fine
tasks to evince that BiRefNet exhibits remarkable perfor- and non-salient features in image objects can be well re-
mance, outperforming task-specific cutting-edge methods flected by obtaining gradient features through derivative op-
across all benchmarks. Our codes are publicly available at erations on the original image. In addition, when certain
[Link] positions exhibit high similarity in color and texture to the
background, the gradient features are probably too weak.
† Peng finished the majority of this work when he was a visiting
For such cases, we further introduce ground-truth (GT) fea-
scholar at Nankai University. tures for side supervision, allowing the framework to learn
∗ Corresponding author (dengpfan@[Link]). the characteristics of these positions. We name the incorpo-
1
Supervision
on grad-maps
…
Enc Dec Enc Dec Enc Dec Enc Dec …
…
(a) (b) (c) (d)
Figure 2. Comparison between our proposed BiRefNet and other existing methods for HR segmentation tasks. (a) Common frame-
work [38]; (b) Image pyramid as input [13, 56]; (c) Scaled images as inward reference [21, 27]; (d) BiRefNet: patches of original images
at original scales as inward reference and gradient priors as outward reference. Enc = encoder, Dec = decoder.
ration of the image reference and the introduction of both 2. Related Works
the gradient and GT references as bilateral reference.
2.1. High-Resolution Class-agnostic Segmentation
We propose a novel progressive bilateral reference net-
work BiRefNet to handle the high-resolution DIS task with High-resolution class-agnostic segmentation has been a typ-
separate localization and reconstruction modules. For the ical computer vision objective for decades, and many re-
localization module, we extract hierarchical features from lated tasks have been proposed and attracted much atten-
vision transformer backbone, which are combined and tion, such as dichotomous image segmentation (DIS) [37],
squeezed to obtain corase predictions in low resolution in high-resolution salient object detection (HRSOD) [53], and
deep layers. For the reconstruction module, we further de- concealed object detection (COD) [10]. To provide stan-
sign the inward and outward references as bilateral refer- dard HRSOD benchmarks, several typical HRSOD datasets
ences (BiRef), in which the source image and the gradi- (e.g., HRSOD [53], UHRSD [46], HRS10K [6]) and nu-
ent map are fed into the decoder at different stages. In- merous approaches [21, 42, 46, 53] have been proposed.
stead of resizing the original images to lower-resolution Zeng et al. [53] employed a global-local fusion of the multi-
versions to ensure consistency with decoding features at scale input in their network. Xie et al. [46] used cross-
each stage [21, 27], we keep the original resolution for in- model grafting modules to process images at different scales
tact detail features in inward reference and adaptively crop from multiple backbones (lightweight [16] and heavy [30]).
them into patches for compatibility with decoding features. Pyramid blending was also used in [21] for a lower compu-
In addition, we investigate and summarize practical strate- tational cost. Concealed objects are difficult to locate due
gies for the training on high-resolution (HR) data, including to similar-looking surrounding distractors [9]. Therefore,
long training and region-level loss for better segmentation image priors, such as frequency [57], boundary [40], gradi-
in parts of fine details, and multi-stage supervision to accel- ent [19], etc, are used as auxiliary guidance to train COD
erate learning of them. models. Furthermore, a higher resolution has been found
Our main contributions are summarized as follows: beneficial for detecting targets [17–19]. To produce more
precise and fine-detail segmentation results, Yin et al. [50]
1. We present a bilateral reference network (BiRefNet), employed progressive refinement with masked separable at-
which is a simple yet strong baseline to perform high- tention. Li et al. [27] incorporated the original images at
quality dichotomous image segmentation. different scales to aid in the refining process.
2. We propose a bilateral reference module, which con- High-resolution DIS is a newly proposed task that fo-
sists of an inward reference with source image guidance cuses more on the complex slender structure of target ob-
and an outward reference with gradient supervision. It jects in high-resolution images, making it even more chal-
shows great efficacy in the reconstruction of the pre- lenging. Qin et al. [37] proposed the DIS5K dataset and
dicted HR results. IS-Net with intermediate supervision to alleviate the loss of
3. We explore and summarize various practical strategies fine areas. In addition, Zhou et al. [58] embedded a fre-
tailored for DIS to easily improve performance, predic- quency prior to their DIS network to capture more details.
tion quality, and convergence acceleration. Pei et al. [34] applied a label decoupling strategy [45] to
4. The proposed BiRefNet shows its excellent performance the DIS task and achieved competitive segmentation per-
and strong generalization capabilities to achieve state- formance in the boundary areas of objects. Yu et al. [52]
of-the-art performance on not only the DIS5K task but used patches of HR images to accelerate the training in a
also on HRSOD and COD with 6.8%, 2.0%, and 5.6% more memory-efficient way. Unlike previous models that
average Sm [7] improvements, respectively. used compressed/resized images to enhance HR segmenta-
tion, we utilized intact HR images as supplementary infor-
2
Localization Module Reconstruction Module Predictor
I Transformer F1l F1d+ M
1x1
1x1
Blocks F1d LBCE, IoU
1x1
Auxiliary ′
Cls F3d
Classification Blocks Ĝ3
Supervision F2d
F2e Outward
on Seg-Maps Feature Reference
Supervision BiRef Block
on Grad-Maps
M2
Supervision F3l F3d+
Transformer
·
1x1
on Class
Reconstruction Block
Multi-scale Blocks
Supervision N
F3e F3d F3d+ {Pk=1 }
· Element-wise
Multiplication
Transformer F
e
ASPP
F d
BiRef Block Feature
Inward
Reference
LBCE
LIoU
LBCE
S Sigmoid Blocks
Cls LCE
Figure 3. Pipeline of the proposed bilateral reference Network (BiRefNet). BiRefNet mainly consists of the localization module (LM)
and the reconstruction module (RM) with bilateral reference (BiRef) blocks. Please refer to Sec. 3.1 for details.
3
d
Fi−1
Fid+ AG
· i Mi Deformable
DFC1x1 BR S DFC
Convolution
Conv Dilate
Adaptatively Crop BR BN+ReLU
DFC3x3 BR Fiθ ′
Conv1x1
Conv1x1
d
… B Fi B Dilate
Dilation
N }
{Pk=1 C
DFC7x7 BR
C R
Mi
R Gm
i · Operation
Supervision
I AvgPool on Grad-Maps
1024×1024 FiG Ĝi Ggt
i
Inward Reference Reconstruction Block Outward Reference C Concatenation
Figure 4. Pipeline of the proposed bilateral reference blocks. The source images at the original scale are combined with decoder
features as the inward reference and fed into the reconstruction block, where deformable convolutions with hierarchical receptive fields are
employed. The aggregated features are then used to predict the gradient maps in the outward reference. Gradient-aware features are then
turned into the attention map to act on the original features.
ously [49], which is important for HR tasks involved, we challenging. To deal with these two main problems, we pro-
employ ASPP modules [3] here for multi-context fusion. pose bilateral reference, consisting of an inward reference
F e is squeezed to F d for transfer to the reconstruction mod- (InRef) and an outward reference (OutRef), which is illus-
ule. trated in Fig. 4. Inward reference and outward reference
play the roles of supplementing HR information and draw-
3.3. Reconstruction Module
ing attention to areas with dense details, respectively.
The setting of the receptive field (RF) has been a challenge In InRef, images I with original high resolution are
N
of HR segmentation. Small RFs lead to inadequate context cropped to patches {Pk=1 } of consistent size with the out-
information to locate the right target on a large background, put features of the corresponding decoder stage. These
whereas large RFs often result in insufficient feature extrac- patches are stacked with the original feature Fid+ to be fed
tion in detailed areas. To achieve balance, we propose the into the RM. Existing methods with similar techniques ei-
reconstruction block (RB) in each BiRef block as a replace- ther add I only at the last decoding stage [34] or resize I to
ment for the vanilla residual blocks. In RB, we employ de- make it applicable with original features in low resolution.
formable convolutions [4] with hierarchical receptive fields Our inward reference avoids these two problems through
(i.e., 1×1, 3×3, 7×7) and an adaptive average pooling layer adaptive cropping and supplies the necessary HR informa-
to extract features with RFs of various scales. These fea- tion at every stage.
tures extracted by different RFs are then concatenated as In OutRef, we use gradient labels to draw more attention
Fiθ , followed by a 1×1 convolution layer and a batch nor- to areas of richer gradient information, which is essential
′
malization layer to generate the output feature of RM Fid . for the segmentation of fine structures. First, we extract the
d
In the reconstruction module, the squeezed feature F is fed gradient maps of the input images as Ggt i . Meanwhile, Fi
θ
G
into the BiRef block for the feature F3d . With F3l , the first is used to generate the feature Fi to produce the predicted
BiRef block predicts coarse maps, which are then recon- gradient maps Ĝi . With this gradient supervision, FiG is
structed into higher-resolution versions through the follow- sensitive to the gradient. It passes through a conv and a
ing BiRef blocks. Following [29], the output feature of each sigmoid layer and is used to generate the gradient referring
d′
BiRef block Fid is added with its lateral feature Fil of the attention AGi , which is then multiplied by Fi to generate
LM at each stage, i.e., {Fid+ = Upsample ↑ (Fid +Fil ), i = d
output of the BiRef block as Fi−1 .
, 3, 2, 1}. Meanwhile, all BiRef blocks generate intermedi- Considering that the background may have non-target
ate predictions {Mi }1i=3 by multi-stage supervision, with noise with a lot of gradient information, we apply a masking
resolutions in ascending order. Finally, the last decoding strategy to alleviate the influence of non-target areas. We
feature F1d+ is passed through an 1×1 convolution layer to perform morphological operations on intermediate predic-
obtain the final predicted maps M ∈ RN ×1×H×W . tions Mi and use dilated Mi as a mask. The mask is used
to multiply the gradient map Ggt m
i to generate Gi , where the
3.4. Bilateral Reference gradients outside the mask area are removed.
In DIS, HR training images are very important for deep
3.5. Objective Function
models to learn details and perform highly accurate seg-
mentation. However, most segmentation models follow pre- In HR segmentation tasks, using only pixel-level supervi-
vious works [29, 38] to design the network architecture in sion (BCE loss) usually results in the deterioration of de-
an encoder-decoder structure with down-sampling and up- tailed structural information in HR data. Inspired by the
sampling, respectively. Besides, due to the large size of great results in [35] which used a hybrid loss, we use BCE,
the input, concentrating on the target objects becomes more IoU, SSIM, and CE losses together to collaborate for the
4
supervision on the levels of pixel, region, boundary, and 3.6. Training Strategies Tailored for DIS
semantic, respectively. The final objective function is a
Due to the high cost of training models on HR data, we
weighted combination of the above losses and can be for-
have explored training tricks for HR segmentation tasks to
mulated as:
improve performance and reduce training costs.
L = Lpixel + Lregion + Lboundary + Lsemantic First, we found that our model converges relatively
(1) quickly in the localization of targets and the segmenta-
= λ1 LBCE + λ2 LIoU + λ3 LSSIM + λ4 LCE ,
tion of rough structures (measured by F-measure [1], S-
measure [7]) on DIS5K (e.g., 200 epochs). However, the
where λ1 , λ2 , λ3 , and λ4 are respectively set to 30, 0.5, 10, performance in segmenting fine parts is still increasing af-
and 5 to keep all the losses on the same quantitative level at ter very long training (e.g., 400 epochs), which is reflected
the beginning of the training. The final objective function in metrics such as Fβω and HCEγ . Second, though long
consists of binary cross-entropy (BCE) loss, intersection training can easily achieve great results in terms of both
over union (IoU) loss, structural similarity index measure structure and edges, it consumes too much computation; we
(SSIM) loss, and cross-entropy (CE) loss. The complete found that multi-stage supervision can dramatically accel-
definition of losses can be found below. erate the learning on segmenting fine details and make the
• BCE loss: pixel-aware supervision, which is used for model achieve similar performance as before but with only
pixel-level supervision for the generation of binary maps: 30% training epochs. Third, we also found that fine-tuning
with only region-level losses can easily improve the bina-
[G(i,j) log(M (i,j))+(1−G(i,j)) log(1−M (i,j))], rization of predicted results and those metric scores (e.g.,
P
LBCE =−
(i,j)
(2) Fβω , Eϕm , HCE) that are closer to practical use. Finally, we
where G(i, j) and M(i, j) denote the value of the GT and used context feature fusion and image pyramid inputs on the
binarized predicted maps, respectively, at pixel (i, j). backbone, which are commonly used tricks to process HR
• IoU loss: region-aware supervision for the enhancement images with deep models. In experiments, these two mod-
of binary map predictions: ifications to the backbone achieved a general improvement
in DIS and similar HR segmentation tasks.
H P
P W As shown in Tab. 1, we show the effectiveness of training
M (i,j)G(i,j)
r=1 c=1 epochs and the multi-stage supervision. As the results show,
LIoU = 1 − H W
. (3)
P P
[M (i,j)+G(i,j)−M (i,j)G(i,j)] our BiRefNet can achieve relatively good results after 200
r=1 c=1
epochs of training. Continuous training in 400 epochs can
increase a small portion of metrics measuring structural in-
• SSIM loss: boundary-aware supervision to improve the
formation (e.g., Fβx , Sm ), while bringing a larger improve-
accuracy in boundary parts. Given GT maps G and
ment in metrics measuring fine details (e.g., HCEγ ).
predicted maps M, y = {yj : j = 1, ..., N 2 } and
Although simple long training can achieve better results,
x = {xj : j = 1, ..., N 2 } represent the pixel values of
the improvement is relatively small, concerning its high
two corresponding N × N patches derived from G and
computational cost on HR data. We investigated the multi-
M, respectively. SSIM (x, y) is defined as:
stage supervision (MSS), which is a widely used training
strategy used in binary segmentation works [37, 54]. Differ-
(2µx µy + C1 )(2σxy + C2 )
LSSIM = 1 − , (4) ent from MSS in these works for higher precision, it plays
(µ2x + µ2y + C1 )(σx2 + σy2 + C2 ) a role in accelerating the training convergence. As the re-
sults in Tab. 1 show, our BiRefNet trained for 200 epochs
where µx , µy and σx , σy are the means and standard de- with MSS can achieve similar performance with it trained
viations of x and y, respectively, σxy is their covariance. for 400 epochs. MSS successfully cut training time in half
C is used to avoid division by zero. and can be used for further HR segmentation tasks for more
• CE loss: semantic-aware supervision, which is used to efficient training.
learn better semantic representation:
4. Experiments
N
X
LCE = − yo,c log(po,c ), (5) 4.1. Datasets
c=1
Training Sets. For DIS, we follow [34, 37, 58] to use
where N is the number of classes, yo,c states whether DIS5K-TR as our training set in experiments. For HRSOD,
class label c is the correction classification for observa- we follow [46] to set different combinations of HRSOD,
tion o, and po,c denotes the predicted probability that o is UHRSD, and DUTS as the training set. For COD, we fol-
of class c. low [10, 18] to use the concealed samples in CAMO-TR
5
Table 1. Quantitative ablation studies of the proposed multi- Table 2. Quantitative ablation studies of the proposed com-
stage supervision for acceleration and training epochs. ponents in the proposed BiRefNet. The ablation studies are con-
ducted on the effectiveness of the proposed components, including
Settings DIS-VD reconstruction module (RM), inward reference (InRef), outward
MSS Epoch Fβx ↑ Fβω ↑ M↓ Sm ↑ Eϕm ↑ HCEγ ↓ reference (OutRef), and their combinations.
200 .875 .848 .041 .886 .914 1207
400 .897 .863 .036 .905 .937 1039 Modules DIS-VD
✓ 200 .892 .858 .037 .901 .932 1043 RM InRef OutRef Fβx ↑ Fβω ↑ M ↓ Sm ↑ Eϕm ↑ HCE ↓
γ
.837 .785 .056 .845 .887 1204
✓ .855 .831 .048 .865 .895 1167
✓ .848 .825 .050 .857 .903 1152
and COD10K-TR as the training set. ✓ ✓ .869 .834 .041 .886 .912 1093
Test Sets. To obtain a complete evaluation of ✓ ✓ .863 .831 .042 .891 .918 1106
our BiRefNet, we tested it on all test sets in DIS5K (DIS- ✓ ✓ .861 .839 .044 .881 .911 1114
✓ ✓ ✓ .889 .851 .038 .900 .924 1065
TE1, DIS-TE2, DIS-TE3, and DIS-TE4). We also con-
ducted an evaluation of BiRefNet on the HRSOD test sets
(DAVIS-S [53], HRSOD-TE [53], and UHRSD-TE [46])
and the COD test sets (CAMO-TE [25], COD10K-TE [9], and globally. E-measure is defined as:
and NC4K [31]). Low-resolution SOD test sets (DUTS-
TE [44] and DUT-OMRON [48]) are additionally used for 1 XX
W H
supplementary experiments. Eξ = ϕξ (x, y), (8)
W H x=1 y=1
4.2. Evaluation Protocol
For a comprehensive evaluation, we employ the widely where ϕξ indicates the enhanced alignment matrix. Sim-
used metrics, i.e., S-measure [7] (Sm ), max/mean/weighted ilarly to the F-measure, we also adopt the maximum E-
F-measure [1] (Fβx /Fβm /Fβω ), max/mean E-measure [8] measure (Eξx ) and also the mean E-measure (Eξm ) as our
(Eξx /Eξm ), mean absolute error (MAE), and relax HCE [37] evaluation metrics.
(HCEγ ) to evaluate performance. Detailed descriptions of • MAE (mean absolute error, ϵ) is a simple pixel-level eval-
these metrics can be found as follows. uation metric that measures the absolute difference be-
tween non-binarized predicted results M and GT maps
• S-measure [7] (structure measure, Sα ) is a structural sim-
G. It is defined as:
ilarity measurement between a saliency map and its cor-
responding GT map. Evaluation with Sα can be obtained W H
1 XX
at high speed without binarization. The Sα -measure is ϵ= |Ŷ (x, y) − G(x, y)|. (9)
computed as: W H x=1 y=1
6
Table 3. Effectiveness of practical strategies for training high- classes according to their label names and added an auxil-
resolution segmentation. The experimental comparison of the iary classification head at the end of the encoder. In the de-
proposed several tricks for HR segmentation tasks is provided coder of the baseline network, each decoder block is made
here, including context feature fusion (CFF), image pyramids in-
up of two residual blocks [16]. All stages of the encoder
put (IPT), regional loss fine-tuning (RLFT), and their combina-
tions. The results are obtained by our final model.
and decoder are connected with an 1×1 convolution, ex-
cept the deepest stage, where an ASPP [3] block is used for
Modules DIS-VD connectivity. With this setup, our baseline network has out-
CFF IPT RLFT Fβx ↑ Fβω ↑ M ↓ Sm ↑ Eϕm ↑ HCE ↓
γ performed existing DIS models in most metrics, as shown
.889 .851 .038 .900 .924 1065 in Tab. 2 and Tab. 4.
✓ .893 .856 .038 .904 .928 1054
✓ .895 .857 .037 .904 .927 1051 Reconstruction Module. As shown in Tab. 2, our model
✓ .890 .861 .036 .899 .932 1043 gains an overall improvement with the proposed RM. The
✓ ✓ ✓ .897 .863 .036 .905 .937 1039 RM provides multi-scale receptive fields on the HR features
for local details and overall semantics. It brings ∼2.2% Fβx
relative improvement with little extra computational cost.
Bilateral Reference. We separately investigate the ef-
fectiveness of the inward reference (InRef, with source im-
ages) and the outward reference (OutRef, with gradient la-
bels) in BiRef. InRef supplemented lossless HR informa-
tion globally, while OutRef drew more attention to the fine-
detail parts to achieve higher precision in those areas. As
shown in Tab. 2, they work jointly to bring 2.9% Fβx rela-
tive improvement to BiRefNet. RM and BiRef are combined
to achieve 6.2% Fβx relative improvement.
Training Strategies. As shown in Tab. 3, the proposed
strategies improve performance from different perspectives.
CCF and IPT improve overall performance, while RLFT
Figure 5. Quantitative comparisons of the pro-
specifically improves precision in edge details, which is re-
posed BiRefNet and the best task-specific models. S-
measure [7] is used for the comparison here. UDUN [34], flected in metrics such as Fβω and HCEγ .
FSPNet [18], PGNet-UH [46], and PGNet-DH [46] are currently
the best models for the DIS, COD, HRSOD, and SOD tasks, 4.5. State-of-the-Art Comparison
respectively. To validate the general applicability of our method, we
conduct extensive experiments on four tasks, i.e., high-
resolution dichotomous image segmentation (DIS), high-
DIS/HRSOD/COD tasks for 600/150/150 epochs, respec- resolution salient object detection (HRSOD), concealed ob-
tively. The model is fine-tuned with the IoU loss for the last ject detection (COD), and salient object detection (SOD).
20 epochs. The initial learning rate is set to 10−4 and 10−5 We compare our proposed BiRefNet with all the latest task-
for DIS and others, respectively. Models are trained with specific models on existing benchmarks [10, 25, 31, 37, 44,
PyTorch [33] on eight NVIDIA A100 GPUs. The batch size 46, 48, 53].
is set to N =4 for each GPU during training.
Quantitative Results. Tab. 4 shows a quantitative com-
parison between the proposed BiRefNet and previous state-
4.4. Ablation Study
of-the-art methods. Our BiRefNet outperforms all previ-
We study the effectiveness of each component (i.e., RM and ous methods in widely used metrics. The complexities of
BiRef) and practical strategies (i.e., CFF, IPT, and RLFT) DIS-TE1∼DIS-TE4 are in ascending order. The metrics for
introduced for our BiRefNet and conduct an investigation structure similarity (e.g., Sα , Eϕx ) focus more on global in-
about their contributions to improved DIS results. Quanti- formation. Pixel-level metrics, such as MAE (M ), empha-
tative results regarding each module and strategy are shown size the precision of details. Metrics based on mean values
in Tab. 2 and 3, respectively. (e.g., Eϕm , Fϕm ) better match the requirements of practical
Baseline. We provide a simple but strong encoder- applications where maps are thresholded. As seen in Tab. 4,
decoder network as the baseline for the DIS task. To capture our BiRefNet outperforms previous methods not only on the
better hierarchical features on various scales, we chose the accuracy of the global shape but also in the details of the
Swin transformer large [30] as our default backbone net- pixels. It is noteworthy that the results are better, especially
work. Then, to obtain a better semantic representation in in metrics that cater more to practical applications.
the DIS task, we divided the images in DIS-TR into 219 Additionally, our BiRefNet outperforms existing task-
7
DIS-TE1
DIS-TE2
DIS-TE3
DIS-TE4
DIS-VD
Image GT Ours UDUN [34] IS-Net [37] U2 Net [36] HRNet [43]
Figure 6. Qualitative comparisons of the proposed BiRefNet and previous methods on the DIS5K benchmark. The results of the
previous methods are from [34], where all models are trained with images in 1024×1024. Zoom in for a better view.
Tiny
Slim
Occluded
Multiple
Figure 7. Visual comparisons of the proposed BiRefNet and other competitors on COD10K benchmark. Samples with different
challenges are provided here to show the superiority of BiRefNet from different perspectives.
specific models on the HRSOD and COD tasks. As on both high-resolution and low-resolution SOD bench-
shown in Tab. 5, BiRefNet achieved much higher accuracy marks. Compared with the previous SOTA method [46],
8
Table 4. Quantitative comparisons between our BiRefNet and the state-of-the-art methods on DIS5K. “↑” (“↓”) means that the higher
(lower) is better. We use the results from [34], where all methods take 1024×1024 input.
BASNet19 [35] .663 .577 .105 .741 .756 155 .738 .653 .096 .781 .808 341 .790 .714 .080 .816 .848 681
U2 Net20 [36] .701 .601 .085 .762 .783 165 .768 .676 .083 .798 .825 367 .813 .721 .073 .823 .856 738
HRNet20 [43] .668 .579 .088 .742 .797 262 .747 .664 .087 .784 .840 555 .784 .700 .080 .805 .869 1049
PGNet22 [46] .754 .680 .067 .800 .848 162 .807 .743 .065 .833 .880 375 .843 .785 .056 .844 .911 797
IS-Net22 [37] .740 .662 .074 .787 .820 149 .799 .728 .070 .823 .858 340 .830 .758 .064 .836 .883 687
FP-DIS23 [58] .784 .713 .060 .821 .860 160 .827 .767 .059 .845 .893 373 .868 .811 .049 .871 .922 780
UDUN23 [34] .784 .720 .059 .817 .864 140 .829 .768 .058 .843 .886 325 .865 .809 .050 .865 .917 658
BiRefNet .860 .819 .037 .885 .911 106 .894 .857 .036 .900 .930 266 .925 .893 .028 .919 .955 569
BiRefNetSwinB .857 .819 .038 .884 .912 110 .890 .854 .037 .898 .930 275 .919 .886 .030 .915 .953 597
BiRefNetSwinT .823 .774 .048 .855 .887 117 .862 .821 .046 .877 .912 290 .899 .860 .036 .897 .942 627
BiRefNetP V T v2b2 .839 .796 .042 .870 .903 111 .881 .842 .040 .888 .925 280 .903 .866 .036 .901 .941 614
DIS-TE4 (500) DIS-TE (1-4) (2,000) DIS-VD (470)
Methods
Fβx ↑ Fβω ↑ M ↓ Sm ↑ Eϕ
m ↑ HCE ↓ F x ↑ F ω ↑ M ↓ S
γ β β
m x ω m
m ↑ Eϕ ↑ HCEγ ↓ Fβ ↑ Fβ ↑ M ↓ Sm ↑ Eϕ ↑ HCEγ ↓
BASNet19 [35] .785 .713 .087 .806 .844 2852 .744 .664 .092 .786 .814 1007 .737 .656 .094 .781 .809 1132
U2 Net20 [36] .800 .707 .085 .814 .837 2898 .771 .676 .082 .799 .825 1042 .753 .656 .089 .785 .809 1139
HRNet20 [43] .772 .687 .092 .792 .854 3864 .743 .658 .087 .781 .840 1432 .726 .641 .095 .767 .824 1560
PGNet22 [46] .831 .774 .065 .841 .899 3361 .809 .746 .063 .830 .885 1173 .798 .733 .067 .824 .879 1326
IS-Net22 [37] .827 .753 .072 .830 .870 2888 .799 .726 .070 .819 .858 1016 .791 .717 .074 .813 .856 1116
FP-DIS23 [58] .846 .788 .061 .852 .906 3347 .831 .770 .047 .847 .895 1165 .823 .763 .062 .843 .891 1309
UDUN23 [34] .846 .792 .059 .849 .901 2785 .831 .772 .057 .844 .892 977 .823 .763 .059 .838 .892 1097
BiRefNet .904 .864 .039 .900 .939 2723 .896 .858 .035 .901 .934 916 .891 .854 .038 .898 .931 989
BiRefNetSwinB .899 .860 .040 .895 .938 2836 .891 .855 .036 .898 .933 954 .881 .844 .039 .890 .925 1029
BiRefNetSwinT .880 .834 .049 .878 .925 2888 .866 .822 .045 .877 .916 980 .862 .819 .045 .874 .917 1070
BiRefNetP V T v2b2 .890 .846 .045 .886 .929 2871 .878 .838 .041 .886 .925 969 .868 .827 .044 .880 .919 1073
Table 5. Quantitative comparisons between our BiRefNet and the state-of-the-art methods in high-resolution and low-resolution
SOD datasets. TR denotes the training set. To provide a fair comparison, we train our BiRefNet with different combinations of training
sets, where 1, 2, and 3 represent DUTS [44], HRSOD [53], and UHRSD [46], respectively.
High-Resolution Benchmarks Low-Resolution Benchmarks
Test Sets
DAVIS-S (92) HRSOD-TE (400) UHRSD-TE (988) DUTS-TE (5,019) DUT-OMRON(5,168)
Methods TR
Sm ↑ Fβx ↑ m
Eϕ ↑ M ↓ Sm ↑ Fβx ↑ m
Eϕ ↑ M ↓ Sm ↑ Fβx ↑ m
Eϕ ↑ M ↓ Sm ↑ Fβx ↑ m
Eϕ ↑ M ↓ Sm ↑ Fβx ↑ Eϕ
m ↑M↓
LDF20 [45] 1 .922 .911 .947 .019 .904 .904 .919 .032 .888 .913 .891 .047 .892 .898 .910 .034 .838 .820 .873 .051
HRSOD19 [53] 1,2 .876 .899 .955 .026 .896 .905 .934 .030 - - - - .824 .835 .885 .050 .762 .743 .831 .065
DHQ21 [42] 1,2 .920 .938 .947 .012 .920 .922 .947 .022 .900 .911 .905 .039 .894 .900 .919 .031 .836 .820 .873 .045
PGNet22 [46] 1 .935 .936 .947 .015 .930 .931 .944 .021 .912 .931 .904 .037 .911 .917 .922 .027 .855 .835 .887 .045
PGNet22 [46] 1,2 .948 .950 .975 .012 .935 .937 .946 .020 .912 .935 .905 .036 .912 .919 .925 .028 .858 .835 .887 .046
PGNet22 [46] 2,3 .954 .957 .979 .010 .938 .945 .946 .020 .935 .949 .916 .026 .859 .871 .897 .038 .786 .772 .884 .058
BiRefNet 1 .967 .966 .984 .008 .957 .958 .972 .014 .931 .933 .943 .030 .939 .937 .958 .019 .868 .813 .878 .040
BiRefNet 1,2 .973 .976 .990 .006 .962 .963 .976 .011 .937 .942 .951 .024 .938 .935 .960 .018 .868 .818 .882 .040
BiRefNet 1,3 .975 .977 .989 .006 .959 .958 .972 .014 .952 .960 .965 .019 .942 .942 .961 .018 .881 .837 .896 .036
BiRefNet 2,3 .976 .980 .990 .006 .956 .953 .967 .016 .952 .958 .964 .019 .933 .928 .954 .020 .864 .810 .879 .040
BiRefNet 1,2,3 .975 .979 .989 .006 .962 .961 .973 .013 .957 .963 .969 .016 .944 .943 .962 .018 .882 .839 .896 .038
our BiRefNet achieved an average improvement of 2.0% shown in Fig. 5, where we run our model and the best task-
Sm . Furthermore, as shown in Tab. 6, in the COD task, specific models on DIS/HRSOD/COD/SOD tasks. As the
BiRefNet also shows a much better performance compared results show, our BiRefNet achieves leading results in all
to the previous SOTA models, with an average improvement four tasks. The other task-specific models show their weak-
of 5.6% Sm on the three widely used COD benchmarks. ness in similar HR segmentation tasks. For example, FSP-
These results show the remarkable generalization ability of Net [18] ranks second in COD benchmarks, while it ranks
our BiRefNet to similar HR tasks. fourth/third/third in DIS/HRSOD/SOD tasks, respectively.
For a clearer illustration of the generalizability and pow- Qualitative Results. Fig. 6 shows segmentation maps
erful performance of BiRefNet, we provide a radar picture produced by the most competitive existing DIS models and
9
Table 6. Comparison of BiRefNet with recent methods. As seen, BiRefNet performs much better than previous methods.
CAMO (250) COD10K (2,026) NC4K (4,121)
Methods
Sm ↑ Fβω ↑ Fβm ↑ m
Eϕ ↑ x
Eϕ ↑ M↓ Sm ↑ Fβω ↑ Fβm ↑ m
Eϕ ↑ x
Eϕ ↑ M↓ Sm ↑ Fβω ↑ Fβm ↑ Eϕ
m ↑ Ex ↑ M ↓
ϕ
SINet20 [9] .751 .606 .675 .771 .831 .100 .771 .551 .634 .806 .868 .051 .808 .723 .769 .871 .883 .058
BGNet22 [40] .812 .749 .789 .870 .882 .073 .831 .722 .753 .901 .911 .033 .851 .788 .820 .907 .916 .044
SegMaR22 [20] .815 .753 .795 .874 .884 .071 .833 .724 .757 .899 .906 .034 .841 .781 .820 .896 .907 .046
ZoomNet22 [32] .820 .752 .794 .878 .892 .066 .838 .729 .766 .888 .911 .029 .853 .784 .818 .896 .912 .043
SINetv222 [10] .820 .743 .782 .882 .895 .070 .815 .680 .718 .887 .906 .037 .847 .770 .805 .903 .914 .048
FEDER23 [14] .802 .738 .781 .867 .873 .071 .822 .716 .751 .900 .905 .032 .847 .789 .824 .907 .915 .044
HitNet23 [17] .849 .809 .831 .906 .910 .055 .871 .806 .823 .935 .938 .023 .875 .834 .853 .926 .929 .037
FSPNet23 [18] .856 .799 .830 .899 .928 .050 .851 .735 .769 .895 .930 .026 .879 .816 .843 .915 .937 .035
BiRefNet .904 .890 .904 .954 .959 .030 .913 .874 .888 .960 .967 .014 .914 .894 .909 .953 .960 .023
Table 7. Comparison of different DIS methods on the performance, efficiency, and model complexity. Full details can be referred
to [Link]
Model Runtime (ms) #Params (MB) MACs (G) DIS-TEs (HCE, Fβω )
BiRefNetSwinL 83.3 215 1143 916, .858
BiRefNetSwinL cp 78.3 215 1143 916, .858
BiRefNetSwinB 61.4 101 561 954, .855
BiRefNetSwinT 40.9 39 231 980, .822
BiRefNetP V T v2b2 47.8 35 195 969, .838
BiRefNetP V T v2b1 36.6 23 147 978, .817
BiRefNetP V T v2b0 32.9 11 89 1013, .806
IS-Net 16.0 44 160 1016, .726
UDUNRes50 33.5 25 142 977, .772
curved edges.
0.78 BiRefNet-SwinT
BiRefNet-PVT_v2_b2 We also provide a qualitative comparison on the COD
0.76 BiRefNet-PVT_v2_b1 task. Fig. 7 shows hard samples with different challenges.
BiRefNet-PVT_v2_b0 For example, in the row of the occluded frog, the area
0.74 IS-Net of the frog is divided by the branch that covers it, while
UDUN our BiRefNet can accurately segment the scattered frag-
0.72
20 30 40 50 60 70 80 ments almost the same as the GT map. In contrast, in the
Runtime (ms) ↓ results of the other methods, fragments are difficult to find
all, let alone to provide precise segmentation maps. For tiny
Figure 8. Comparison of the efficiency, size, complexity, and per- and slim objects, BiRefNet shows a better ability to find the
formance of BiRefNet and existing DIS methods. right target. Our BiRefNet also shows superiority in finding
multiple concealed objects.
the proposed BiRefNet. As the results show, we provide Efficiency and Complexity Comparison. We equip
samples of all test sets and one validation set. BiRefNet our BiRefNet with different backbones to obtain models in
outperforms the previous DIS methods from two perspec- different sizes. Runtime, number of parameters, MACs, and
tives, i.e., the location of target objects and the more accu- performance of them are further tested to provide a com-
rate segmentation of the details of the objects. For example, prehensive comparison between them and other methods.
in the samples of DIS-TE4 and DIS-TE2, there are neigh- First, we provide a quantitative comparison in Tab. 7. The
boring distractors that attract the attention of other models FPS of the largest BiRefNet can be more than 10, which is
10
A D F
E
B
Figure 9. Potenial applications and selected existing third-party applications based on BiRefNet, and visual comparisons on social
media. (A) Potential application #1. Building crack detection for the maintenance of architecture health. (B) Potential application #2.
Highly accurate object extraction in high-resolution natural images. (C) A project by viperyl first packs our BiRefNet as a ComfyUI node
and makes this SOTA model easier to use for everyone. (D) ZHO also provides a ComfyUI-based project to further improve the UI for
our BiRefNet, especially for video data. (E) [Link] encapsulates our BiRefNet online with more useful options in UI and API to call the
model. (F) ZHO provides a visual comparison between our BiRefNet and previous SOTA method BRIA RMGB v1.4 which has extra
training on their private training dataset. (G) Toyxyz conducted a comparison between our BiRefNet and previous competitive human
matting methods (e.g., BRIAAI, RemBG, Robust Video Matting, and Person YOLOv8s) with both videos and images.
acceptable in most practical applications. We also used the results when target objects have too high shape complexi-
compiled version (BiRefNetSwinL cp ) by PyTorch 2.0 [33] ties [35, 36] or need manual guidance (e.g., scribble, point,
on BiRefNetSwinL to accelerate its inference by 13%. In ad- and coarse mask) for more accurate segmentation [5, 26].
dition, we draw the performance and runtime of each model The proposed BiRefNet trained on DIS5K can generate re-
in Fig. 8 for a clearer display. BiRefNet with different back- sults with much higher resolution and segment thin threads
bones are evaluated and compared with existing DIS meth- at the hair level without a mask, as shown in Fig. 9 (B).
ods on DIS-TEs and DIS-VD. Different methods are drawn On the basis of such refined results, there may be numerous
in different colors and markers. All tests are conducted on a successful downstream applications in the future.
single NVIDIA A100 GPU and an AMD EPYC 7J13 CPU.
6. Third-Party Creations
5. Potential Applications
Since the release of our project on Mar 7, 2024, it has at-
We envisage that generated fine segmentation maps have the tracted much attention from many researchers and develop-
potential to be utilized in various practical applications. ers in the community to promote it spontaneously. Further-
Potential Application #1 Crack Detection. The qual- more, great third-party applications have also been made
ity of the walls is important for the health of the architec- based on our BiRefNet. Due to the rapid growth of relevant
ture [59]. However, segmentation models trained on com- works, we only list some typical ones.
monly used datasets (e.g., COCO [28]) can only segment #1 Practical Applications. Because of the excellent per-
regular foreground objects. The proposed BiRefNet trained formance of our BiRefNet, more and more third-party ap-
on the DIS5K dataset is more aware of the fine details and plications have been created by developers in the commu-
can also segment targets with higher shape complexities. As nity12 . As shown in Fig. 9 (C and D), some developers have
shown in Fig. 9 (A), our BiRefNet can accurately find cracks integrated our BiRefNet into the ComfyUI as a node, which
in the walls and help maintain when to repair them. helps a lot in matting foreground segmentation to better pro-
Potential Application #2 Highly Accurate Object Ex- cessing in the subsequent stable diffusion models. For bet-
traction. Foreground object extraction and background re- 1 [Link]
moval have been popular applications in recent years. How- 2 https : / / github . com / ZHO - ZHO - ZHO / ComfyUI -
11
ter online access, [Link] has established an online demo We also show that the techniques of BiRefNet can be trans-
of our BiRefNet running on an A6000 GPU34 , as shown ferred and used in many practical applications. We hope
in Fig. 9 (E). In addition to the common prediction of re- that the proposed framework can encourage the develop-
sults, this online application also provides an API service ment of unified models for various tasks in the academic
for easy use with HTTP requests. community and that our model can empower and inspire
#2 Social Media. In recent days, our BiRefNet has drawn the developer community to create more great works.
attention from the community. Many tweets have been
posted on the X platform (formerly Twitter)5 . ZHO pro- References
vides a visual comparison between our BiRefNet and other [1] Radhakrishna Achanta, Sheila Hemami, Francisco Estrada,
methods, as given in Fig. 9 (F). BiRefNet achieves compet- and Sabine Susstrunk. Frequency-tuned salient region detec-
itive results with the previous SOTA method BRIA RMGB tion. In IEEE / CVF Computer Vision and Pattern Recogni-
v1.4 in their tests6 . It should be noted that our BiRefNet tion Conference, 2009. 5, 6
was trained on the training set of the open-source dataset [2] Ali Borji, Ming-Ming Cheng, Huaizu Jiang, and Jia Li.
DIS5K [37] under MIT license, while the other one was Salient object detection: A benchmark. IEEE Transactions
trained on their carefully selected private data and cannot on Image Process., 24:5706–5722, 2015. 6
be used for commercial use. As shown in Fig. 9 (G), more [3] Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian
comparisons on both video and image data have been pro- Schroff, and Hartwig Adam. Encoder-decoder with atrous
separable convolution for semantic image segmentation. In
vided by Toyxyz on X between our BiRefNet and previous
European Conference on Computer Vision Workshop, 2018.
great foreground human matting methods7 . In addition to
4, 7
these posts, PurzBeats has also made an animation with [4] Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong
our BiRefNet and uploaded relevant videos8 . A video tuto- Zhang, Han Hu, and Yichen Wei. Deformable convolutional
rial can also be found on YouTube by ‘AI is in wonderland’ networks. In IEEE / CVF International Conference on Com-
in Japanese about how to use our BiRefNet in ComfyUI9 . puter Vision, 2017. 4
[5] Linhui Dai, Xiang Song, Xiaohong Liu, Chengqi Li, Zhi-
7. Conclusions hao Shi, Martin Brooks, and Jun Chen. Enabling trimap-free
image matting with a frequency-guided saliency-aware net-
This work proposes a BiRefNet framework equipped with work via joint learning. IEEE Transactions on Multimedia,
a bilateral reference, which can perform dichotomous im- 25:4868–4879, 2022. 11
age segmentation, high-resolution (HR) salient object de- [6] Xinhao Deng, Pingping Zhang, Wei Liu, and Huchuan Lu.
tection, and concealed object detection in the same frame- Recurrent multi-scale transformer for high-resolution salient
object detection. In ACM International Conference on Mul-
work. With the comprehensive experiments conducted, we
timedia, 2023. 2
find that unscaled source images and a focus on regions of
[7] Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali
rich information are vital to generating fine and detailed ar- Borji. Structure-measure: A new way to evaluate foreground
eas in HR images. To this end, we propose the bilateral maps. In IEEE / CVF International Conference on Computer
reference to fill in the missing information in the fine parts Vision, 2017. 2, 5, 6, 7
(inward reference) and guide the model to focus more on [8] Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, Ming-
regions with richer details (outward reference). This signif- Ming Cheng, and Ali Borji. Enhanced-alignment measure
icantly improves the model’s ability to capture tiny-pixel for binary foreground map evaluation. In International Joint
features. To alleviate the high training cost of HR data Conference on Artificial Intelligence, 2018. 6
training, we also provide various practical tricks to deliver [9] Deng-Ping Fan, Ge-Peng Ji, Guolei Sun, Ming-Ming Cheng,
higher-quality prediction and faster convergence. Competi- Jianbing Shen, and Ling Shao. Camouflaged object detec-
tive results on 13 benchmarks demonstrate outstanding per- tion. In IEEE / CVF Computer Vision and Pattern Recogni-
tion Conference, 2020. 2, 6, 10
formance and strong generalization ability of our BiRefNet.
[10] Deng-Ping Fan, Ge-Peng Ji, Ming-Ming Cheng, and Ling
3 [Link] Shao. Concealed object detection. IEEE Transactions on
4 Thanks to [Link] for providing us with additional computation re- Pattern Analysis and Machine Intelligence, 44(10):6024–
sources for further explorations on more practical applications. 6042, 2022. 1, 2, 5, 7, 10
5 https : / / twitter . com / search ? q = birefnet & src = [11] Deng-Ping Fan, Ge-Peng Ji, Peng Xu, Ming-Ming Cheng,
typed_query Christos Sakaridis, and Luc Van Gool. Advances in deep
6 https : / / twitter . com / ZHOZHO672070 / status /
concealed scene understanding. Visual Intelligence, 1(1):16,
1771026516388041038 2023. 1
7 https : / / twitter . com / toyxyz3 / status /
1771413245267746952 [12] Deng-Ping Fan, Jing Zhang, Gang Xu, Ming-Ming Cheng,
8 https : / / twitter . com / i / status / and Ling Shao. Salient objects in clutter. IEEE Transactions
1772323682934775896 on Pattern Analysis and Machine Intelligence, 45(2):2344–
9 [Link] 2366, 2023. 1
12
[13] William I Grosky and Ramesh Jain. A pyramid-based ap- image matting. International Journal of Computer Vision,
proach to segmentation applied to region matching. IEEE 130(2):246–266, 2022. 3, 11
Transactions on Pattern Analysis and Machine Intelligence, [27] Xiaofei Li, Jiaxin Yang, Shuohao Li, Jun Lei, Jun Zhang, and
8(5):639–650, 1986. 2, 3 Dong Chen. Locate, refine and restore: A progressive en-
[14] Chunming He, Kai Li, Yachao Zhang, Longxiang Tang, Yu- hancement network for camouflaged object detection. In In-
lun Zhang, Zhenhua Guo, and Xiu Li. Camouflaged object ternational Joint Conference on Artificial Intelligence, 2023.
detection with feature decomposition and edge reconstruc- 2
tion. In IEEE / CVF Computer Vision and Pattern Recogni- [28] Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James
tion Conference, 2023. 10 Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and
[15] Chunming He, Kai Li, Yachao Zhang, Yulun Zhang, Chenyu C. Lawrence Zitnick. Microsoft coco: Common objects in
You, Zhenhua Guo, Xiu Li, Martin Danelljan, and Fisher context. In European Conference on Computer Vision Work-
Yu. Strategic preys make acute predators: Enhancing cam- shop, 2014. 11
ouflaged object detectors by generating camouflaged objects. [29] Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He,
In International Conference on Learning Representations, Bharath Hariharan, and Serge Belongie. Feature pyramid
2023. 3 networks for object detection. In IEEE / CVF Computer Vi-
[16] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. sion and Pattern Recognition Conference, 2017. 4
Deep residual learning for image recognition. In IEEE / CVF [30] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng
Computer Vision and Pattern Recognition Conference, 2016. Zhang, Stephen Lin, and Baining Guo. Swin transformer:
2, 7 Hierarchical vision transformer using shifted windows. In
[17] Xiaobin Hu, Shuo Wang, Xuebin Qin, Hang Dai, Wenqi Ren, IEEE / CVF International Conference on Computer Vision,
Donghao Luo, Ying Tai, and Ling Shao. High-resolution 2021. 2, 3, 7
iterative feedback network for camouflaged object detection. [31] Yunqiu Lv, Jing Zhang, Yuchao Dai, Aixuan Li, Bowen Liu,
In AAAI Conference on Artificial Intelligence, 2023. 2, 10 Nick Barnes, and Deng-Ping Fan. Simultaneously localize,
[18] Zhou Huang, Hang Dai, Tian-Zhu Xiang, Shuo Wang, Huai- segment and rank the camouflaged objects. In IEEE / CVF
Xin Chen, Jie Qin, and Huan Xiong. Feature shrinkage pyra- Computer Vision and Pattern Recognition Conference, 2021.
mid for camouflaged object detection with transformers. In 6, 7
IEEE / CVF Computer Vision and Pattern Recognition Con- [32] Youwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang,
ference, 2023. 5, 7, 9, 10 and Huchuan Lu. Zoom in and out: A mixed-scale triplet
[19] Ge-Peng Ji, Deng-Ping Fan, Yu-Cheng Chou, Dengxin Dai, network for camouflaged object detection. In IEEE / CVF
Alexander Liniger, and Luc Van Gool. Deep gradient learn- Computer Vision and Pattern Recognition Conference, 2022.
ing for efficient camouflaged object detection. Machine In- 10
telligence Research, 20(1):92–108, 2023. 2 [33] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer,
[20] Qi Jia, Shuilian Yao, Yu Liu, Xin Fan, Risheng Liu, and James Bradbury, Gregory Chanan, Trevor Killeen, Zem-
Zhongxuan Luo. Segment, magnify and reiterate: Detecting ing Lin, Natalia Gimelshein, Luca Antiga, et al. PyTorch:
camouflaged objects the hard way. In IEEE / CVF Computer An imperative style, high-performance deep learning library.
Vision and Pattern Recognition Conference, 2022. 10 Advances in Neural Information Processing Systems, 2019.
[21] Taehun Kim, Kunhee Kim, Joonyeong Lee, Dongmin Cha, 7, 11
Jiho Lee, and Daijin Kim. Revisiting image pyramid struc- [34] Jialun Pei, Zhangjun Zhou, Yueming Jin, He Tang, and Heng
ture for high resolution salient object detection. In Asian Pheng-Ann. Unite-divide-unite: Joint boosting trunk and
Conference on Computer Vision, 2022. 2 structure for high-accuracy dichotomous image segmenta-
[22] Diederik P. Kingma and Jimmy Ba. Adam: A method tion. In ACM International Conference on Multimedia, 2023.
for stochastic optimization. In International Conference on 1, 2, 3, 4, 5, 7, 8, 9
Learning Representations, 2015. 6 [35] Xuebin Qin, Zichen Zhang, Chenyang Huang, Chao Gao,
[23] Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming- Masood Dehghan, and Martin Jagersand. Basnet: Boundary-
Hsuan Yang. Deep laplacian pyramid networks for fast and aware salient object detection. In IEEE / CVF Computer
accurate super-resolution. In IEEE / CVF Computer Vision Vision and Pattern Recognition Conference, 2019. 3, 4, 9, 11
and Pattern Recognition Conference, 2017. 3 [36] Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood De-
[24] Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming- hghan, Osmar R Zaiane, and Martin Jagersand. U2-net: Go-
Hsuan Yang. Fast and accurate image super-resolution with ing deeper with nested u-structure for salient object detec-
deep laplacian pyramid networks. IEEE Transactions on Pat- tion. Pattern Recognition, 106:107404, 2020. 8, 9, 11
tern Analysis and Machine Intelligence, 41(11):2599–2613, [37] Xuebin Qin, Hang Dai, Xiaobin Hu, Deng-Ping Fan, Ling
2018. 3 Shao, et al. Highly accurate dichotomous image segmen-
[25] Trung-Nghia Le, Tam V Nguyen, Zhongliang Nie, Minh- tation. In European Conference on Computer Vision Work-
Triet Tran, and Akihiro Sugimoto. Anabranch network for shop, 2022. 1, 2, 3, 5, 6, 7, 8, 9, 12
camouflaged object segmentation. Computer Vision and Im- [38] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net:
age Understanding, 184:45–56, 2019. 6, 7 Convolutional networks for biomedical image segmentation.
[26] Jizhizi Li, Jing Zhang, Stephen J Maybank, and Dacheng In Medical Image Computing and Computer Assisted Inter-
Tao. Bridging composite and real: towards end-to-end deep ventions, 2015. 2, 3, 4
13
[39] Tiancheng Shen, Yuechen Zhang, Lu Qi, Jason Kuen, [52] Qian Yu, Xiaoqi Zhao, Youwei Pang, Lihe Zhang,
Xingyu Xie, Jianlong Wu, Zhe Lin, and Jiaya Jia. High qual- and Huchuan Lu. Multi-view aggregation network
ity segmentation for ultra high-resolution images. In IEEE / for dichotomous image segmentation. arXiv preprint
CVF Computer Vision and Pattern Recognition Conference, arXiv:2404.07445, 2024. 2
2022. 3 [53] Yi Zeng, Pingping Zhang, Jianming Zhang, Zhe Lin, and
[40] Yujia Sun, Shuo Wang, Chenglizhao Chen, and Tian-Zhu Xi- Huchuan Lu. Towards high-resolution salient object detec-
ang. Boundary-guided camouflaged object detection. In In- tion. In IEEE / CVF Computer Vision and Pattern Recogni-
ternational Joint Conference on Artificial Intelligence, 2022. tion Conference, 2019. 2, 6, 7, 9
2, 10 [54] Zhao Zhang, Wenda Jin, Jun Xu, and Ming-Ming Cheng.
[41] Chufeng Tang, Hang Chen, Xiao Li, Jianmin Li, Zhaoxi- Gradient-induced co-saliency detection. In European Con-
ang Zhang, and Xiaolin Hu. Look closer to segment bet- ference on Computer Vision Workshop, 2020. 5
ter: Boundary patch refinement for instance segmentation. In [55] Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang
IEEE / CVF Computer Vision and Pattern Recognition Con- Wang, and Jiaya Jia. Pyramid scene parsing network. In
ference, 2021. 3 IEEE / CVF Computer Vision and Pattern Recognition Con-
[42] Lv Tang, Bo Li, Yijie Zhong, Shouhong Ding, and Mofei ference, 2017. 3
Song. Disentangled high quality salient object detection. In [56] Hengshuang Zhao, Xiaojuan Qi, Xiaoyong Shen, Jianping
IEEE / CVF Computer Vision and Pattern Recognition Con- Shi, and Jiaya Jia. Icnet for real-time semantic segmenta-
ference, 2021. 2, 9 tion on high-resolution images. In European Conference on
[43] Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Computer Vision Workshop, 2018. 2, 3
Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui [57] Yijie Zhong, Bo Li, Lv Tang, Senyun Kuang, Shuang Wu,
Tan, Xinggang Wang, et al. Deep high-resolution represen- and Shouhong Ding. Detecting camouflaged object in fre-
tation learning for visual recognition. IEEE Transactions quency domain. In IEEE / CVF Computer Vision and Pattern
on Pattern Analysis and Machine Intelligence, 43(10):3349– Recognition Conference, 2022. 2
3364, 2020. 8, 9
[58] Yan Zhou, Bo Dong, Yuanfeng Wu, Wentao Zhu, Geng
[44] Lijun Wang, Huchuan Lu, Yifan Wang, Mengyang Feng, Chen, and Yanning Zhang. Dichotomous image segmenta-
Dong Wang, Baocai Yin, and Xiang Ruan. Learning to de- tion with frequency priors. In International Joint Conference
tect salient objects with image-level supervision. In IEEE / on Artificial Intelligence, 2023. 1, 2, 3, 5, 9
CVF Computer Vision and Pattern Recognition Conference,
[59] Qin Zou, Zheng Zhang, Qingquan Li, Xianbiao Qi, Qian
2017. 6, 7, 9
Wang, and Song Wang. Deepcrack: Learning hierarchical
[45] Jun Wei, Shuhui Wang, Zhe Wu, Chi Su, Qingming Huang,
convolutional features for crack detection. IEEE Transac-
and Qi Tian. Label decoupling framework for salient ob-
tions on Image Process., 28(3):1498–1512, 2018. 11
ject detection. In IEEE / CVF Computer Vision and Pattern
Recognition Conference, 2020. 2, 9
[46] Chenxi Xie, Changqun Xia, Mingcan Ma, Zhirui Zhao, Xi-
aowu Chen, and Jia Li. Pyramid grafting network for one-
stage high resolution saliency detection. In IEEE / CVF
Computer Vision and Pattern Recognition Conference, 2022.
2, 5, 6, 7, 8, 9
[47] Ning Xu, Brian Price, Scott Cohen, and Thomas Huang.
Deep image matting. In IEEE / CVF Computer Vision and
Pattern Recognition Conference, 2017. 3
[48] Chuan Yang, Lihe Zhang, Huchuan Lu, Xiang Ruan, and
Ming-Hsuan Yang. Saliency detection via graph-based man-
ifold ranking. In IEEE / CVF Computer Vision and Pattern
Recognition Conference, pages 3166–3173, 2013. 6, 7
[49] Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan
Yang. Denseaspp for semantic segmentation in street scenes.
In IEEE / CVF Computer Vision and Pattern Recognition
Conference, pages 3684–3692, 2018. 4
[50] Bowen Yin, Xuying Zhang, Qibin Hou, Bo-Yuan Sun, Deng-
Ping Fan, and Luc Van Gool. Camoformer: Masked sep-
arable attention for camouflaged object detection. arXiv
preprint arXiv:2212.06570, 2022. 2
[51] Qihang Yu, Jianming Zhang, He Zhang, Yilin Wang, Zhe
Lin, Ning Xu, Yutong Bai, and Alan Yuille. Mask guided
matting via progressive refinement network. In IEEE / CVF
Computer Vision and Pattern Recognition Conference, 2021.
3
14
Regional loss fine-tuning (RLFT) enhances precision specifically in edge details by focusing the learning process on these critical areas. This strategy results in significant gains in metrics such as F ω β and HCEγ, underlining its effectiveness in refining the sharpness and accuracy of edge segmentation in BiRefNet .
The Swin transformer large serves as an effective backbone in BiRefNet by capturing hierarchical features at various scales which are critical for effective semantic representation in high-resolution dichotomous image segmentation. This architectural choice facilitates improved performance across multiple metrics in comparison to existing DIS models .
The Reconstruction Module (RM) contributes to improvements in F x β metrics by providing multi-scale receptive fields for better capturing of local details and overall semantics, yielding a 2.2% improvement. The Bilateral Reference (BiRef) module adds another 2.9% improvement by using inward and outward references to enhance precision globally and in fine-detail areas. The combined effect of RM and BiRef results in a total of 6.2% F x β relative improvement, indicating significant enhancements in capturing detailed image segmentation with minimal computational expense .
In the Bilateral Reference module, the inward reference (InRef) provides global lossless high-resolution information, while the outward reference (OutRef) focuses more on fine-detail areas to achieve higher precision. When used together, they bring about 2.9% F x β relative improvement to the BiRefNet. Combined with the Reconstruction Module, they lead to a total 6.2% improvement in F x β .
The usage of NVIDIA A100 GPUs in the training setup of BiRefNet allows for efficient handling of high-resolution image data. The model is trained on eight GPUs with a batch size of four per GPU, accommodating large-scale data processing and accelerating computational operations which are crucial for optimizing model performance and achieving state-of-the-art results .
The training strategies enhance BiRefNet's performance from various angles. CFF and IPT boost overall performance, while RLFT is specifically focused on enhancing precision in edge details, leading to improvements reflected in metrics such as F ω β and HCEγ. The combined use of these strategies ensures a more comprehensive enhancement of the network's capabilities in handling high-resolution segmentation tasks .
Context Feature Fusion (CFF) improves overall performance in high-resolution segmentation tasks by efficiently merging contextual information across different scales, enabling BiRefNet to better capture intricate details and semantic consistency across an image. This contributes to improved results in metrics like F x β, illustrating its effectiveness .
BiRefNet outperforms previous task-specific models in high-resolution salient object detection (HRSOD) and concealed object detection (COD) tasks. It achieves higher accuracy on both high-resolution and low-resolution benchmarks, showing superior performance in capturing the global shape and pixel detail levels compared to the previous state-of-the-art methods .
The Reconstruction Module (RM) in BiRefNet provides an overall improvement to the model by offering multi-scale receptive fields that enhance both local details and overall semantics. It contributes to a roughly 2.2% F x β relative improvement with minimal additional computational cost, aiding in capturing better hierarchical features for high-resolution image tasks .
Metrics such as Sα (structure similarity), MAE (mean absolute error), and Em ϕ (enhanced metric) are crucial for evaluating BiRefNet. Sα focuses on global information capture, MAE emphasizes precision in details, and Em ϕ combines mean values and thresholds for practical applications. BiRefNet outperforms previous state-of-the-art methods significantly in these metrics, achieving better accuracy in both global shape and pixel-level detail, thereby proving its efficiency and applicability in practical scenarios .