0% found this document useful (0 votes)
19 views7 pages

LightStereo: Efficient 2D Cost Aggregation

LightStereo is a novel stereo-matching network designed to enhance efficiency in 2D cost aggregation by utilizing a 3D cost volume, focusing on the channel dimension to improve performance. It achieves competitive results with minimal computational demands, requiring only 22 GFLOPs and 17 ms of runtime, while outperforming existing state-of-the-art methods in speed and accuracy. The paper introduces inverted residual blocks and a Multi-Scale Convolutional Attention module to optimize cost aggregation, paving the way for real-time stereo vision applications.

Uploaded by

wxlu613
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views7 pages

LightStereo: Efficient 2D Cost Aggregation

LightStereo is a novel stereo-matching network designed to enhance efficiency in 2D cost aggregation by utilizing a 3D cost volume, focusing on the channel dimension to improve performance. It achieves competitive results with minimal computational demands, requiring only 22 GFLOPs and 17 ms of runtime, while outperforming existing state-of-the-art methods in speed and accuracy. The paper introduces inverted residual blocks and a Multi-Scale Convolutional Attention module to optimize cost aggregation, paving the way for real-time stereo vision applications.

Uploaded by

wxlu613
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

LightStereo: Channel Boost Is All You Need

for Efficient 2D Cost Aggregation


∗ ∗
Xianda Guo1, , Chenming Zhang2,3, , Youmin Zhang4,5 , Wenzhao Zheng6 ,

Dujun Nie7,8 , Matteo Poggi4 , Long Chen7,2,3,
1
School of Computer Science, Wuhan University 2 IAIR, Xi’an Jiaotong University
3 4 5 6
Waytous University of Bologna Rock Universe AI University of California, Berkeley
7 8
Institute of Automation, Chinese Academy of Sciences Metoak
xianda guo@[Link]; chmzhang@[Link]; [Link]@[Link]
arXiv:2406.19833v3 [[Link]] 26 Feb 2025

Abstract— We present LightStereo, a cutting-edge stereo-


matching network crafted to accelerate the matching process.
Departing from conventional methodologies that rely on aggre-
gating computationally intensive 4D costs, LightStereo adopts
the 3D cost volume as a lightweight alternative. While similar
approaches have been explored previously, our breakthrough
lies in enhancing performance through a dedicated focus on the
channel dimension of the 3D cost volume, where the distribution
of matching costs is encapsulated. Our exhaustive exploration
has yielded plenty of strategies to amplify the capacity of the
pivotal dimension, ensuring both precision and efficiency. We Fig. 1: Performence vs. Time trade-off. Results on Scene-
compare the proposed LightStereo with existing state-of-the-art Flow (Left) and KITTI15 (Right) datasets.
methods across various benchmarks, which demonstrate its su-
perior performance in speed, accuracy, and resource utilization. right images and introduces intra-scale and cross-scale 2D
LightStereo achieves a competitive EPE metric in the SceneFlow cost aggregation modules to enhance the efficiency and
datasets while demanding a minimum of only 22 GFLOPs and
17 ms of runtime, and ranks 1st on KITTI 2015 among real-time
accuracy of cost aggregation. MobilenetStereo-2D [17] in-
models. Our comprehensive analysis reveals the effect of 2D troduces MobileNet [20], [21] blocks to reduce the cost,
cost aggregation for stereo matching, paving the way for real- but the performance is far from satisfaction. Overall, these
world applications of efficient stereo systems. Code is available methods based on 2D cost aggregation perform poorly. Cost
at [Link] aggregation is critical to accuracy and efficiency in stereo
I. I NTRODUCTION matching yet existing methods require a compromise be-
tween accuracy and speed. However, existing methodologies
Stereo matching is a pivotal task in computer vision,
often necessitate a trade-off between accuracy and processing
which aims to ascertain correspondences between points in
speed. This paper poses the question: Is it possible to design
stereo image pairs to retrieve depth information. This pro-
a lightweight 2D encoder-decoder aggregation net to achieve
cess underpins numerous applications, including autonomous
precise disparity estimation?
driving, robotic navigation, and augmented reality. Despite
From the perspective of cost aggregation in stereo match-
substantial advancements, achieving real-time stereo match-
ing, focusing on the disparity channel dimension presents
ing without sacrificing accuracy or efficiency remains a
several advantages. Firstly, it allows for more direct modeling
formidable challenge, especially on embedded platforms.
of the disparities between corresponding image points, which
Most developments [1], [2], [3], [4], [5], [6], [7], [8],
is crucial for accurate disparity estimation. By focusing
[9], [10], [11], [12], [13], [14] have primarily focused
on this dimension, it becomes possible to more effectively
on leveraging deep learning techniques for accurate stereo
capture and assimilate the critical information needed for
matching. Most methods are based on 3D CNN for cost
robust stereo matching. In this paper, we propose LightStereo
aggregation: because of the disparity dimension being mod-
as a positive solution. We explore 2D cost aggregation for
eled in the constructed 4D cost volume, the network can
stereo matching and leverage inverted residual blocks to
perform exhaustively aggregation for accurate matching, yet
enhance both accuracy and computational efficiency. The
at the expense of high memory consumption and runtime.
model is specifically designed to address the challenges of
Nevertheless, the overall runtime of these iterative-based
real-time stereo vision applications, focusing on reducing
methods is over 100ms on a custom GPU.
computational demands without compromising the quality of
There are also efforts [15], [16], [17], [18], [19] focusing
disparity estimation. Specifically, we utilize inverted residual
on lightweight design for stereo matching. AANet [18]
blocks for 2D cost aggregation in stereo matching, which
constructs a 3D cost volume by correlating the left and
focuses on the disparity channel dimension of the 3D cost
* These authors contributed equally to this work. volume rather than the dimension of height and width. In
† Corresponding Author inverted residual blocks, the expansion phase increases the
MSCA
MSCA

1/4 x4 Final Disparity


1/16
D x16
1
1
1
1 16
8 1 ...
Left Image 8 H Cost
4 4
W

2D Regularization Network

Context Network
Right Image

Fig. 2: The diagram above illustrates the architecture of LightStereo.


channel dimension, allowing the network to learn richer neural architecture search (NAS) framework to enhance deep
features at a reduced computational cost before compressing stereo matching performance. IGEV [30] constructs a unified
them back. In addition, inspired by the effectiveness of large geometry encoding volume, integrating geometry, contex-
kernel and strip convolutions in image segmentation [22], tual cues, and local matching intricacies. OpenStereo [14]
[23], [24], we propose a Multi-Scale Convolutional Attention conducted a comprehensive benchmark with a focus on
Module (MSCA) module for enhancing cost aggregation by practical applicability and introduced StereoBase, which
extracting features from the left image. By leveraging multi- further elevates the performance ceiling of stereo matching.
scale image features to excite the channel dimension of the Selective-IGEV [31] proposes the Selective Recurrent Unit
cost volume, we utilize semantic information inherent in the to help integrate disparity information across frequencies,
images (such as object-level semantic details) to guide the minimizing loss during iterative steps. In addition, work [32]
cost aggregation process. When encountering discontinuities improves the loss function by introducing ADL, an adaptive
in disparity, the network halts propagation. multi-modal cross-entropy loss, to guide network learning of
Our main contributions are as follows: diverse pixel distribution patterns.
• We propose LightStereo, which achieves a competitive Real-Time-Focused Stereo Matching. A multitude of mod-
End-Point-Error (EPE) in the SceneFLow [25] datasets els have been proposed to target practical applications. Stere-
while demanding a minimum of only 22 Gflops with an oNet [33] uses color inputs to guide hierarchical refinement
inference time of 17 ms. and can recover high-frequency details. DeepPruner[34] pro-
• We propose inverted residual blocks for 2D cost aggre- poses a PatchMatch module with learnable parameters to
gation in stereo matching. save memory and computation by gradually trimming the
• We propose the MSCA module for enhancing cost disparity space to be searched for each pixel. AnyNet[35]
aggregation by extracting features from the left image. has a 2D image convolutional network and a tiny 3D cost
• We verify the effectiveness of LightStereo, achieving aggregation network. FADNet[15] leverages efficient 2D cor-
state-of-the-art performance on the Sceneflow [25] and relation layers with residual structures to perform multi-scale
KITTI [26], [27] benchmark within the lightweight predictions. Based on the learned bilateral grid, BGNet[36]
stereo-matching methods. designs an edge-preserving cost volume up-sampling mod-
ule, allowing computationally expensive operations such as
II. R ELATED W ORK
3D convolution to be performed at low resolution. MADNet
We review the literature concerning deep learning-based series [37], [38] relies on fully 2D modules, maximizing
stereo matching, identifying two categories of methods. We speed at the expense of accuracy. CoEx[39] uses image
refer the reader to [28], [29] for a broader overview. feature-guided weights and cost volume to excite 3D CNN to
Accuracy-Focused Stereo Matching. Many contemporary extract relevant geometric features. Building upon the foun-
stereo-matching methods are dedicated to enhancing accu- dation of StereoNet [33], MobileStereoNet[17] introduces
racy through various techniques and optimizations. Among two models suitable for resource-constrained devices.
these, 3D end-to-end networks have introduced significantly However, these methods still have considerable room for
heightened precision in disparity estimation within stereo improvement in terms of computational efficiency and pa-
matching [2], [3], [7], [12], [30], [14]. PSMNet [2] adopts an rameter scale. After carefully considering and testing various
architecture that leverages spatial pyramid pooling and 3D network structures, we propose a lightweight network, called
convolutional neural networks to learn disparity estimation LightStereo. This model innovatively explores 2D cost aggre-
from stereo images. GwcNet [3] introduces a novel approach gation for stereo matching, leveraging advanced techniques
for constructing the cost volume in stereo matching using to enhance both accuracy and computational efficiency.
group-wise correlation. By dividing left and right features
into groups along the channel dimension, correlation maps III. M ETHOD
are computed within each group to generate multiple match- In the deployment of stereo matching on edge devices,
ing cost proposals. LEAStereo [7] employs a hierarchical constructing a 4D cost volume and using 3D CNNs for cost
1×1
Conv 1x1,
Conv 3x3 Add
Linear Conact&Projection Channel
Mixing
input
Conv 1x1,
head1 head
2 ...
(a) Regular Linear Self-attention
Self-attention
DW 3x3,
Stride=2, Token
d,7×1 d,11×1 d,21×1
Conv 1x1, Token Interaction
Relu6 Interaction
Relu6 DW 3x3,
Q K V
Relu6 Q K V

DW 3x3 d,1×7 d,1×11 d,1×21


Stride=s,
Relu6
Conv 1x1,
Relu6
Conv 1x1,
Relu6 ...
split
Multi-scale
input input input input d,1×1
Stride=1 Stride=2
Feature
block block
(b) V1 block (c) V2 block (d) VIT block Fig. 4: Multi-Scale Convolutional Attention (MSCA).
Fig. 3: Comparison of different blocks used for cost aggre-
gation – DW refers to depthwise separable convolution; V1 where Wproject represents the weights of the projection
Block represents the depthwise separable convolution [20]; convolution. If the input and output dimensions match, a
V2 Block represents an inverted residual block [21]; ViT skip connection is added to improve gradient flow:
Block refers to the block used in EfficientViT [40].
out = out + x if dimensions match. (4)
aggregation prove to be extremely inefficient. Our objective The structure of the inverted residual block signifi-
is to construct a 3D cost volume, where single scores are cantly reduces computational complexity, making it ideal for
stored along the channel dimension, and utilize 2D CNNs resource-constrained environments. Our experimental results
augmented with channel boost for cost aggregation to bal- will demonstrate its superiority over regular CNN block,
ance efficiency and accuracy. We are now going to introduce V1 block [20], and Vision Transformer (ViT) block [40] –
LightStereo, whose overview is shown in Figure 2, starting depicted in Figure 3 (a), (b) and (d) respectively.
with its main components – the Inverted Residual Blocks
[21] and the new Multi-Scale Convolutional Attention – and B. Multi-Scale Convolutional Attention Module
concluding with the overall architectural design. Inspired by the effectiveness of large kernel and strip con-
volutions in image segmentation [22], [23], [24], we propose
A. Inverted Residual Blocks for 2D Cost Aggregation a Multi-Scale Convolutional Attention (MSCA) module for
Previous efforts, such as MobilenetStereo-2D [17], have enhancing cost aggregation by extracting features from the
incorporated MobileNet [20], [21] blocks to decrease com- left image. Figure 4 illustrates the architecture of such a
putational costs. Despite these efforts, the results have not module, designed to capture and integrate features at multiple
met expectations, with an EPE of 1.14 on SceneFLow [25] scales and enhance their representation for cost aggregation.
still being reported. To address this issue, our study intro- The MSCA incorporates a series of depthwise separable
duces inverted residual blocks to boost disparity estimation convolutions with varying kernel sizes, specifically 1 × 1,
accuracy, as shown in Figure 3 (c). Given the cost volume in 7 × 1, 1 × 7, 11 × 1, 1 × 11, 21 × 1, and 1 × 21, capturing
Figure 2 whose size is H W Disp both horizontal and vertical strip-like features, which are
4 × 4 × 4 , the key idea is first to
boost the number of disparity channels, then apply depthwise crucial for identifying elongated structures within the image.
convolution, and finally project the expanded features back to These depth-wise strip convolutions are lightweight and can
a lower-dimensional space, enhancing feature representation replace, for instance, a standard 7 × 7 convolution with 7 × 1
significantly. Inverted residual blocks are utilized at resolu- and 1 × 7 convolutions. Thus, strip convolutions serve as
tions of 14 , 18 , and 16
1
, each corresponding to different blocks: a complement to grid convolutions and aid in extracting
initially, the cost volume C is first passed through a 1 × 1 strip-like features of left images. Utilizing multi-scale image
convolution to increase the number of disparity channel: features, MSCA enhances the channel dimension of the cost
volume, by incorporating semantic information embedded
y = σ(Wexpand ∗ C), (1) within the images to guide the aggregation process. The
network is designed to cease propagation upon detecting
where Wexpand represents the weights of the expansion disparity discontinuities.
convolution, ∗ denotes the convolution operation, and σ is the Given a stereo image of size H × W × 3, feature maps
ReLU6 activation function. Next, the expanded feature map at 14 , 81 , and 16
1
of the original resolution, and processed
y undergoes a 3 × 3 depthwise convolution, which operates by the MSCA module to further extract horizontal and
independently on each channel to capture spatial features: vertical strip-like features. The outputs are then concatenated
to form a comprehensive multi-scale feature representation,
z = σ(Wdepthwise ∗ y), (2)
subsequently processed by a 1 × 1 convolution, which acts
where Wdepthwise are the weights of the depthwise convolu- as a channel mixer. The 1 × 1 convolution blends the multi-
tion. Finally, the result z is then passed through another 1×1 scale features and recalibrates the feature channels, allowing
convolution to restore the original number of channels: the network to focus on relevant information across different
scales. The final output is combined with the aggregated cost
out = Wproject ∗ z, (3) by multiplication, enhancing the effect of cost aggregation.
C. LightStereo – Network Architecture TABLE I: Results on SceneFlow [25]. Comparison with
state-of-the-art models pursuing speed over accuracy.
We now describe the overall LightStereo architecture,
Method FLOPs(G) Params(M) EPE (px) Time(ms)
which consists of four components: feature extraction, cost
DeepPruner-Fast[34] 219.12 7.47 1.25 39
computation, cost aggregation, and disparity prediction. StereoNet[43] 85.93 0.40 1.10 20
Multi-scale Feature Extraction. For a stereo image pair 2D-MobileStereoNet[17] 128.84 2.23 1.11 73
FADNet++[16] 148.21 124.26 0.85 21
with dimensions H × W × 3, our network harnesses a CoEx[39] 53.39 2.72 0.67 36
MobileNetV2 [39], [30] model, previously trained on the Fast-ACVNet [44] 79.34 3.08 0.64 22
Fast-ACVNet+ [44] 93.08 3.20 0.59 27
ImageNet [41] dataset. We enable the extraction of feature AANet[18] 152.86 2.97 0.87 93
maps across four distinct scales, effectively reducing the res- HITNet [45] 50.23 0.42 0.55 36
IINet[46] 90.16 19.54 0.54 26
olution to 14 , 81 , and 16
1 1
and 32 of the initial size, respectively. LightStereo-S(Ours) 22.71 3.44 0.73 17
Subsequently, upsampling blocks with skip connections are LightStereo-M(Ours) 36.36 7.64 0.62 23
LightStereo-L(Ours) 91.85 24.29 0.59 37
utilized to restore these feature maps to 14 scale, which is LightStereo-H(Ours) 159.26 45.63 0.51 54
used to construct the cost volume.
Cost Volume. We construct a correlation volume from the IV. E XPERIMENTS
previous left fl,4 and right features fr,4 . For each disparity d
within the range of 0 to D − 1, the similarity between fl,4 A. Datasets and Evaluation Metrics
and fr,4 shifted by d pixels is computed as:
SceneFlow[25] is a synthetic stereo collection that pro-
vides 35 454 and 4 370 image pairs for training and testing,
C respectively. The dataset has a resolution of 960×540 and
1 X
Ccor (d, h, w) = fl,4 (h, w) · fr,4 (h, w − d). (5) provides dense disparity maps as ground truth. Here, we use
C c=1
the End-Point-Error (EPE) as the evaluation metric.
KITTI. KITTI 2012 [26] and 2015 [27] are real-world
For d = 0, the cost is computed as the average of the
datasets, counting 194/195 and 200/200 image pairs for
element-wise product of feature vectors from fl,4 and fr,4 at
training and testing, respectively, for which sparse ground-
the same spatial location, across all channels.
truth disparities are collected with a LiDAR system. We both
Cost Aggregation. We use inverted residual blocks to
report the results obtained from the online benchmarks, for
aggregate the cost volume at 14 , 18 , and 16 1
resolutions,
which we report the official metrics, as well as assess cross-
and at each resolution, we apply multi-scale convolutional
domain generalization performance, measuring the percent-
attention using the left image, as described in Section III-A
age of pixels with error >3px in this latter case.
and Section III-B. We develop four variants of LightStereo
Middlebury 2014. [47] It collects high-resolution indoor
based on block and expansion variations: LightStereo-S for
scenes, providing 15 image pairs for training and 15 for test-
small-scale applications, LightStereo-M for medium-scale
ing, respectively. We involved this dataset for cross-domain
tasks, LightStereo-L for large-scale tasks, and LightStereo-H
generalization experiments, using half-resolution images and
for huge-scale tasks. In LightStereo-S, the inverted residual
measuring the percentage of pixels with error >2px.
blocks are configured as (1, 2, 4) with an expansion factor of
4. For LightStereo-M, the blocks are set as (4, 8, 14) with an
expansion factor of 4. Finally, LightStereo-L uses blocks (8, B. Implementation Details
16, 32) with an expansion factor of 8. LightStereo-H denotes
a LightStereo-L variant with EfficientnetV2 [42] backbone. LightStereo is implemented in PyTorch and trained on
Disparity Regression. We use soft-argmax [1], [2] to 8 NVIDIA RTX 3090 GPUs. For the SceneFlow [25]
predict the final disparity map d, ˆ determined by summing dataset, the batch sizes for LightStereo-S, LightStereo-M,
each disparity d weighted by its probability σ(cd ), with cd LightStereo-L, and LightStereo-H are set at 24, 12, 8, and
being the predicted cost and σ(·) a softmax layer: 6, [Link] use the AdamW optimizer coupled with
OneCycleLR scheduling, where the maximum learning rate
D
was set to 0.0001 multiplied by the batch size. LightStereo
max
X underwent training for 90 epochs. Only random crop (320 ×
dˆ = d × σ(cd ). (6)
736) is employed. For the KITTI dataset, we fine-tuned
d=0
the pre-trained models on the SceneFlow dataset [25] for
Training Loss. We train LightStereo with smooth L1 loss 500 epochs using a mixed training set comprising KITTI
2012 [26] and KITTI2015 [27] training datasets. Batch size
N is set to 2. OneCycleLR scheduling is used with a max
ˆ = 1 X learning rate of 0.0002. For the generalization experiments,
L(d, d) smoothL1 (di − dˆi ), (7)
N i=1 we employ data augmentation techniques including color
jitter, random erase, random scale, and random crop. In the
where N is the number of labeled pixels, d represents the ablation study, these LightStereo models are trained over 50
ground-truth disparity, and dˆ denotes the predicted disparity. epochs.
EPE: 0.85 EPE: 0.67 EPE: 0.59 EPE: 0.54 EPE: 0.73 EPE: 0.62 EPE: 0.59 EPE: 0.51

RGB FADNet++ CoEx Fast-ACVNet+ IINet LightStereo-S(Ours) LightStereo-M(Ours) LightStereo-L(Ours) LightStereo-H(Ours)

Fig. 5: Qualitative results on SceneFlow [25].

TABLE II: Ablation study on SceneFlow [25] – Conv.


Block Selection. We underline our final choice.
Conv. Type Kernel Block EPE (px) Flops(G) Param(M) Time(ms) Avg EPE: 2.110 Avg EPE: 3.016 Avg EPE: 0.919 Avg EPE: 1.574

Regular 3×3 (4 8 16) 0.7652 36.27 8.04 16.70


Regular 5×5 (4 8 16) 0.7979 71.16 18.62 17.32
Regular 7×7 (4 8 16) 0.8190 123.51 34.49 19.68 RGB Regular V1 block V2 block ViT block
Regular 11 × 11 (4 8 16) 0.8672 280.53 82.10 33.60
Fig. 6: Qualitative results on SceneFlow [25]. Comparison
V1 Block [20] DW 3 × 3 (30 60 120) 0.7801 34.86 7.57 54.21
V2 Block [21] DW 3 × 3 (4 8 16) 0.7144 35.82 7.54 22.93 between LightStereo blocks.
ViT Block [40] - (3 6 9) 0.7149 34.48 6.53 51.14
TABLE IV: Ablation study on SceneFlow [25] – Block
TABLE III: Ablation study on SceneFlow [25] – Backbone structure design. The last row refers to LightStereo-L.
selection. Comparison among efficient feature extractors.
Block Exp. factor SE [51] MSCA EPE (px) Flops(G) Param(M) Time(ms)
Backbone Type EPE (px) Flops(G) Param(M) Time(ms) a (4 8 16) (2 2 2) 0.7557 26.23 4.81 22.44
b (4 8 16) (4 4 4) 0.7144 35.82 7.54 22.93
MobilenetV2 [21] CNN 0.7144 35.82 7.54 22.93 c (4 8 16) (8 8 8) 0.6853 54.99 12.99 23.59
MobilenetV3 [48] CNN 0.7292 35.72 9.16 25.02 d (4 8 16) (16 16 16) 0.6650 93.34 23.90 36.67
StarNet [49] CNN 0.7247 40.21 8.98 26.63 e (1 2 4) (4 4 4) 0.8317 22.17 3.34 16.65
EfficientnetV2 [42] CNN 0.6130 103.14 28.87 46.83 f (2 4 8) (4 4 4) 0.7464 26.72 4.74 20.04
RepVIT [50] Transformer 0.6823 50.45 9.56 28.65 g (4 8 16) (4 4 4) 0.7144 35.82 7.54 22.93
h (8 16 32) (4 4 4) 0.6973 54.02 13.14 31.09
i (4 8 16) (4 4 4) 0.7144 35.82 7.54 22.93
j (4 8 16) (4 4 4) ✓ 0.7036 35.90 12.76 30.14
C. Comparisons with State-of-the-art on SceneFlow k (4 8 16) (4 4 4) ✓ 0.6809 36.36 7.64 23.14
l (4 8 16) (4 4 4) ✓ ✓ 0.6810 36.44 12.86 30.82
Table I compares LightStereo with several state-of-the-art m (1 2 4) (4 4 4) ✓ 0.7899 22.71 3.44 17.59
n (4 8 16) (4 4 4) ✓ 0.6809 36.36 7.64 23.14
stereo matching approaches on SceneFlow [25]. LightStereo- o (8 16 32) (8 8 8) ✓ 0.6382 91.85 24.29 37.55

S runs in 17ms only, being substantially faster than other


methods. Regarding model complexity, LightStereo-S strikes balance, achieving the lowest EPE of 0.7144, and FLOPs of
a favorable balance with only 22.71Gflops, comparable to 35.82G, with an inference time of 22.93ms. This superior
StereoNet [43] (85.93Gflops). This reflects our commitment performance is attributed to the structure of the V2 block,
to maintaining a lightweight model while ensuring com- which incorporates expansion convolutions in the disparity
petitive performance. In terms of accuracy, LightStereo- dimension. The ViT block showed a similar EPE to the
H achieves an EPE of 0.51, resulting in more accuracy V2 block but had a lower parameter count (6.53M) and
than Fast-ACVNet+ [44] (0.59) and IINet [46](0.54), yet higher inference time (51.14ms). Ultimately, the V2 block
remaining competitive in terms of complexity. Overall, our was chosen for its optimal balance between accuracy and
LightStereo framework, particularly the LightStereo-S con- computational efficiency. Figure 6 shows a visual comparison
figuration, presents a compelling solution for real-time stereo with different conv. type. It can be observed that for the
matching, offering a favorable trade-off between computa- occluded area in the lower left corner of the left image,
tional efficiency and depth estimation accuracy. Figure 5 the cost aggregation based on the V2 block achieves better
presents a qualitative comparison of the results by the four results than the other 3 conv. type.
proposed models. Backbone Selection. Table III explores the use of classic,
lightweight models as backbones. MobilenetV2 [21] demon-
D. Ablation Study strates a balanced performance with an EPE of 0.7144,
Conv. Block Selection. In Table II, we highlight the FLOPs of 35.82G, 7.54M parameters, and an inference time
critical importance of disparity dimensions in the cost ag- of 22.93ms. Although MobilenetV3 [48] had slightly lower
gregation process of stereo matching, rather than spatial FLOPs at 35.72G, it showed a higher EPE of 0.7292 and
expansions in the height and width dimensions. On top, we required more parameters (9.16M) with a longer inference
can notice how regular convolutions with larger kernels yield time (25.02ms). StarNet [49] exhibited an EPE of 0.7247
both higher EPE and complexity. This finding indicates that with higher FLOPs (40.21G) and inference time (26.63ms).
merely increasing the spatial extent of convolutions does EfficientnetV2 [42] achieved the lowest EPE of 0.6130, but
not effectively enhance the accuracy of stereo-matching; at the cost of significantly higher computational resources
on the contrary, putting the focus on disparity dimensions (103.14G FLOPs) and parameters (28.87M), with an infer-
is crucial. The V1 block with a 3 × 3 kernel showed ence time of 46.83ms. RepVIT [50], a transformer-based
an EPE of 0.7801, indicating moderate performance with model, showed an EPE of 0.6876 but required 50.45GFLOPs
reduced FLOPs (34.86G), but it had a higher inference and 9.56M parameters with an inference time of 28.65ms.
time (54.21ms). In contrast, the V2 block provided the best Therefore, MobilenetV2 is chosen for its overall efficiency
TABLE V: Results on KITTI online benchmarks. Methods TABLE VII: Generalization performance on multiple
are classified based on runtime (higher or lower than 100ms). datasets. All the models are trained on the SceneFlow only.
We highlight best and second best results. ∗ means time KITTI12 KITTI15 Middle
Method
measured on our RTX 3090 GPU. D1(%) D1(%) 2px(%)
DPF [34] 16.8 15.9 30.83
KITTI 2012[26] KITTI 2015[27] Time BGNet [36] 24.8 20.1 37.00
Method
3-all EPE-all D1-bg D1-fg D1-all (ms) CoEx [39] 13.5 11.6 25.51
GANet[4] 1.60 0.91 1.48 3.46 1.81 1800 FastACV [44] 12.4 10.6 20.13
LaC+GANet[52] 1.42 0.80 1.44 2.83 1.67 1800 IINet [46] 11.6 8.5 19.57
CFNet[53] 1.58 0.92 1.54 3.56 1.88 180 LightStereo-S(Ours) 11.6 9.0 19.63
SegStereo[54] 2.03 1.25 1.88 4.07 2.25 600 LightStereo-M(Ours) 7.0 6.6 17.69
SSPCVNet [55] 1.90 1.08 1.75 3.89 2.11 900 LightStereo-L(Ours) 6.4 6.4 17.51
EdgeStereo-V2[56] 1.83 1.07 1.84 3.30 2.08 320 LightStereo-H(Ours) 7.2 7.3 14.27
LEAStereo[7] 1.45 0.83 1.40 2.91 1.65 300
CREStereo[10] 1.46 0.90 1.45 2.86 1.69 410
ACVNet [12] 1.47 0.86 1.37 3.07 1.65 200
DispNetC[25] 4.11 2.77 4.32 4.41 4.34 60 from 0.7144 to 0.7036. However, this improvement comes
StereoNet [43] - - 4.30 7.45 4.83 22∗
DeepPrunerFast[34] - - 2.32 3.91 2.59 50 with an increase in parameters from 7.54M to 12.76M and
AANet[18] 2.42 1.46 1.99 5.39 2.55 62
DecNet[57] - - 2.07 3.87 2.37 50 a slight increase in FLOPs and inference time.
HITNet[45] 1.41 1.14 1.74 3.20 1.98 54
CoEx[39]
Fast-ACVNet[44]
1.55
1.68
1.15
1.23
1.79
1.82
3.82
3.93
2.13
2.17
33
39
Runtime Analysis. In, Table VI, LightStereo demonstrates
Fast-ACVNet+[44]
LightStereo-S (Ours)
1.45
1.88
1.06
1.30
1.70
2.00
3.53
3.80
2.01
2.30
45
17∗
commendable efficiency across its constituent modules. As
23∗
LightStereo-M (Ours)
LightStereo-L (Ours)
1.56
1.55
1.10
1.10
1.81
1.78
3.22
2.64
2.04
1.93 34∗
the number of blocks increases, the time required for cost ag-
LightStereo-H (Ours) 1.34 0.96 1.60 2.92 1.82 49∗
gregation in LightStereo-S, LightStereo-M, and LightStereo-
TABLE VI: Runtime breakdown. L also increases. For LightStereo-H, the time spent on feature
extraction is extended due to the adoption of a more complex
Feature Cost Disparity Total
Module Cost
Extraction Aggregation Regression Time EfficientNetV2. These runtime breakdowns underscore the
LightStereo-S
LightStereo-M
10.39
10.39
1.98
1.98
3.98
9.59
1.48
1.49
17.83
23.45
effectiveness of the LightStereo framework in achieving real-
LightStereo-L
LightStereo-H
10.39
27.17
1.98
1.98
23.64
23.64
1.49
1.49
37.50
54.28
time stereo-matching capabilities.

E. Comparison with State-of-the-art on Real Datasets


and balance between accuracy and computational cost.
Block Structure Analysis. In Table IV, we explore the KITTI benchmarks. We evaluate the results of our
impact of different block structures while maintaining a LightStereo variants on the KITTI 2012 and 2015 on-
constant expansion factor of 4. Configurations ’e’ to ’h’ line benchmarks. As illustrated in Table V, LightStereo-
explore blocks (1, 2, 4), (2, 4, 8), (4, 8, 16), and (8, 16, 32) S runs the fastest. Notably, LightStereo-H surpasses all
respectively. The results show that larger blocks generally other lightweight state-of-the-art stereo-matching networks
lead to better performance. For instance, configuration ’e’ across every evaluation metric on KITTI 2012. Additionally,
with the smallest block (1, 2, 4) has the highest EPE of LightStereo-H achieved the best results in terms of D1-bg
0.8317, while configuration ’h’ with the largest block (8, 16, and D1-all among lightweight models on KITTI 2015.
32) improves EPE, but detailed metrics are not provided. Cross-domain generalization. We evaluated the general-
There is a trade-off between accuracy and computational ization performance of our model on real-world datasets such
cost, as larger blocks increase FLOPs and parameters. as KITTI12 [26], KITTI15 [27], Middlebury [47], as shown
Expansion Factor Analysis. Still in Table IV, we analyze in Table VII. LightStereo shows superior generalization over
the effect of varying expansion factors while keeping the other lightweight methods, which indicates that our approach
block structure constant. As the expansion factor increases, a has moderate complexity to avoid overfitting.
consistent decrease in EPE is observed, indicating improved
accuracy. The analysis further proves that the critical im- V. C ONCLUSION
portance of disparity dimensions cannot be overstated in the This paper designs a lightweight 2D encoder-decoder
cost aggregation process of stereo matching. The disparity aggregation network to achieve precise and fast disparity
dimension is pivotal in effectively aggregating cost volumes, estimation, called LightStereo. While similar approaches
leading to more accurate and reliable depth estimations. have been explored previously, our novel contribution lies
However, this improvement in accuracy comes at the cost in optimizing performance through a targeted emphasis on
of significantly increased computational complexity and pa- the disparity channel dimension within the 3D cost volume,
rameters, with flops increasing from 26.23G to 93.34G and which encapsulates the distribution of matching costs. Our
the inference time increases from 22.44ms to 36.67ms. exhaustive exploration has led to the development of nu-
MSCA Module Analysis. As illustrated in Table IV, the merous strategies to enhance the capacity of this crucial
baseline configuration (i) with blocks (4, 8, 16) achieved dimension, ensuring both precision and efficiency in disparity
an EPE of 0.7144. Incorporating MSCA (configuration k) estimation. LightStereo offers a compelling solution for
reduced the EPE further to 0.6809 with minimal changes in accelerating the matching process while maintaining high
FLOPs and parameters. This suggests that MSCA provides levels of accuracy and efficiency.
the most significant improvement in accuracy with minimal Acknowledgements. This work was supported by the
impact on computational efficiency. We also explored the use National Natural Science Foundation of China under Grant
of the SE module [51]. Configurations ’i’ to ’j’ show that 62373356 and the Joint Funds of the National Natural
although the SE module improves accuracy, reducing EPE Science Foundation of China under U24B20162.
R EFERENCES [29] F. Tosi, L. Bartolomei, and M. Poggi, “A survey on deep
stereo matching in the twenties,” 2024. [Online]. Available:
[1] A. Kendall, H. Martirosyan, S. Dasgupta, P. Henry, R. Kennedy, [Link]
A. Bachrach, and A. Bry, “End-to-end learning of geometry and [30] G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encoding
context for deep stereo regression,” in ICCV, 2017. volume for stereo matching,” in CVPR, 2023.
[2] J.-R. Chang and Y.-S. Chen, “Pyramid stereo matching network,” in [31] X. Wang, G. Xu, H. Jia, and X. Yang, “Selective-stereo: Adaptive
CVPR, 2018. frequency information selection for stereo matching,” in CVPR, 2024.
[3] X. Guo, K. Yang, W. Yang, X. Wang, and H. Li, “Group-wise [32] P. Xu, Z. Xiang, C. Qiao, J. Fu, and T. Pu, “Adaptive multi-modal
correlation stereo network,” in CVPR, 2019. cross-entropy loss for stereo matching,” in CVPR, 2024.
[4] F. Zhang, V. Prisacariu, R. Yang, and P. H. Torr, “Ga-net: Guided [33] S. Khamis, S. Fanello, C. Rhemann, A. Kowdle, J. Valentin, and
aggregation net for end-to-end stereo matching,” in CVPR, 2019. S. Izadi, “Stereonet: Guided hierarchical refinement for real-time edge-
[5] F. Zhang, X. Qi, R. Yang, V. Prisacariu, B. Wah, and P. Torr, “Domain- aware depth prediction,” in ECCV, 2018.
invariant stereo matching networks,” in ECCV, 2020. [34] S. Duggal, S. Wang, W.-C. Ma, R. Hu, and R. Urtasun, “Deeppruner:
Learning efficient stereo matching via differentiable patchmatch,” in
[6] X. Gu, Z. Fan, S. Zhu, Z. Dai, F. Tan, and P. Tan, “Cascade cost
ICCV, 2019.
volume for high-resolution multi-view stereo and stereo matching,” in
[35] Y. Wang, Z. Lai, G. Huang, B. H. Wang, L. van der Maaten,
CVPR, 2020.
M. Campbell, and K. Q. Weinberger, “Anytime stereo image depth
[7] X. Cheng, Y. Zhong, M. Harandi, Y. Dai, X. Chang, H. Li, T. Drum-
estimation on mobile devices,” in ICRA, 2019.
mond, and Z. Ge, “Hierarchical neural architecture search for deep
[36] B. Xu, Y. Xu, X. Yang, W. Jia, and Y. Guo, “Bilateral grid learning
stereo matching,” in NeurIPS, 2020.
for stereo matching networks,” in CVPR, 2021.
[8] H. Wang, R. Fan, P. Cai, and M. Liu, “Pvstereo: Pyramid voting [37] A. Tonioni, F. Tosi, M. Poggi, S. Mattoccia, and L. D. Stefano, “Real-
module for end-to-end self-supervised stereo matching,” ICRA, 2021. time self-adaptive deep stereo,” in CVPR, 2019.
[9] X. Song, G. Yang, X. Zhu, H. Zhou, Z. Wang, and J. Shi, “Adastereo: [38] M. Poggi and F. Tosi, “Federated online adaptation for deep stereo,”
A simple and efficient approach for adaptive stereo matching,” in in CVPR, 2024.
CVPR, 2021. [39] A. Bangunharcana, J. W. Cho, S. Lee, I. S. Kweon, K.-S. Kim, and
[10] J. Li, P. Wang, P. Xiong, T. Cai, Z. Yan, L. Yang, J. Liu, H. Fan, S. Kim, “Correlate-and-excite: Real-time stereo matching via guided
and S. Liu, “Practical stereo matching via cascaded recurrent network cost volume excitation,” in IROS, 2021.
with adaptive correlation,” in CVPR, 2022. [40] X. Liu, H. Peng, N. Zheng, Y. Yang, H. Hu, and Y. Yuan, “Efficientvit:
[11] B. Liu, H. Yu, and Y. Long, “Local similarity pattern and cost self- Memory efficient vision transformer with cascaded group attention,”
reassembling for deep stereo matching networks,” in AAAI, 2022. in CVPR, 2023.
[12] G. Xu, J. Cheng, P. Guo, and X. Yang, “Attention concatenation [41] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei,
volume for accurate and efficient stereo matching,” in CVPR, 2022. “ImageNet: A large-scale hierarchical image database,” in CVPR,
[13] G.-Y. Nie, M.-M. Cheng, Y. Liu, Z. Liang, D.-P. Fan, Y. Liu, and 2009.
Y. Wang, “Multi-level context ultra-aggregation for stereo matching,” [42] M. Tan and Q. Le, “Efficientnetv2: Smaller models and faster training,”
in CVPR, 2019. in ICML, 2021.
[14] X. Guo, J. Lu, C. Zhang, Y. Wang, Y. Duan, T. Yang, Z. Zhu, and [43] S. Khamis, S. Fanello, C. Rhemann, A. Kowdle, J. Valentin, and
L. Chen, “Openstereo: A comprehensive benchmark for stereo match- S. Izadi, “Stereonet: Guided hierarchical refinement for real-time edge-
ing and strong baseline,” arXiv preprint arXiv:2312.00343, 2023. aware depth prediction,” in ECCV, 2018.
[15] Q. Wang, S. Shi, S. Zheng, K. Zhao, and X. Chu, “FADNet: A fast [44] G. Xu, Y. Wang, J. Cheng, J. Tang, and X. Yang, “Accurate and
and accurate network for disparity estimation,” in ICRA, 2020. efficient stereo matching via attention concatenation volume,” TPAMI,
[16] ——, “Fadnet++: Real-time and accurate disparity estimation with 2023.
configurable networks,” arXiv preprint arXiv:2110.02582, 2021. [45] V. Tankovich, C. Hane, Y. Zhang, A. Kowdle, S. Fanello, and
[17] F. Shamsafar, S. Woerz, R. Rahim, and A. Zell, “Mobilestereonet: S. Bouaziz, “Hitnet: Hierarchical iterative tile refinement network for
Towards lightweight deep networks for stereo matching,” in WACV, real-time stereo matching,” in CVPR, 2021.
2022. [46] X. Li, C. Zhang, W. Su, and W. Tao, “Iinet: Implicit intra-inter
[18] H. Xu and J. Zhang, “Aanet: Adaptive aggregation network for information fusion for real-time stereo matching,” in AAAI, 2024.
efficient stereo matching,” in CVPR, 2020. [47] D. Scharstein, H. Hirschmüller, Y. Kitajima, G. Krathwohl, N. Nešić,
[19] G. Xu, H. Zhou, and X. Yang, “Cgi-stereo: Accurate and real-time X. Wang, and P. Westling, “High-resolution stereo datasets with
stereo matching via context and geometry interaction,” arXiv preprint subpixel-accurate ground truth,” in GCPR, 2014.
arXiv:2301.02789, 2023. [48] A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan,
[20] M. M. Mijwil, R. Doshi, K. K. Hiran, O. J. Unogwu, and I. Bala, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al., “Searching for
“Mobilenetv1-based deep learning model for accurate brain tumor mobilenetv3,” in ICCV, 2019.
classification,” Mesopotamian Journal of Computer Science, 2023. [49] X. Ma, X. Dai, Y. Bai, Y. Wang, and Y. Fu, “Rewrite the stars,” CVPR,
[21] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, 2024.
“Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, [50] A. Wang, H. Chen, Z. Lin, H. Pu, and G. Ding, “Repvit: Revisiting
2018. mobile cnn from vit perspective,” CVPR, 2024.
[22] C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel matters– [51] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in
improve semantic segmentation by global convolutional network,” in CVPR, 2018.
CVPR, 2017. [52] B. Liu, H. Yu, and Y. Long, “Local similarity pattern and cost self-
[23] Q. Hou, L. Zhang, M.-M. Cheng, and J. Feng, “Strip pooling: reassembling for deep stereo matching networks,” in AAAI, 2022.
Rethinking spatial pooling for scene parsing,” in CVPR, 2020. [53] Z. Shen, Y. Dai, and Z. Rao, “Cfnet: Cascade and fused cost volume
for robust stereo matching,” arXiv preprint arXiv:2104.04314, 2021.
[24] M.-H. Guo, C.-Z. Lu, Q. Hou, Z. Liu, M.-M. Cheng, and S.-M.
[54] G. Yang, H. Zhao, J. Shi, Z. Deng, and J. Jia, “Segstereo: Exploiting
Hu, “Segnext: Rethinking convolutional attention design for semantic
semantic information for disparity estimation,” in ECCV, 2018.
segmentation,” NeurIPS, 2022.
[55] Z. Wu, X. Wu, X. Zhang, S. Wang, and L. Ju, “Semantic stereo
[25] N. Mayer, E. Ilg, P. Hausser, P. Fischer, D. Cremers, A. Dosovitskiy,
matching with pyramid cost volumes,” in ICCV, 2019.
and T. Brox, “A large dataset to train convolutional networks for
[56] X. Song, X. Zhao, L. Fang, H. Hu, and Y. Yu, “Edgestereo: An
disparity, optical flow, and scene flow estimation,” in CVPR, 2016.
effective multi-task learning network for stereo matching and edge
[26] A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous detection,” IJCV, 2020.
driving? the kitti vision benchmark suite,” in CVPR, 2012. [57] C. Yao, Y. Jia, H. Di, P. Li, and Y. Wu, “A decomposition model for
[27] M. Menze and A. Geiger, “Object scene flow for autonomous vehi- stereo matching,” in CVPR, 2021.
cles,” in CVPR, 2015.
[28] M. Poggi, F. Tosi, K. Batsos, P. Mordohai, and S. Mattoccia, “On the
synergies between machine learning and binocular stereo for depth
estimation from images: a survey,” IEEE Transactions on Pattern
Analysis and Machine Intelligence, 2021.

Common questions

Powered by AI

LightStereo achieves competitive computational efficiency through a combination of techniques. Its architecture includes MobileNetV2 for lightweight feature extraction, inverted residual blocks for cost aggregation, and a disparity regression method using soft-argmax for final output prediction. LightStereo-S, the most efficient configuration, boasts only 22.71 Gflops, with an inference time of 17ms, demonstrating a remarkable balance between speed and resource requirements while maintaining a competitive End-Point-Error (EPE).

MobileNetV2 is employed within LightStereo for its efficient convolutional architecture capable of extracting high-quality feature maps at multiple scales with reduced computational complexity, which is crucial for lightweight models. It is preferred over other models due to its superior balance of accuracy, speed, and resource efficiency, as evidenced by its satisfactory EPE and FLOPs in comparisons. This balance is essential for enabling effective real-time stereo vision applications without heavy computational demands .

LightStereo overcomes MobileNetStereo-2D's limitations by integrating inverted residual blocks and efficient feature extraction techniques like multi-scale convolutional attention, optimizing both the accuracy and efficiency of disparity estimation. While MobileNetStereo-2D suffered from poor performance, LightStereo's architecture focuses on disparity channels and utilizes advanced residual mechanisms, facilitating better feature learning at reduced FLOPs. It offers substantial improvements in real-time application capabilities, evidenced by its superior EPE and computational benchmarks .

Focusing on the disparity channel dimensions allows LightStereo to directly model disparities, thus effectively capturing the correspondence between stereo image pairs' points. This focus aids in maintaining essential disparity information and improving the robustness against spatial convolutions' inaccuracies that result from focusing solely on height and width. This approach minimizes error propagation and allows for more accurate and efficient stereo depth estimations, addressing common hurdles in traditional stereo vision methods .

The LightStereo model addresses two key challenges in stereo matching: the trade-off between computational efficiency and accuracy, and the propagation of errors at disparity discontinuities. LightStereo employs inverted residual blocks to enhance computational efficiency while maintaining disparity estimation accuracy, enabling real-time application potential. It incorporates a Multi-Scale Convolutional Attention Module inspired by image segmentation techniques to improve cost aggregation by leveraging semantic features from input images. The use of disparity channel dimensions instead of height and width further enhances the robustness of stereo matching .

LightStereo-S exhibits significantly higher efficiency with a computing requirement of only 22.71 Gflops, setting a fast benchmark at 17ms of inference time, notably outperforming many state-of-the-art methods in terms of speed. Moreover, LightStereo-H achieves superior accuracy with an EPE of 0.51, surpassing models like Fast-ACVNet+ (EPE of 0.59) while remaining competitive in computational complexity. These benchmarks reinforce LightStereo's position as a feasible solution for balancing efficiency and accuracy in real-time stereo matching .

Inverted residual blocks are significant in LightStereo's cost aggregation as they facilitate learning richer features at a reduced computational cost. These blocks expand channel dimensions temporarily to capture more detailed information before compressing them, thereby enabling accurate modeling of the disparity channel dimension rather than just the spatial dimensions. This design optimizes the accuracy of disparity estimation without a trade-off in processing speed, which is essential for real-time stereo applications .

The MSCA module enhances cost aggregation by exploiting semantic information from the left image to effectively excite the disparity channel dimension of the cost volume. This approach utilizes multi-scale feature extraction, which excites crucial semantic details such as object-level information, to guide the aggregation process more precisely and handle disparity discontinuities more robustly. The MSCA's design supports improved feature learning at reduced computational overhead, contributing to both accuracy and speed .

LightStereo employs the smooth L1 loss function for training, which balances between L1 and L2 loss, making it robust to outliers while maintaining sensitivity to small disparity errors. This loss is calculated as the average of disparity prediction errors across all labeled pixels. It is particularly suitable given stereo matching tasks' inherent challenges—ensuring a balance between computational stability and sensitivity required for precise depth recovery, especially in regions with subtle disparity variations .

Disparity regression in LightStereo is implemented using a soft-argmax approach where the final disparity map is predicted by summing each potential disparity value weighted by its predicted probability. The advantage of this approach is the precise calculation of continuous disparity values instead of discrete estimates, enhancing the overall prediction accuracy and handling of sub-pixel disparities, which is vital for accurate depth estimation in stereo matching .

You might also like