0% found this document useful (0 votes)
3 views13 pages

2025 Rss Semrob.github.io Paper5

RayFronts is a novel real-time semantic mapping system designed for open-set semantic mapping, enabling robots to efficiently understand and explore environments both within and beyond their depth perception range. It integrates dense mapping with ray-based representations, significantly reducing search volumes and improving localization capabilities for distant entities. The system operates at high efficiency, achieving state-of-the-art performance in zero-shot 3D semantic segmentation while maintaining real-time processing capabilities.

Uploaded by

francissmith72
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views13 pages

2025 Rss Semrob.github.io Paper5

RayFronts is a novel real-time semantic mapping system designed for open-set semantic mapping, enabling robots to efficiently understand and explore environments both within and beyond their depth perception range. It integrates dense mapping with ray-based representations, significantly reducing search volumes and improving localization capabilities for distant entities. The system operates at high efficiency, achieving state-of-the-art performance in zero-shot 3D semantic segmentation while maintaining real-time processing capabilities.

Uploaded by

francissmith72
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

RayFronts: Open-Set Semantic Ray Frontiers for

Online Scene Understanding and Exploration


Omar Alama, Avigyan Bhattacharya, Haoyang He, Seungchan Kim, Yuheng Qiu, Wenshan Wang,
Cherie Ho, Nikhil Keetha, Sebastian Scherer
Carnegie Mellon University, Pittsburgh, Pennsylvania 15213
Email: {oalama, avigyanb, hhe2, seungch2, yuhengq, wenshanw, cherieh, nkeetha, basti}@[Link]

Mission: Find (1) Significant Search Volume (2) Online Mapping:


Text Query
Reduction with Open-Set
Red Building Pushing
Semantic Ray Fronts RayFronts

Image
Query Water Tower

Radio Tower
Metal Stairs

Time
Red Building
Road Cracks

Dense Green Canopy

(3) Fine-Grained
Open-Set Querying Wooden Picnic Tables To
Voxels

Fig. 1: RayFronts is a real-time semantic mapping system that enables fine-grained scene understanding both within and beyond
the depth perception range. Given an example mission through multi-modal queries to locate red buildings & a water tower, RayFronts
enables: (1) Significant search volume reduction for online exploration (as shown by the red and blue cones at the top) and localization
of far-away entities (e.g., the water & radio tower). (2) Online semantic mapping, where prior semantic ray frontiers evolve into semantic
voxels as entities enter the depth perception range (e.g., the red buildings query on the right side). (3) Multi-objective fine-grained open-set
querying supporting various open-set prompts such as “Road Cracks”, “Metal Stairs”, and “Green Dense Canopy”.

Abstract—Open-set semantic mapping is crucial for open- I. I NTRODUCTION


world robots. Current mapping approaches either are limited by
the depth range or only map beyond-range entities in constrained Open-set semantic mapping is essential for robotic systems
settings, where overall they fail to combine within-range and
to reason, search, and navigate in open-world environments.
beyond-range observations. Furthermore, these methods make
a trade-off between fine-grained semantics and efficiency. We The task requires capturing both fine-grained local details and
introduce RayFronts, a unified representation that enables both distant beyond-range semantic cues in real-time. For instance,
dense and beyond-range efficient semantic mapping. RayFronts as shown in Fig. 1, an aerial or ground robot may need to
encodes task-agnostic open-set semantics to both in-range voxels localize the water or radio towers over 100 meters beyond
and beyond-range rays encoded at map boundaries, empowering
its depth perception capability, as well as locate any hazards
the robot to reduce search volumes significantly and make
informed decisions both within & beyond sensory range, while (road cracks) or interesting structures along the way (red
running at 8.84 Hz on an Orin AGX. Benchmarking the within- building). This work explores what the most effective semantic
range semantics shows that RayFronts’s fine-grained image mapping system would be to capture this information and
encoding provides 1.34⇥ zero-shot 3D semantic segmentation enable reasoning within and beyond depth sensing limitations.
performance while improving throughput by 16.5⇥. Tradition-
ally, online mapping performance is entangled with other system Although there is a growing body of literature on open-set
components and planners, complicating evaluation. We propose metric semantic mapping [3, 14, 32, 8, 24], these methods
a planner-agnostic evaluation framework that captures the utility focus primarily on offline mapping for downstream usage in
for online beyond-range search and exploration on prerecorded limited environments ignoring efficiency and depth-sensing
mapping runs, and show RayFronts reduces search volume limitations. Such representations cannot guide the robot in
2.2⇥ more efficiently than the closest online baselines.
search and exploration tasks as they provide no information
about the unmapped region. Other works change the way
semantics are typically encoded in a map (point clouds, abstract textual concepts and images. These Visual Language
voxels, and bounding boxes) to representations that can guide Models (VLMs) initially aligned textual descriptions and im-
exploration (i.e semantic frontiers [36, 4], and semantic poses ages as a whole and not to particular regions or pixels. Subse-
[30]). However, existing semantic frontier maps are limited quently, follow-up work based on supervised and unsupervised
to 2D indoor environments and have limited semantics due regimes has attempted to address this issue [33]. A recurring
to whole-image encoding [36] or using closed-set models [4], theme in these methods is the trade-off between efficiency and
whereas semantic poses [30] lack fine-grained reconstructions accuracy, with the most performant approaches often using
and can only recognize prominent objects in an image. multiple foundation models like DINOv2 [23], Grounding
In this context, we explore the question, “How to design DINO [19], and SAM [18]. This is not optimal for online
an efficient online mapping representation that facilitates real-world deployment and hence we explore the applicability
fine-grained scene understanding, and be aware of beyond- of RADIO [26], a foundation model aligned with various
range semantic entities?” We introduce RayFronts, a dense visual foundation models. While RADIO’s language
semantic map representation which seamlessly integrates tra- alignment is to the image as a whole, we find that employing a
ditional within-depth mapping with ray-based representations, simple attention trick [10] with its SIGLIP adapter enables us
facilitating both dense mapping within observed depth ranges to achieve state-of-the-art pixel-level language alignment and
and perception beyond them. Unlike conventional represen- real-time performance on embedded hardware.
tations truncated at the depth range, multi-directional seman-
tic ray frontiers retain coarse-grained far-range information, B. Offline & Bounded Open-Set Semantic Mapping
enabling downstream planners (e.g., object-search) to reduce
their search volume significantly. Additionally, to assess the Traditional semantic mapping systems have relied on
utility of the proposed representation, we construct a planner- learning-based methods to detect and segment a fixed set of
agnostic benchmark and propose a new metric to measure concepts, with performance limited by vocabulary size and
how effectively an online mapping strategy reduces the search training distribution [20, 27, 7, 21, 35, 13, 11]. With the rise of
space for fast object localization and exploration. Finally, to dense 2D open-set semantics, interest has grown in open-vocab
avoid single vector image encoding and expensive pipelines, semantic mapping systems using representations like point
we introduce a language-aligned and spatially-dense encoder clouds, voxels, and scene graphs [12]. These systems have
that achieves state-of-the-art performance on zero-shot 3D shown strong open-world capabilities for navigation, manipu-
semantic segmentation enabling a computationally efficient, lation, and scene understanding [3, 14, 32, 8, 24, 16, 34, 15].
open-world, and deployable 3D online mapping system. However, most focus on offline database maps and lack online
Our key contributions are as follows: utility for robotics, with many design choices making real-
C1: Unified 3D Map Representation for Within-Depth and time deployment infeasible. To address this, we introduce a
Beyond-Depth Perception: We develop the first-of-its-kind computationally efficient, fast, and deployable 3D mapping
open-set semantic ray frontier 3D map, which enables robots system for online scene understanding.
to reason in open environments achieving up to 1.85x mIoU
in offline zero-shot performance, and are 2.2x more efficient C. Online & Unbounded Open-Set Semantic Mapping
in reducing search volume in online mapping than the closest
offline & online baselines respectively. While offline and bounded semantic mapping has excelled
C2: Planner-Agnostic Online Semantic Mapping Evalua- in indoor scenes, it struggles with outdoor, unbounded, and
tion Framework: We showcase that online semantic mapping unstructured environments, where limited depth perception
systems can be evaluated on their fundamental utility for becomes a challenge. An effective online semantic mapping
exploration, without being tightly coupled with a planner, by system must support both efficient exploration and fine-grained
developing a metric that assesses “correctly reduced search scene understanding. VLFM [36] addresses this by encoding
volume”. semantics on 2D frontiers for object goal navigation, but it
C3: Efficient real-time open-set online mapping system: can is limited to a single object at a time and only works in
run end to end at 8.84 Hz on an ORIN AGX and our efficient indoor settings. Similarly, Embedding Pose Graph (EPG) [30]
dense vision-language encoder is 16.5x faster than the closest encodes semantics into rays from pose nodes, but lacks fine-
baseline and achieves state-of-the-art on open-vocab zero-shot grained mapping and condenses the entire image into one
3D semantic segmentation mIoU. feature vector, risking the loss of subtle details.
In contrast, we propose a novel representation combining
II. R ELATED W ORK
semantic voxels with ray-based frontiers, capturing multiple
A. Dense 2D Open-Set Semantics viewing directions and open-set features. This approach en-
The rapid rise of foundation models [2] has spearheaded ables efficient online search and rough triangulation of distant
progress in tasks requiring fine-grained open-set concepts objects, allowing us to capture both in-range and beyond-range
which are hard to capture with a fixed taxonomy of semantic semantic entities. Our synergy of metric-map-based semantic
classes [23, 18]. CLIP [25] and its subsequent variants such voxels and direction-based ray frontiers supports fine-grained
as SIGLIP [37] have shown impressive alignment between scene understanding and efficient exploration.
High
Dense Language VDB Occupancy Map Semantic Voxels
Aligned Features VDB Occupancy Map Ot Semantic Voxel Map Vt
Image
Query

Cosine Similarity
Empty Occupied First 3 PCA Features
First 3 PCA Features
RayFronts Text Query Text Query
Frontier Map Ft Semantic Ray Fronts Rt Thick Forest Trash Container
Posed RGBD Frontier Map Semantic Ray Fronts

Low
Frontiers First 3 PCA Features

45#+61*0"3--2$/%072+8'/-9$1*01,
!"#$% &'()*+"%,-."/0"1-2'##0"3 )0"'/-2'# :0%;0"-'"8-<1(+"8-&'"31-

Fig. 2: Overview of our online mapping system, RayFronts is designed for multi-objective & multi-modal open-set querying of both
in-range and beyond-range semantic entities. Given posed RGB-D images, we first extract dense features with our fast language-aligned
image encoder. Then, posed depth information and features are used to construct a semantic voxel map for in-range queries. In parallel,
RayFronts also maintains a VDB-based occupancy map to generate frontiers, which are further associated with multi-directional semantic
rays. These semantic ray fronts enable us to perform beyond-range querying of open-set concepts in the unobserved region.

III. M ETHOD vanilla ViT [6], it struggles with fine-grained localization


We present RayFronts, a unified 3D semantic mapping of visual features, a critical challenge in semantic scene
system for multi-modal open-set semantic querying of both in- understanding. To address this, we integrate the explicit spatial
range and beyond-range semantic entities. RayFronts main- attention mechanism proposed by NACLIP [10] and modify
tains a semantic voxel map Vt containing voxel coordinates the RADIO encoder accordingly. Specifically, we augment the
and semantic features for within-range entities, an occupancy attention layer of the final ViT block by introducing a locality
VDB map Ot , a set of frontiers Ft denoting subsampled constraint via an unnormalized multivariate Gaussian kernel
boundary voxels between observed and unobserved spaces, centered around each patch essentially pushing the model to
and semantic ray fronts Rt , a ray-based representation on the attend to its neighboring patches and improving locality.
frontiers, which contains features for beyond-range semantic To densely align the RADIO feature space with language,
reasoning. we explore the available pre-trained MLP-based adaptor heads
RayFronts operates in four steps: (1) extracting dense, provided by RADIO. Simply following RADIO’s original
language-aligned features from RGB input through our ef- distillation approach– projecting spatial features onto CLIP
ficient encoding pipeline, (2) fusing within-range featurized or SIGLIP space using their respective adapters–yields subpar
points into a sparse semantic voxel map, (3) maintaining performance. Instead, we use the SIGLIP summary feature
an occupancy map for frontier computation and semantic adapter to project spatial features to the SIGLIP CLS token
voxel pruning, and (4) ray casting beyond-range semantics space, thus resulting in a spatially consistent and language-
onto frontiers to semantically reason beyond the observed aligned feature map and observe significant performance im-
map. Our system is optimized for parallel computing and provements over existing methods.
online mapping, leveraging PyTorch tensors for Vt , Ft , Rt
B. Semantic Voxels for Dense Within-Depth Queries
on the GPU and OpenVDB [22] for Ot on the CPU. This
design ensures efficient querying, seamless feature integration, Given a pose Pt 2 SE3 and depth map Dt 2 RH⇥W , we
and adaptability to evolving environments. The pipeline and initialize a voxel grid retaining only points within the frustrum,
outputs of RayFronts are illustrated in Fig. 2. yielding Qt . We transform the points Qt into the camera
frame and classify its occupancy based on depth Dt . For each
A. Extracting Dense Language-Aligned Features occupied point, we find the associated feature via nearest-
There has been a rapid growth of methods that extract neighbor interpolation yielding Ptlocal = {(pi , fi )}M i=1 where
dense language aligned features from RGB images. However, pi 2 R3 are point coordinates and fi 2 R3+D+1 (3 for RGB, D
existing methods fall short by (1) lacking generalization due to is feature dimension, and 1 for the hit count). Local updates are
limited supervision, (2) sacrificing efficiency with multi-model accumulated into a buffer of m frames before being voxelized
multi-stage pipelines, or (3) prioritizing efficiency and gener- at resolution a and integrated into the global map V .
alization at the cost of segmentation quality. In this work, we Feature Fusion and Aggregation: Rather than complex
adopt RADIO[26], a vision foundation model that distills key fusion methods used in [14, 32], we employ a simple weighted
features from CLIP[25], DINOv2[23], and SAM[18] yielding average, where each voxel’s hit count serves as the weight
a richer representation. However, since RADIO leverages when fusing features within the same voxel. To achieve this,
we concatenate coordinate and feature tensors of accumulated semantic leakage at object boundaries, and used to select the
local updates Ptlocal with those of global voxel map Vt . A semantic pixels to propagate as rays Rtlocal = {(or , dr , fr )}H t
r=0
parallel scatter-reduce operation fuses features at the same 3
where or 2 R is the ray and camera origin, dr 2 R is 3

discretized coordinates into a single voxel. the normalized direction vector, and fr represents semantic
features.
C. Occupancy Mapping for Frontiers and Pruning Associate (Matching Rays to Frontiers): In the presence
To represent occupancy map Ot efficiently, we employ of depth information, rather than keeping rays at the robot’s
OpenVDB [22], recently used in modern 3D frontier-based origin as in [30], we leverage the mapped area to push rays
exploration works [1, 9, 17] for its sparse tree representation closer to their observed entities, improving localization. For
and multi-resolution capability. Following standard practice, each semantic ray (or , dr , fr ), we select a frontier from the
we store log-odds occupancy o j in a signed byte. To better candidate set Ft+1 through a two-step filtering process. First,
tolerate dynamic environments and to avoid overflow, we we prune frontiers by (1) removing those not in front of the ray
limit probocc (o j ) to lower and upper limits. Fig. 2 shows (2) computing the shortest orthogonal distance dortho between
the OpenVDB map with large free voxels showing the multi- the ray and frontiers, discarding those where dortho > b
resolution aspect of the occupancy representation. (exceeding the frontier grid cell size), and (3) calculating the
Pruning Semantic Voxels: When accumulating voxels over distance from ray origin or to frontier origin p, obtaining dorig
long distances and time periods, odometry drifts and dynamic and removing frontiers where dorig > 4 ⇥ depth range.
objects can introduce inconsistencies, not to mention the Next, for the remaining k candidate frontiers, we compute
growing memory consumption. To mitigate this, we prune a cost function
invalid semantic voxels by querying the occupancy map Ot !
and removing those with occupancy below 0.5. dortho dorig
dcost = + /2, dcost 2 [0, 1]
r
max({dortho }kr=0 ) max({dorig
r }k )
r=0
D. Finding the “Fronts”: Computing 3D Frontiers (1)
We identify frontiers by iterating over all free observed We select the frontier with the minimum dcost as the best
voxels using efficient OpenVDB iterators and examining their match. We qualitatively find that utilizing both dortho and dorig
neighbors. A voxel is considered a frontier if its neighbors improves results and prevents distant frontiers from receiving
meet the minimum thresholds for unobserved (minunobsrv ), noisy semantics.
occupied (minocc ), and free (min f ree ) counts, allowing us to For further refinement, we optionally apply ray tracing,
emphasize frontiers near surfaces or open space as needed. To marching each ray through the occupancy map Ot+1 until
reduce density, we subsample the frontier map using a coarser it reaches its assigned frontier or encounters occupied or
voxel grid of size b . Fine-grid frontiers are accumulated into unobserved (possibly occupied) cells. At this stage, each
a coarser grid, and cells with enough frontiers remain as semantic ray is associated with a frontier, updating its origin
frontiers. (or ) to the corresponding frontier origin p. Since we lack depth
information about the underlying semantic entity, we maintain
E. Semantic Ray Frontiers for Beyond-Depth Mapping the ray’s direction dr when shifting its origin.
Need for richer frontiers: Existing semantic frontier meth- Discretize and Accumulate (“Ray Binning”): Similar
ods have fundamentally constrained beyond-range semantic to voxelization techniques, we organize semantic rays into
encoding, where only a single object can be pursued at a angle bins with a resolution of y degrees. The normalized
time due to feature collisions from distinct objects observed ray directions dr are converted to spherical angles using:
through the same frontier. To enable multi-object semantic qr = atan2(dr1 , dr0 ), fr = acos(dr2 ) where atan2 is the four-
guidance for search and exploration, we transition from con- quadrant inverse tangent. We then discretize these angles and
ventional semantic frontiers Fsem = {(pk , fk )}Fk=1 to semantic merge rays that correspond to the same frontier and the same
ray frontiers Rsem = {(or , qr , fr , fr )}Rr=1 , where or is ray angle bin from both the local update Rtlocal and the global set
origin, qr 2 [ p, p) and fr 2 [0, p) are azimuthal and zenith Rt . We use 1 dcost for weighing the features while merging,
angles, and fr are semantic features. This shift drastically assigning lower trust to high-cost associations. This yields the
enhances the mapping system by allowing efficient storage of updated ray-frontier Rt+1 .
rich multi-object semantics with minimal feature collisions, Pushing the ray fronts onward: Semantic ray frontiers
enabling rough triangulation of object locations, and reducing must be updated as new areas are mapped. After computing
the search space volume needed for exploration. We discuss the frontier update Ft+1 , we use a voxel grid to perform a set
the ray mapping process (observe, associate, discretize & intersection between all ray origins and frontier origins, similar
accumulate) and how rays are pruned and propagated below. to pruning voxels, and remove rays no longer associated
Observe: To identify out-of-range regions in the feature with active frontiers. If ray tracing is enabled, removed rays
map Ft we compute a boolean mask Mt 2 RH⇥W from the are added back to the ray accumulation buffer to be re-
depth Dt (obtained via stereo, LiDAR, or monocular depth es- cast in the next iteration. This preserves previously observed
timation). The mask encompasses either +• values from depth semantics that may no longer be directly visible as the frontier
sensors or far low-certainty values. Mt is eroded to prevent shifts (e.g., from a side view). However, without ray tracing,
Scene Bound • Semantic Poses (Sem Pose): Emulates EPG [30] by using
Unmapped Region
Mapped global encoding for the image, resulting in a single ray
Search volume
Region per frame located at the robot origin.
• Semantic Voxels (Sem Voxels): Subsumes representa-
Traditional
Object of
Interest
tions that encode only within-range semantics [32, 14, 8].
Semantic • Semantic Frontiers: Emulates in 3D the 2D approaches
Segmentation different semantic
mapping capability that paint frontiers with semantics [36, 4]. We recog-
nize that there are two ways semantic frontiers can be
Search Volume Recall Search Cut Volume (SCV) interpreted; (1) As a spherical region encompassing the
semantic entity (i.e Spherical Sem Fronts), or (2) as a

* =
Search Cut
Volume Recall single ray pointing away from the observed region (i.e
TP Added to allow (SCVR) Unidirectional Sem Fronts). We evaluate both.
SCV to be 1 for
perfect prediction We define search volume as the unmapped region unless
Fig. 3: An illustration of our proposed planner-agnostic metric further evidence is provided. Ray-based approaches cast search
(Search Cut Volume Recall) for open-world online search bench- cones while Spherical Sem Fronts define a sphere volume
marking. Intuitively, the metric captures “How much of the search extending to the nearest frontier. Multiple search volumes
volume is eliminated correctly?” An optimal mapper should promptly are summed in a voxel grid, counts are normalized, and
and accurately reduce the search space, enabling fast multi-object
thresholded at 0.05 to get the final search volume for a
localization and exploration.
class. For Unidirectional Sem Fronts, frontier directions are
inferred using the occupancy map Ot by computing a weighted
combination of all directions around a frontier in a 3x3x3
removed rays would continue to propagate indefinitely, so we
window where mapped voxels have a weight of -1 (Pushing
disable the behavior under that setting. In our experiments, we
away) and unmapped voxels have a weight of +1 (Pulling
use ray tracing unless otherwise stated.
toward).
Evaluation Protocol: We ask the question “Can an online
IV. E XPERIMENTAL S ETUP semantic mapping system’s utility for search and exploration
A good online mapping system should (1) intelligently be assessed independently of specific planners ?” Yes, the
guide the robot toward regions of interest in any environment, key is examining how accurately and efficiently the map
eliminating irrelevant volumes early, (2) accurately capture constrains the search space. Traditional mapping metrics such
fine-grained open-set semantics within a metric map, and (3) as mIoU, mAcc, F1 measure fine-grained semantic localization
do so efficiently. In this section, we first introduce our pro- but overlook search volume efficiency, as they ignore true
posed online mapping evaluation framework, which assesses negatives. In search and exploration, a high true negative
the utility of a mapping system in guiding exploration with- rate in the unobserved region reduces wasted search time.
out a planner in-the-loop, and introduce competitive baseline Therefore, for beyond-range search volume estimation, we
representations. We then present our extensive offline map introduce a novel metric below, and to evaluate within-range
evaluation following established protocols. Finally, we conduct fine-grained online performance, we use the area under the
a deployability and throughput analysis. mIoU-time curve.
Search Cut Volume Recall Metric: Our proposed metric,
A. Planner-Agnostic Online Semantic Mapping Evaluation shown in Figure 3, measures how accurately and efficiently a
mapping system cuts search volume. The intuitive definition is
Dataset: Originally designed to challenge visual SLAM to compute total unmapped volume volunmapped and subtract
with large, cluttered, long-tail objects in indoor and the search volume from it, however to avoid punishing true
outdoor environments, TartanAirV2[31] serves as a stress positives, we define search cut volume (SCV) as:
test for our representation. To simulate scenarios with
FPunmapped
severely limited depth, we choose four large outdoor scenes SCV = 1 , SCV 2 [0, 1] (2)
AbandonedCableday, Factory, Downtown and volunmapped
ConstructionSiteOvercast where bounding boxes To temper the metric against incorrectly cutting down volume,
span approximately 8 million m3 with a 50m range cutoff. We we compute Recall in the unmapped region and multiply it
generate ground truth occupancy (defining the scene volume) with SCV yielding the Search Cut Volume Recall (SCVR)
and a semantic label map at 1-meter voxel resolution from metric:
the provided posed RGBD input. T Punmapped
Baselines: There are no established online mapping baselines SCV R = SCV ⇤ , SCV R 2 [0, 1]
FNunmapped + T Punmapped
for 3D open-world environments. Therefore, we take inspira- (3)
tion from existing works and design the following baselines, The SCVR metric is robust to both naive cases (1) not
keeping the encoder fixed to isolate the impact of our mapping constraining the search volume, or (2) constraining it to 0
approach: volume, yielding 0 for both. For an aggregate, we compute
TABLE I: Online & Unbounded Semantic Mapping Benchmarking on TartanAirV2 [31]. Ranking shown as first , second , and third .
0m Depth (AUC) " 10m Depth (AUC) " 20m Depth (AUC) "
Methods mIoU(%) SCV(%) Recall(%) SCVR(%) mIoU(%) SCV(%) Recall(%) SCVR(%) mIoU(%) SCV(%) Recall(%) SCVR(%)
Sem Poses 0.00 11.37 91.91 4.02 – – – – – – – –
Sem Voxels – – – – 20.49 0.00 100.00 0.00 13.03 0.00 100.00 0.00
Spherical Sem Fronts – – – – 20.49 18.00 82.35 0.40 13.03 13.33 87.02 0.41
Unidirectional Sem Fronts – – – – 20.49 16.12 85.93 3.15 13.03 11.58 89.54 2.07
RayFronts (Ours) 0.00 36.59 75.37 16.27 20.49 22.94 81.15 7.08 13.03 14.32 88.69 4.56

the area under the SCVR-time curve, stopping time for each ranges. Sem Poses fails to capture any fine-grained reconstruc-
class when 50% of it has entered the mapped region. To further tions scoring 0 mIoU-AUC, while Sem Voxels fails to provide
assess the robustness of RayFronts, we vary depth sensing any information about the unmapped region scoring 0 SCVR-
range at 0m, 10m, and 20m. AUC. At 0 depth range, RayFronts attaches dense semantic
rays at each pose as opposed to the global encoding scheme
B. Offline 3D Open-Vocabulary Semantic Segmentation employed by Sem Poses. This allows us to encode non-
Datasets: We follow prior work [14, 8, 32] and evalu- prominent objects seamlessly and results in a ⇠ 4⇥ SCVR-
ate on Replica (office[0-4], room[0-2]) and ScanNet AUC than Sem Poses. This observation is illustrated in Fig. 4
(scene[0011,0050,0231,0378,0518]). In line with where for a simple prominent object such as “building”, both
previous protocols, we report results while ignoring back- Sem Poses and RayFronts perform similarly. However,
ground classes (“floor”, “wall”, “ceiling”, “door”, “window”). for a more distant object like “chimney”, Sem Poses fails
However, we additionally evaluate across all classes to demon- to capture its semantics entirely. At higher depth ranges, we
strate our ability to handle background seamlessly. More- observe that RayFronts consistently outperforms semantic
over, to showcase RayFronts’s effectiveness in outdoor, frontier baselines at ⇠ 2.2⇥ the SCVR-AUC. We attribute this
unstructured, “in-the-wild” environments, we further evaluate to (1) less semantic collisions as distinct objects are unlikely to
on the TartanAirV2 [31] scenes referenced in IV-A excluding be fused in the same ray unlike semantic frontiers which can
methods that cannot function outdoors. have many collisions, (2) better preservation of the angle at
Baselines: We compare our method with two categories of which the semantic entity was observed from, and (3) allowing
approaches: (1) vision-language representations that create 3D each frontier to have multiple rays attached, increasing the
semantic maps, namely, ConceptFusion [14], ConceptGraphs density of beyond-range semantics. RayFronts is superior
[8], and HOV-SG[32]; and (2) zero-shot semantic segmen- to all baselines across depth ranges empowering both fine-
tation encoders, namely, NACLIP[10] and Trident[28]. We grained localization and beyond-range guidance.
extend the latter encoder-based methods to 3D using the same B. Offline 3D Semantic Segmentation
projection and fusion method as our system.
Evaluation Protocol: We follow standard open-vocabulary Table II provides a detailed comparison of the performance
semantic segmentation evaluation protocols. We generate 3D between our framework and other zero-shot approaches, out-
segmentations by running HOV-SG and ConceptGraph code lined in Section IV-B. RayFronts consistently outperforms
ensuring an accurate representation of their scene graph the baselines in mIoU, and achieves SOTA performance
method. For all others, we generate segmentations by com- beating the next best baselines by +18.07% and +9.63% mIoU
puting the cosine similarity between the embedded feature and on Replica and Scannet, respectively, excluding background.
the class-name text embedding, making a voxel prediction if RayFronts is also able to handle background seamlessly
its softmax probability exceeds 0.1. We encode class names with its single-forward pass approach while segment-and-
using each method’s specified templates. For our approach, encode approaches fall short.
we follow NACLIP and use 80 templates [10], with a prompt For outdoor in-the-wild performance on TartanAirV2, Ta-
denoising [38] threhsold of 0.5 to suppress irrelevant classes. ble III shows that RayFronts exceeds the performance of
We also apply k-NN matching (k=5) following HOV-SG [32] the baselines by 3.36% mIoU. While Trident-3D serves as a
protocol, assigning each GT voxel the majority label. All close second to our approach and achieves a slightly higher
baselines use the ViT-L model architecture for consistency. f-mIoU on TartanAirV2 by a marginal 0.13%, it does so at
We resize images to 480x640, apply a frame skip of 10, 5cm the cost of integrating multiple foundational models into their
voxels for Replica and ScanNet and 1m voxels for TartanAir. pipeline, which significantly reduces efficiency—an essential
factor for online semantic mapping.
V. R ESULTS & D ISCUSSION
C. Encoder & Mapping Throughput Analysis
A. Online Semantic Mapping To assess deployability, we run RayFronts on an NVIDIA
Table I summarizes online performance of the five meth- Jetson AGX Orin and perform a quantitative comparison of
ods in their respective operating ranges. We observe that image encoder throughput shown in Fig. 5. Our mapping
RayFronts excels and is the upper bound across depth system achieves SOTA performance in 3D open-set semantic
20m Depth 0m Depth
RayFronts Spherical Sem Fronts Unidirectional Sem Fronts RayFronts Sem Poses

Building
Chimney

Fig. 4: RayFronts consistently surpasses baselines for online semantic mapping. Two query scenarios are shown: (1) querying for a
prominent object (i.e Building) that enters depth range, and (2) a distant object (i.e Chimney) that remains beyond range. Through unified
dense voxel mapping, and beyond-range semantic ray frontiers, RayFronts sets the upper-bound in both scenarios.

TABLE II: Offline 3D Semantic Segmentation Benchmarking on Indoor Datasets.


Replica [29] ScanNet [5]
Without Background With Background Without Background With Background
Methods mIoU(%) f-mIoU(%) Acc(%) mIoU(%) f-mIoU(%) Acc(%) mIoU(%) f-mIoU(%) Acc(%) mIoU(%) f-mIoU(%) Acc(%)
ConceptFusion [14] 21.07 31.51 35.65 20.38 35.75 41.58 21.76 26.71 34.13 18.57 23.06 28.77
ConceptGraphs [8] 11.63 16.61 19.80 11.72 21.35 28.28 21.62 24.32 31.04 20.83 23.61 35.80
HOV-SG [32] 16.93 31.45 34.74 19.29 30.64 35.17 26.79 36.05 44.17 23.48 28.92 38.52
NACLIP-3D [10] 20.37 35.08 47.47 15.30 16.98 26.23 31.66 39.03 51.65 22.32 24.32 33.46
Trident-3D [28] 21.30 43.34 54.79 20.63 38.53 50.31 29.97 37.62 51.06 24.80 27.77 38.43
RayFronts (Ours) 39.37 62.03 68.80 27.73 43.37 54.45 41.29 46.42 56.76 32.29 39.04 49.15

TABLE III: Offline 3D Semantic Segmentation Benchmarking on an


Outdoor Dataset (TartanAirV2 [31]).

Methods mIoU (%) f-mIoU (%) Acc (%)


ConceptFusion [14] 5.84 32.78 39.76
NACLIP-3D [10] 9.66 40.82 54.10
Trident-3D [28] 9.86 43.56 55.34
RayFronts (Ours) 13.22 43.43 57.26

segmentation with 1.34x the mIoU of Trident while being


16.5x faster, running at 17.5 Hz, and with only 46% of the
parameters. While NACLIP has similar throughput, we surpass
it by a significant 1.81x in mIoU. ConceptFusion’s 0.03 Hz
throughput makes it impractical for real-time use.
Furthermore, we test the end-to-end throughput of
RayFronts on a real-world outdoor scene using pre-
recorded data from a mobile ground robot. We use a reso- Fig. 5: RayFronts provides state-of-the-art mIoU & 17.5 Hz
lution of 224x224, 30cm voxel size, and the base encoder throughput on an AGX Orin. It surpasses Trident with 1.34x higher
model, while compressing features to top 100 PCA compo- mIoU and a 16.5x speedup, while achieving 1.81x higher mIoU than
NACLIP, which operates at a similar throughput.
nents (retaining ⇠ 80% variance), and disabling ray-tracing.
RayFronts runs real-time at 8.84 Hz on Orin AGX.
RayFronts accurately reconstructs fine-grained details (e.g.,
D. Qualitative Real-World Study
“road cracks”) while detecting far-range objects (e.g., “water
To evaluate RayFronts in unbounded, open-world set- tower”), demonstrating that RayFronts empowers robots
tings, we record a run through an unstructured fire train- within and beyond depth-sensing limitations in open-world
ing facility with a Zed-X camera. As shown in Fig. 1, environments.
VI. C ONCLUSION Training-free embodied object goal navigation with se-
mantic frontiers, 2023.
We present RayFronts, a real-time semantic mapping
[5] Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal-
system for multi-modal open-set scene understanding for
ber, Thomas Funkhouser, and Matthias Nießner. Scannet:
both within- and beyond-range mapping. Our key insight,
Richly-annotated 3d reconstructions of indoor scenes. In
semantic ray frontiers, enables open-set queries about observa-
Proceedings of the IEEE conference on computer vision
tions beyond depth mapping by associating beyond-depth ray
and pattern recognition, pages 5828–5839, 2017.
features with the map’s frontiers. This allows RayFronts
[6] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov,
to significantly reduce search volumes while retaining fine-
Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner,
grained within-range scene understanding. RayFronts im-
Mostafa Dehghani, Matthias Minderer, Georg Heigold,
proves open-set image encoding with an efficient language-
Sylvain Gelly, et al. An image is worth 16x16 words:
aligned encoder, and introduces a new planner-agnostic metric
Transformers for image recognition at scale. arXiv
for open-world search. We achieve state-of-the-art results in
preprint arXiv:2010.11929, 2020.
3D open-set semantic segmentation, strong performance in
[7] Margarita Grinvald, Fadri Furrer, Tonci Novkovic,
online mapping, and efficient encoder throughput. Our future
Jen Jen Chung, Cesar Cadena, Roland Siegwart, and
work aims to include instance differentiation in RayFronts
Juan Nieto. Volumetric instance-aware semantic mapping
and planning integration to facilitate online exploration.
and 3d object discovery. IEEE Robotics and Automation
ACKNOWLEDGMENTS Letters, 4(3):3037–3044, 2019.
[8] Qiao Gu, Ali Kuwajerwala, Sacha Morin,
This work was supported by Defense Science and Technol- Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya
ogy Agency (DSTA) Contract #DST000EC124000205, King Agarwal, Corban Rivera, William Paul, Kirsty Ellis,
Abdulaziz University, and National Institute on Disability, Rama Chellappa, et al. Conceptgraphs: Open-vocabulary
Independent Living, and Rehabilitation Research (NIDILRR) 3d scene graphs for perception and planning. In 2024
Grant #90IFDV0042. We thank Krishna Murthy Jatavallab- IEEE International Conference on Robotics and
hula, Jacob Yeung, Sam Triest and Ayush Jain for insightful Automation (ICRA), pages 5021–5028. IEEE, 2024.
discussions and feedback, and Katerina Nikiforova for early [9] Raphael Hagmanns, Thomas Emter, Marvin Grosse-
encoder exploration. Besselmann, and Jürgen Beyerer. Efficient global occu-
L IMITATIONS pancy mapping for mobile robots using openvdb, 2022.
URL [Link]
While RayFronts is the upper-bound of the online map- [10] Sina Hajimiri, Ismail Ben Ayed, and Jose Dolz.
ping baselines in correct search volume reduction, it also has Pay attention to your neighbours: Training-free open-
the highest memory consumption. However, RayFronts can vocabulary semantic segmentation. arXiv preprint
be tuned down by reducing ray angle bins down until a single arXiv:2404.08181, 2024.
bin (becoming “semantic frontiers”), or reducing depth range [11] Cherie Ho, Jiaye Zou, Omar Alama, Sai Mitheran Ja-
down until 0, giving the flexibility to achieve the best trade-off gadesh Kumar, Benjamin Chiang, Taneesh Gupta, Chen
for an application. Wang, Nikhil Keetha, Katia Sycara, and Sebastian
Scherer. Map it anywhere (mia): Empowering bird’s
R EFERENCES eye view mapping using large-scale public data. arXiv
[1] Graeme Best, Rohit Garg, John Keller, Geoffrey A preprint arXiv:2407.08726, 2024.
Hollinger, and Sebastian Scherer. Resilient multi-sensor [12] Yafei Hu, Quanting Xie, Vidhi Jain, Jonathan Francis,
exploration of multifarious environments with a team of Jay Patrikar, Nikhil Keetha, Seungchan Kim, Yaqi Xie,
aerial robots. In Robotics: Science and Systems (RSS), Tianyi Zhang, Hao-Shu Fang, et al. Toward general-
2022. purpose robots via foundation models: A survey and
[2] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ meta-analysis. arXiv preprint arXiv:2312.08782, 2023.
Altman, Simran Arora, Sydney von Arx, Michael S Bern- [13] Nathan Hughes, Yun Chang, and Luca Carlone. Hy-
stein, Jeannette Bohg, Antoine Bosselut, Emma Brun- dra: A real-time spatial perception system for 3d scene
skill, et al. On the opportunities and risks of foundation graph construction and optimization. arXiv preprint
models. arXiv preprint arXiv:2108.07258, 2021. arXiv:2201.13360, 2022.
[3] Boyuan Chen, Fei Xia, Brian Ichter, Kanishka Rao, [14] Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala,
Keerthana Gopalakrishnan, Michael S Ryoo, Austin Qiao Gu, Mohd Omama, Tao Chen, Alaa Maalouf,
Stone, and Daniel Kappler. Open-vocabulary queryable Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil
scene representations for real world planning. In 2023 Keetha, et al. Conceptfusion: Open-set multimodal 3d
IEEE International Conference on Robotics and Automa- mapping. arXiv preprint arXiv:2302.07241, 2023.
tion (ICRA), pages 11509–11522. IEEE, 2023. [15] Christina Kassab, Matias Mattamala, Lintong Zhang, and
[4] Junting Chen, Guohao Li, Suryansh Kumar, Bernard Maurice Fallon. Language-extended indoor slam (lexis):
Ghanem, and Fisher Yu. How to not train your dragon: A versatile system for real-time visual scene under-
standing. In 2024 IEEE International Conference on Learning transferable visual models from natural lan-
Robotics and Automation (ICRA), pages 15988–15994. guage supervision. In International conference on ma-
IEEE, 2024. chine learning, pages 8748–8763. PMLR, 2021.
[16] Nikhil Keetha, Avneesh Mishra, Jay Karhade, Kr- [26] Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo
ishna Murthy Jatavallabhula, Sebastian Scherer, Madhava Molchanov. Am-radio: Agglomerative vision foundation
Krishna, and Sourav Garg. Anyloc: Towards universal model reduce all domains into one. In CVPR, pages
visual place recognition. IEEE Robotics and Automation 12490–12500, 2024.
Letters, 9(2):1286–1293, 2023. [27] Martin Runz, Maud Buffier, and Lourdes Agapito. Mask-
[17] Seungchan Kim, Micah Corah, John Keller, Graeme fusion: Real-time recognition, tracking and reconstruc-
Best, and Sebastian Scherer. Multi-robot multi-room tion of multiple moving objects. In 2018 IEEE in-
exploration with geometric cue extraction and circular ternational symposium on mixed and augmented reality
decomposition. IEEE Robotics and Automation Letters, (ISMAR), pages 10–20. IEEE, 2018.
9(2):1190–1197, 2023. [28] Yuheng Shi, Minjing Dong, and Chang Xu. Harnessing
[18] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi vision foundation models for high-performance, training-
Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, free open vocabulary segmentation. arXiv preprint
Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, arXiv:2411.09219, 2024.
et al. Segment anything. In CVPR, pages 4015–4026, [29] Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen,
2023. Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-
[19] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Artal, Carl Ren, Shobhit Verma, et al. The replica
Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei dataset: A digital replica of indoor spaces. arXiv preprint
Yang, Hang Su, et al. Grounding dino: Marrying dino arXiv:1906.05797, 2019.
with grounded pre-training for open-set object detection. [30] Hugues Thomas, Mouli Sivapurapu, and Jian Zhang.
In European Conference on Computer Vision, pages 38– Embedding pose graph, enabling 3d foundation model
55. Springer, 2024. capabilities with a compact representation. arXiv preprint
[20] John McCormac, Ankur Handa, Andrew Davison, and arXiv:2403.13777, 2024.
Stefan Leutenegger. Semanticfusion: Dense 3d semantic [31] Wenshan Wang, Delong Zhu, Xiangwei Wang, Yaoyu
mapping with convolutional neural networks. In 2017 Hu, Yuheng Qiu, Chen Wang, Yafei Hu, Ashish Kapoor,
IEEE International Conference on Robotics and automa- and Sebastian Scherer. Tartanair: A dataset to push the
tion (ICRA), pages 4628–4635. IEEE, 2017. limits of visual slam. In 2020 IEEE/RSJ International
[21] John McCormac, Ronald Clark, Michael Bloesch, An- Conference on Intelligent Robots and Systems (IROS),
drew Davison, and Stefan Leutenegger. Fusion++: Vol- pages 4909–4916. IEEE, 2020.
umetric object-level slam. In 2018 international confer- [32] Abdelrhman Werby, Chenguang Huang, Martin Büchner,
ence on 3D vision (3DV), pages 32–41. IEEE, 2018. Abhinav Valada, and Wolfram Burgard. Hierarchi-
[22] Ken Museth, Jeff Lait, John Johanson, Jeff Budsberg, cal Open-Vocabulary 3D Scene Graphs for Language-
Ron Henderson, Mihai Alden, Peter Cucka, David Hill, Grounded Robot Navigation. In Proceedings of Robotics:
and Andrew Pearce. Openvdb: an open-source data Science and Systems, Delft, Netherlands, July 2024. doi:
structure and toolkit for high-resolution volumes. In ACM 10.15607/[Link].077.
SIGGRAPH 2013 Courses, SIGGRAPH ’13, New York, [33] Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan,
NY, USA, 2013. Association for Computing Machinery. Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang,
ISBN 9781450323390. doi: 10.1145/2504435.2504454. Yunhai Tong, Xudong Jiang, et al. Towards open
URL [Link] vocabulary learning: A survey. IEEE Transactions on
[23] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Pattern Analysis and Machine Intelligence, 46(7):5092–
Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fer- 5113, 2024.
nandez, Daniel Haziza, Francisco Massa, Alaaeldin El- [34] Quanting Xie, So Yeon Min, Tianyi Zhang, Kedi Xu,
Nouby, et al. Dinov2: Learning robust visual features Aarav Bajaj, Ruslan Salakhutdinov, Matthew Johnson-
without supervision. arXiv preprint arXiv:2304.07193, Roberson, and Yonatan Bisk. Embodied-rag: General
2023. non-parametric embodied memory for retrieval and gen-
[24] Songyou Peng, Kyle Genova, Chiyu Jiang, Andrea eration. arXiv preprint arXiv:2409.18313, 2024.
Tagliasacchi, Marc Pollefeys, Thomas Funkhouser, et al. [35] Binbin Xu, Wenbin Li, Dimos Tzoumanikas, Michael
Openscene: 3d scene understanding with open vocabu- Bloesch, Andrew Davison, and Stefan Leutenegger. Mid-
laries. In Proceedings of the IEEE/CVF conference on fusion: Octree-based object-level multi-instance dynamic
computer vision and pattern recognition, pages 815–824, slam. In 2019 International Conference on Robotics and
2023. Automation (ICRA), pages 5231–5237. IEEE, 2019.
[25] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya [36] Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang
Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Wang, and Bernadette Bucher. Vlfm: Vision-language
Amanda Askell, Pamela Mishkin, Jack Clark, et al. frontier maps for zero-shot semantic navigation. In
2024 IEEE International Conference on Robotics and A2. O FFLINE S EMANTIC M APPING V ISUALIZATIONS
Automation (ICRA), pages 42–48. IEEE, 2024. Fig. A.3 showcases open-vocabulary semantic segmentation
[37] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and samples across different datasets, while Fig. A.4 highlights the
Lucas Beyer. Sigmoid loss for language image pre- open-vocabulary capabilities of RayFronts by showing the
training. In CVPR, pages 11975–11986, 2023. segmentations of multiple long-tail classes.
[38] Chong Zhou, Chen Change Loy, and Bo Dai. Extract
free dense labels from clip. In European Conference on
Computer Vision, pages 696–712. Springer, 2022.

A PPENDIX

A1. O NLINE S EMANTIC M APPING V ISUALIZATIONS

Fig. A.1 shows illustrations of the different baselines men-


tioned in Section IV-A. The top left part of the figure highlights
how RayFronts can avoid semantic feature collisions (the
case where different semantic features are forced to be fused
together) by utilizing multiple rays to describe the different
semantic entities. Whereas semantic frontier approaches (irre-
spective of the unidirectional/spherical search volume method)
can fail when semantically different entities are observed
through the same frontier. The top right part of the figure
emphasizes where global encoding approaches like Sem Poses
can fail to capture non-prominent objects in the presence
of a large centered entity. The illustration provides further
explanation for Sem Poses’s inability to capture chimneys in
the AbandonedCableDay scene as shown in Figs. A.2 and 4.
Finally, the bottom row illustrates how each baseline computes
its search volume. Spherical Sem Fronts can fail to capture
a distant object, with increasing radius cubically increasing
search volume. Unidirectional Sem Fronts is highly sensitive
to the mapped region topology since it uses it to infer the
semantic ray direction, and in the illustrated case, it fails. Sem
Poses fails to utilize depth information to push the ray further
onto the mapped region boundary for better localization. In
contrast with all baselines, RayFronts is able to accurately
determine the direction of the ray and limit the search volume
efficiently, utilizing depth information if available.
Fig. A.2 visualizes the two query scenarios shown in
Fig. 4 with ground truth generated at an 80m cutoff (Highest
value that fits in our memory) for further clarity. The top
block shows search volumes for building at a particular
time step. At 20m depth sensing range, it is clear that
RayFronts achieves the best search volume, having 1.35 ⇥
higher SCVR than Unidirectional Sem Fronts. Spherical Sem
Fronts struggles to cast a search volume that encompasses
big objects, whereas Unidirectional Sem Fronts has some
erroneously inferred ray directions. At 0m range, both Sem
Poses and RayFronts perform similarly. Furthermore,
the bottom block shows the search volume of distant non-
prominent objects that never come into the depth sensing
range. At 0m range, Sem Poses fails to capture the semantics
and fails to reduce search volume, yielding an SCVR of 0,
whereas RayFronts provides meaningful areas to explore.
RayFronts vs Sem Fronts RayFronts (No-Depth) vs Sem Poses

Dense encoding Image level encoding supresses


captures all objects non-prominent objects
Features preserved with distinct rays Feature Collision = Non-Identifiable

RayFronts Spherical Sem Fronts Unidirectional Sem Fronts RayFronts (No-Depth) / Sem Poses

Frontier

Inferred-direction away
from mapped region

Mapped Region Frontier Robot Objects of Interest Search Volume

Fig. A.1: Top left shows how RayFronts is able to avoid feature collisions through the use of multiple rays that capture distinct semantics
observed through the same frontier, where semantic frontier approaches [4, 36] fail. The top right illustrates that even with no depth
information, RayFronts dense language-aligned encoding can allow it to capture non-prominent semantics where semantic pose approaches
[30] fail. The bottom row highlights that RayFronts is the upper bound in accurately reducing search volume.
Building
Spherical Sem Fronts (20m) Unidirectional Sem Fronts (20m) Sem Poses (0m)

RayFronts (20m) RayFronts (0m)

Chimney
Spherical Sem Fronts (20m) Unidirectional Sem Fronts (20m) Sem Poses (0m)

RayFronts (20m) RayFronts (0m)

Ground Truth Search Volume Mapped Voxels Activated Voxels Rays Activated Rays

Fig. A.2: Two query scenarios are shown with GT generated at 80m as opposed to 50m cutoff for more clarity: (1) querying for a prominent
object (i.e, Building) that enters depth range, and (2) a distant object (i.e, Chimney) that remains beyond range. Through unified dense voxel
mapping and beyond-range semantic ray frontiers, RayFronts sets the upper bound in both scenarios.
Fig. A.3: Sample visualizations of offline semantic mapping generated by RayFronts for scenes from Replica [29] (room0 and office2),
ScanNet [5](scene0050 and scene00378), and the chosen four scenes from TartanAir [31]. “RGB”, “GT” and “PRED” refer to the RGB
scene reconstruction, Ground Truth semantics, and semantic segmentation prediction by RayFronts, respectively, for each corresponding
scene. RayFronts achieves SOTA mIoU for 3D open-vocabulary semantic segmentation.

Fig. A.4: Examples of long-tail classes segmented by RayFronts across outdoor scenes from TartanAir [31]. We set the voxel size to 0.5
(50cm) for the visualizations. For each set, we present the RGB image, the corresponding 3D reconstructed view, and the classified voxels
left to right respectively. RayFronts effectively segments long-tail concepts.

You might also like