0% found this document useful (0 votes)
30 views21 pages

NOVA: High-Speed UAV Target Tracking

NOVA is a novel framework for autonomous aerial target tracking in GPS-denied and unstructured environments, utilizing only a stereo camera and an IMU for robust navigation and collision avoidance. By formulating perception, estimation, and control in the target's reference frame, NOVA eliminates the need for global localization and enables high-speed tracking exceeding 50 km/h in challenging conditions. The system has been validated in various real-world scenarios, demonstrating consistent performance without reliance on external infrastructure or pre-mapped environments.

Uploaded by

samgpt321
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
30 views21 pages

NOVA: High-Speed UAV Target Tracking

NOVA is a novel framework for autonomous aerial target tracking in GPS-denied and unstructured environments, utilizing only a stereo camera and an IMU for robust navigation and collision avoidance. By formulating perception, estimation, and control in the target's reference frame, NOVA eliminates the need for global localization and enables high-speed tracking exceeding 50 km/h in challenging conditions. The system has been validated in various real-world scenarios, demonstrating consistent performance without reliance on external infrastructure or pre-mapped environments.

Uploaded by

samgpt321
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

.

NOVA: Navigation via Object-Centric


Visual Autonomy for High-Speed
Target Tracking in Unstructured
GPS-Denied Environments
Alessandro Saviolo and Giuseppe Loianno
1
New York University, New York, NY 11201 USA
Corresponding author: Alessandro Saviolo (email: [Link]@[Link]).
“This work was supported by the NSF CAREER Award 2145277, DARPA YFA Grant D22AP00156-00, Qualcomm Research, Nokia, and
NYU Wireless.”
arXiv:2506.18689v2 [[Link]] 7 Jul 2025

ABSTRACT Autonomous aerial target tracking in unstructured and GPS-denied environments remains a
fundamental challenge in robotics. Many existing methods rely on motion capture systems, pre-mapped
scenes, or feature-based localization to ensure safety and control, limiting their deployment in real-world
conditions. We introduce NOVA, a fully onboard, object-centric framework that enables robust target
tracking and collision-aware navigation using only a stereo camera and an IMU. Rather than constructing
a global map or relying on absolute localization, NOVA formulates perception, estimation, and control
entirely in the target’s reference frame. A tightly integrated stack combines a lightweight object detector
with stereo depth completion, followed by histogram-based filtering to infer robust target distances under
occlusion and noise. These measurements feed a visual-inertial state estimator that recovers the full 6-DoF
pose of the robot relative to the target. A nonlinear model predictive controller (NMPC) plans dynamically
feasible trajectories in the target frame. To ensure safety, high-order control barrier functions (CBFs) are
constructed online from a compact set of high-risk collision points extracted from depth, enabling real-
time obstacle avoidance without maps or dense representations. We validate NOVA across challenging
real-world scenarios, including urban mazes, forest trails, and repeated transitions through buildings with
intermittent GPS loss and severe lighting changes that disrupt feature-based localization. Each experiment
is repeated multiple times under similar conditions to assess resilience, showing consistent and reliable
performance. NOVA achieves agile target following at speeds exceeding 50 km/h. These results show that
high-speed, vision-based tracking is possible in the wild using only onboard sensing, with no reliance on
external localization or assumptions on the environment structure. Video: [Link]

INDEX TERMS Target tracking, Aerial robotics, Vision-based navigation, Model predictive control, Control
barrier functions, Visual-inertial odometry, Depth completion, Object detection, GPS-denied environments.

I. INTRODUCTION This capability is key to diverse applications. In search and

T ARGET tracking is the ability of an autonomous system


to detect, follow, and predict the motion of an object of
interest over time. For Unmanned Aerial Vehicles (UAVs),
rescue, tracking allows UAVs to follow moving individuals
through unstructured terrain, maintaining proximity without
relying on maps or GPS [1]. In human–robot interaction, it
this capability is critical in dynamic and unstructured en- enables drones to interpret and respond to human motion in
vironments where pre-mapped routes or fixed waypoints real time, facilitating adaptive behavior [2], [3]. In inspection
are insufficient. In such scenarios, robust tracking allows and surveillance, tracking ensures continuous observation
the UAV to remain persistently coupled to targets such as of mobile assets, even under occlusion or changing view-
people, vehicles, or infrastructure, despite challenges like points [4]–[6]. Moreover, tracking is foundational in multi-
scene changes, occlusions, or unpredictable target motion. robot systems, where relative positioning between agents

1
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

Figure 1: Satellite imagery of representative NOVA flight missions in real-world environments. Each overlaid UAV
trajectory illustrates a distinct experimental scenario designed to test specific failure modes in visual-inertial target tracking
and control. Red: High-speed tracking in a forest trail over 1 km, with target speeds exceeding 50 km/h and motion blur.
Blue: Indoor–outdoor transition with complete GPS loss, exposure collapse, and perceptual ambiguity due to structural
occlusions and distractors. Orange: Elevated tracking from a 10 m offset, challenging stereo depth perception, feature
consistency, and planning geometry. Green: Urban container maze with narrow corridors, cluttered turns, and minimal
texture. Insets show raw onboard RGB frames during each flight. NOVA maintains real-time target lock, depth estimation,
obstacle avoidance, and closed-loop control using only a stereo camera and IMU. GPS is not used by NOVA and appears
here only for visualization of flight paths.

is essential. UAVs in leader–follower formations or aerial ambiguous semantics, and degraded visibility. As a result,
swarms must track teammates to maintain coordinated mo- feature-based localization quickly becomes unreliable [24].
tion [7]–[9]. In autonomous landing, a UAV must descend A notable example of this failure was NASA’s Inge-
onto a moving platform, requiring precise tracking [10]–[12]. nuity Mars Helicopter [25]–[27], which suffered a navi-
Visual tracking methods for UAVs traditionally fall into gation failure while flying over smooth, featureless sand
two categories: Image-Based Visual Servoing (IBVS) and ripples [28]. Without distinctive landmarks, its visual-inertial
Position-Based Visual Servoing (PBVS) [13]–[17]. IBVS op- system produced erroneous velocity estimates, leading to a
erates directly on image-plane measurements, allowing low- crash 20 seconds after takeoff. Had Ingenuity been able to
latency control and robustness to calibration errors. However, localize relative to a known object in the scene, such as the
its effectiveness is limited by the camera’s field of view and Perseverance rover, it might have maintained stable flight and
sensitivity to target occlusion or rapid motion. PBVS instead avoided the crash. This example highlights a broader insight:
reconstructs the target’s 3D position, typically in a globally when a known and observable target is available, it can serve
referenced frame, and uses it to guide control. While PBVS as a reliable frame of reference than the surrounding terrain.
mitigates many of IBVS’s limitations, it introduces a new This paper introduces NOVA, a unified framework for
dependency: accurate consistent global localization. Navigation via Object-centric Visual Autonomy. NOVA en-
In GPS-denied settings such as forests, urban mazes, ables real-time aerial target tracking and obstacle-aware
indoor spaces, or extraterrestrial terrain, global localization navigation using only onboard sensing, with no dependence
typically relies on Visual-Inertial Odometry (VIO) [18]–[20] on GPS, external maps, or motion capture. Unlike traditional
or Simultaneous Localization and Mapping (SLAM) [21]– VIO or SLAM-based methods that require globally consis-
[23], which estimate the robot’s pose by tracking visual tent localization, NOVA formulates perception, estimation,
features across image sequences. These methods assume and control directly in the target’s frame of reference. This
sufficient texture, consistent lighting, and a relatively static object-centric approach enables tracking in environments
environment to maintain reliable feature correspondences. where localization is unreliable or unavailable.
However, unstructured environments often violate these as- At the core of NOVA is a tightly integrated perception-
sumptions due to cluttered geometry, dynamic elements, to-control stack that couples a custom object detector, depth
completion module, and visual-inertial state estimator. The

2
object detector operates in real time on low-resolution inputs to movement. This method is efficient and effective for tasks
using an adaptive zoom strategy to maintain tracking at long where the target remains in view and image-space motion
ranges. Depth estimates are computed using a fusion of corresponds smoothly to inputs.
stereo and monocular cues, then processed via histogram- PBVS, on the other hand, estimates the full 3D pose of
based filtering to produce robust target distance estimates, the target using stereo vision, depth sensors, or geometric
even under occlusion or noise. From this, a full 6-DoF state priors, and computes control actions in Euclidean space. It
of the robot relative to the target is inferred in real time. aims to minimize the spatial error between the current and
Crucially, this perception system is designed to operate under desired poses of the robot relative to the target. This spatial
realistic sensing conditions, such as degraded lighting, partial reasoning improves robustness when the target leaves the
occlusions, and unstructured terrain. field of view. However, PBVS depends on accurate global
These target-centric estimates are passed to a Nonlinear pose estimates of both the robot and the target, requiring
Model Predictive Controller (NMPC) that optimizes dynam- consistent global localization across frames. In GPS-denied
ically feasible trajectories directly in the target’s reference or visually degraded environments, this dependency often
frame. To ensure safe operation in cluttered environments, becomes a critical point of failure.
obstacle avoidance is enforced online using high-order Con- While both paradigms are effective in structured, static en-
trol Barrier Functions (CBFs), derived from a compact set vironments [29]–[33], their assumptions often break down in
of high-risk collision points identified in the depth map. aerial target tracking. UAVs operate under rapid motion and
This eliminates the need for dense mapping or environmental wide viewpoint changes, which distort target appearance and
priors, enabling real-time, map-free collision avoidance. invalidate interaction models. During aggressive maneuvers,
We validate NOVA across a suite of challenging real-world UAVs experience frequent occlusions, motion blur, and target
experiments, illustrated in Figure 1. Each scenario targets a scale variation, causing visual features to become unreliable.
specific failure mode common to traditional visual tracking IBVS fails when the target exits the frame, undergoes large
pipelines. In the forest trail mission, the UAV follows a appearance shifts, or when image-space motion no longer
moving target at speeds exceeding 50 km/h over more than correlates predictably with control. PBVS struggles in GPS-
1 km, inducing severe motion blur and visual degradation denied environments or texture-sparse scenes where global
from airborne dust. The building transition scenario involves pose estimation is not available or stable.
abrupt lighting changes and full GPS loss as the target moves These challenges are further exacerbated by the limited
from an open parking lot into an enclosed space and back computational resources available onboard UAVs and the
outdoors. In the height-offset trial, the UAV maintains a stringent real-time requirements for flight control. In prac-
vertical height of 10 m, challenging feature matching due tice, neither classical IBVS nor PBVS alone can reliably
to reduced parallax and resolution. handle the demands of high-speed, long-range target tracking
Across all experiments, NOVA maintains stable target in unstructured, dynamic environments.
lock, plans feasible trajectories, and avoids obstacles in real
time. To evaluate repeatability, each mission is conducted
B. Data-driven Aerial Target Tracking
multiple times under similar conditions, with consistent per-
formance observed throughout. All sensing and computation Recent approaches in aerial target tracking have leveraged
are performed fully onboard using only a stereo camera the power of deep learning to enhance perception and
and an IMU. Unlike prior methods that rely on motion decision-making under challenging conditions. One promi-
capture, artificial landmarks, or pre-built maps, NOVA op- nent direction is the use of end-to-end learning, where a
erates autonomously in unstructured environments with no neural network directly maps raw sensory inputs to low-
external infrastructure. To the best of our knowledge, this level control actions. These models, often trained via rein-
is the first demonstration of real-time, high-speed object- forcement learning or imitation learning, have demonstrated
centric navigation that operates fully onboard and in-the- fast reaction times and adaptability to perceptual noise [34]–
wild, without external infrastructure, explicit mapping, or [37]. However, this tight coupling of perception and control
dependency on structured priors. sacrifices interpretability, and poses significant challenges in
generalization to unseen environments. Moreover, training
such models requires extensive data collected in diverse,
II. RELATED WORKS unstructured settings, an effort that is not only costly but
A. Traditional Visual Servoing for Target Tracking also difficult to scale. Unstructured outdoor environments,
Visual servoing provides a well-established framework for in particular, exhibit extreme variability in appearance, ge-
controlling robotic motion with respect to a target, typically ometry, and semantics, making it impractical to fully cover
categorized into IBVS and PBVS [13]–[17]. In IBVS, control their distribution during training [38].
commands are computed by minimizing the error between To address these limitations, a more modular paradigm
observed and desired image-space features. These features has emerged where perception and control are treated as
are extracted from the camera view and mapped to robot distinct but tightly integrated subsystems. In this setting, the
motion using an interaction matrix that links image variation perception module is responsible for detecting and localizing

3
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

the target, while the control module translates this infor- More onboard-focused designs have emerged, such as
mation into safe and efficient motion commands. Extensive [60], which uses a monocular camera and a blob detector
research has been devoted to each side of the pipeline. On to track a single spherical obstacle. The framework inte-
the perception front, the computer vision community has grates perception and control onboard, but the simplicity of
produced a range of approaches for object detection and the sensing pipeline imposes its own constraints. Obstacle
tracking, from compact architectures optimized for real-time avoidance is implemented as a soft cost, tuned manually
inference on embedded platforms [39]–[42], to large-scale against the tracking objective, and the system assumes only
foundation models capable of open-set recognition and zero- one obstacle is visible at a time. The drone must be launched
shot generalization [43]–[46]. Techniques including temporal manually, and no mechanism is provided for adapting to
attention, depth prediction, and visual transformers have clutter or dynamic changes. While a step forward, the system
improved robustness to occlusion and visual degradation. is difficult to generalize to real-world scenarios involving
Alongside advances in perception, control design has dense geometry or multiple obstacles.
emerged as a key focus in enabling robust aerial target Multi-agent frameworks like CoNi-MPC [61] take a dif-
following. While model-free approaches such as reinforce- ferent approach, where a UAV follows a ground robot by
ment learning have demonstrated flexibility and general- optimizing a relative objective. This eliminates the need for
ization [47]–[50], they often struggle to provide consistent global SLAM, but again depends on full-state telemetry from
safety or performance guarantees, particularly under distri- the target and is tested under motion capture. The absence
bution shifts or in dynamic settings. In contrast, model- of onboard perception and low operational speeds (under
based methods, especially those based on NMPC [51]–[54], 0.4 m/s) further limit real-world relevance.
offer a structured way to embed physical constraints, safety Later extensions add LiDAR-based obstacle sensing [62],
margins, and task-specific objectives directly into the control improving situational awareness. However, core limitations
policy. By explicitly optimizing trajectories over a finite persist. The system still relies on external tracking for the
horizon, NMPC enables predictive and context-aware behav- target and operates in structured, static environments at low
iors, which are critical for maintaining target visibility and speeds (below 1.5 m/s). Visual perception remains absent,
avoiding obstacles during high-speed, reactive flight [55]. and the setup assumes favorable sensing conditions.
This separation of perception and control not only im- Together, these efforts reflect growing interest in object-
proves system modularity and interpretability, but also allows relative control, but also reveal a gap: current systems are not
for leveraging the complementary strengths of data-driven designed for fast, reactive tracking in complex environments
learning and analytical planning. However, the efficacy of using only onboard sensing. Our framework is built to
such hybrid systems depends critically on how well the two address exactly that gap. By combining real-time object
modules are coupled. detection, depth-completed state estimation, and constraint-
aware planning in a fully onboard pipeline, we enable agile,
target-relative control in GPS-denied, cluttered environment,
C. Object-Centric Visual Autonomy
without reliance on external infrastructure or prior maps.
Traditional navigation frameworks rely on global localization
through GPS, VIO, or SLAM, followed by waypoint-based
III. METHODOLOGY
trajectory generation [56]–[58]. While effective in struc-
A. System Overview
tured environments, these methods often fail in GPS-denied
or visually degraded settings, where feature tracking and NOVA is a fully onboard framework for vision-based aerial
map consistency degrade. Object-centric navigation offers target tracking and collision avoidance in GPS-denied and
a promising alternative: instead of referencing a global visually degraded environments. It requires no maps, infras-
frame, control objectives are defined relative to specific scene tructure, or prior tuning, enabling deployment in unknown
objects, such as people or vehicles. This allows robots to and cluttered settings. The system is structured around two
maintain situational awareness and track targets even when tightly integrated modules: perception and control (Figure 2).
global localization is unreliable or unavailable. The perception stack processes visual and inertial input to
Several recent works have pursued this idea, though typi- estimate the target’s relative state and detect potential colli-
cally under restrictive assumptions that limit applicability in sion risks in real time. These estimates feed into a NMPC
high-speed or unstructured environments. that dynamically tracks the target while enforcing safety
An early example is [59], which formulates control di- through system dynamics and real-time obstacle constraints.
rectly in the target’s frame of reference. While conceptually The entire pipeline runs onboard, supporting reactive agile
aligned with object-centric reasoning, the system relies en- flight with no external dependency.
tirely on motion capture to provide perfect state estimates
for both robot and obstacle. No onboard sensing is used, B. Perception
and full target velocity is assumed known. This removes The perception pipeline estimates the target’s position and
perception from the equation, making the setup impractical surrounding obstacles in real time using only onboard visual
outside controlled environments. and inertial sensors. At each frame, the system detects the

4
Figure 2: System architecture overview. The framework consists of two tightly integrated modules: perception and control.
The perception stack detects the target using a lightweight object detector, estimates its depth via stereo-monocular disparity
alignment, and processes obstacle information for collision avoidance. These outputs are fused with inertial data (omitted
from figure for simplicity) to produce a smooth, high-frequency target-centric odometry. The control module uses NMPC,
augmented with high-order CBFs, to generate safe and agile thrust and angular rate commands for tracking the target.

target, estimates its depth, and fuses this information with


IMU data to maintain a full relative state. Simultaneously,
it identifies potential collision points using completed depth
maps. These estimates form the basis for downstream target-
centric planning and control.

1) Target Detection
A central challenge in target tracking lies in balancing de- Figure 3: Adaptive zoom strategy. The system dynamically
tection range with computational efficiency. High-resolution crops and rescales the input image to center the target while
inputs, such as 640 × 480 pixels, support target detection suppressing irrelevant context. This targeted zoom improves
at distances beyond 40 m but exceed the processing limits detection confidence and robustness, particularly at long
of embedded platforms. In contrast, lower-resolution inputs range or in the presence of visually similar objects, by
like 320 × 240 pixels are more computationally efficient but reducing ambiguity and limiting false positives.
constrain detection to shorter ranges, often under 10 m. To
resolve this trade-off, we introduce an adaptive zooming
mechanism that dynamically crops and rescales the input
image around a region of interest (Figure 3). We employ a lightweight, custom-trained object detector
The zoom module takes the full-resolution RGB frame based on the YOLOv11-small architecture to enable onboard
It ∈ RH×W ×3 and a bounding box bt−1 = [x, y, w, h], real-time detection with limited computational resources.
where (x, y) denotes the center coordinate and (w, h) the The input to the detector is the zoomed-in crop Itcrop , and
width and height, from the previous detection. This box the output is a 2D bounding box bt in crop coordinates,
is enlarged by a factor α > 1 to introduce a margin which is then projected back into the full-resolution frame.
that accounts for motion and detection uncertainty. A crop The object detector is trained from scratch on the COCO
window is then centered on the enlarged box and shaped dataset [63] with domain-specific augmentations designed to
to match the detector’s aspect ratio. The zoomed-in region mirror the visual effects introduced by our zooming strategy.
is extracted from It and resized to the detector’s input In particular, randomized zooming and resizing simulate the
resolution, yielding Itcrop . pixelation and scale variations that occur during aggressive
If detection fails at time t, the crop size is progressively cropping, ensuring the detector remains effective even when
increased to widen the search region. Once the target is targets appear coarse or low-resolution. Additional augmen-
reacquired, the zoom readjusts to tightly frame the detection. tations include pitch and roll rotations and lighting variations
This mechanism preserves long-range detection capabilities to improve robustness under aerial tracking conditions.
at lower resolutions while improving robustness. By nar-
rowing the detector’s receptive field to a localized region,
the system avoids distractions from background clutter and 2) Depth Estimation
allows the network to focus on the target, increasing confi- Accurate depth perception is essential for estimating the
dence and accuracy under occlusion and motion. target distance as well as ensuring safe navigation. While

5
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

Figure 4: Depth completion pipeline. Monocular (top) and


stereo (bottom) disparity maps are fused to produce a dense,
absolute-depth estimate. Monocular predictions capture rela-
tive structure, while stereo provides scale in textured regions.
A polynomial alignment module merges the two in disparity
space for accurate, absolute depth. Figure 5: Histogram-based mode filtering for target depth
estimation. Depth values inside the detected bounding box
are collected and binned. The most frequent bin is selected to
stereo cameras provide absolute depth maps Dtabs ∈ RH×W , estimate the target distance, improving robustness to noise,
these measurements are often sparse or noisy in areas with occlusion, and background clutter.
low texture, specular surfaces, or insufficient baseline [55].
To overcome these limitations, we employ a learning-based
depth completion approach that fuses monocular and stereo partial occlusions. For instance, the target may remain visible
cues via a disparity-domain alignment (Figure 4). and correctly detected, while the center pixel falls on a
We compute the absolute disparity from stereo depth as: background object or an occluding surface, yielding an
f ·B incorrect depth estimate.
∆tabs = t , (1) To address this, we apply histogram-based mode filtering.
Dabs + ϵ
We extract all valid depth values within bt from Dtcom , bin
where f represents the camera’s focal length, B the stereo
them into fixed-width depth intervals, and select the most
baseline, and ϵ a small constant used for numerical stability.
frequent bin as the representative estimate, denoted z . This
In parallel, a monocular depth estimation neural network
approach leverages the observation that the dominant mode
Nmde (It ) [64] predicts a relative disparity map ∆trel .
in the depth distribution corresponds to the visible target. If
To align scales, we fit a second-order polynomial between
this was not the case, the detector would likely fail to localize
the two disparity maps using valid stereo pixels defined by
the target reliably. As illustrated in Figure 5, this mode-based
a binary mask Mt ∈ {0, 1}H×W . Let n be the number
filtering improves robustness to noise, clutter, and occlusion
of valid pixels in Mt . We construct a vector y ∈ Rn
without requiring pixel-level semantic segmentation.
containing the absolute disparities ∆tabs (i, j) at valid pixel
locations. We also construct a design matrix X ∈ Rn×3 ,
where each row Xk = [∆trel (ik , jk )2 , ∆trel (ik , jk ), 1]. The
3) Object-Centric State Estimation
vector of polynomial coefficients is denoted θ = [a, b, c]⊤ .
The polynomial fit is defined as a least squares problem To enable smooth and reactive control, we estimate the full 6-
DoF relative state of the quadrotor with respect to the target
min |Xθ − y|2 , (2) by fusing low-rate visual detections with high-rate inertial
θ∈R3
measurements. Each visual update provides the 2D center
with the closed-form solution
coordinate (x, y) of the detected bounding box and the target
θ ∗ = (X⊤ X)−1 X⊤ y. (3) depth z computed via histogram-based filtering.
The resulting polynomial parameters are applied across the As illustrated in Figure 6, the detected 2D center coordi-
full relative disparity map to generate a completed disparity nate is back-projected into a 3D point in the camera frame C
estimate: using the known camera intrinsics (fx , fy , cx , cy ) obtained
through camera calibration [65]:
∆tcom = a · (∆trel )2 + b · ∆trel + c, (4)  
(x − cx )z/fx
and converted back into a dense completed depth map:
pC =  (y − cy )z/fy  . (6)
f ·B z
Dtcom = t . (5)
∆com + ϵ
The resulting 3D point is transformed into the body frame
To estimate the target’s depth, a naive approach might
B using the static extrinsics between the camera and IMU:
extract a single value from the center of the detected
bounding box bt . However, this is prone to failure under pB = qBC ⊙ pC + tBC , (7)

6
We first generate a TTC map Tt ∈ RH×W by projecting
robot’s velocity vector vCt , expressed in the camera frame,
along the viewing ray direction of each pixel. These ray
directions Rt (i, j) are unit-length vectors pointing from the
camera center through each pixel (i, j), and are computed
from the intrinsic calibration parameters. The TTC map is
then computed as:

Dtcom (i, j)
Tt (i, j) = . (9)
∥vCt · Rt (i, j)∥
Figure 6: Coordinate frame convention. The object-centric
To reduce the dimensionality of Tt , we divide it into non-
frame T is defined at the target location, with its orientation
overlapping grid cells Pu,v of size P ×P , and select the pixel
fixed and aligned with gravity using the IMU. A 2D detection
in each cell with the minimum TTC:
(x, y) in the camera frame C is back-projected into 3D and
transformed into the body frame B using known extrinsics
( )
t
(qBC , tBC ). The resulting position is then rotated into the Sttc = (i, j) ∈ Tt | (i, j) = arg min Tt (i, j), ∀ u, v .
(i,j)∈Pu,v
target frame T using the IMU-derived attitude qT B . This
(10)
defines the relative pose pT , which serves as the reference
We then apply a filtering step to retain only those points
for target-centric perception and control.
likely to intersect the quadrotor’s projected dimensions at
their corresponding depths:
( )
where qBC is the camera-to-body rotation (as a quaternion), t
t t |i − c x | ≤ Qx f x /[2D com (i, j)],
tBC the translation, and ⊙ the quaternion-vector product. Sfilt = (i, j) ∈ Sttc ,
|j − cy | ≤ Qy fy /[2Dtcom (i, j)]
We define a moving target-centric frame T , centered on (11)
the target and aligned with the ground plane. Although this where Qx , Qy represent the robot’s width and height.
frame may translate over time, we assume it remains parallel From this filtered set, we extract the top-K most critical
to the ground, and we model its orientation using the IMU- collision points:
derived attitude of the robot qT B . The relative position of X
the robot in this target frame is given by: t
Stop- K = arg min Tt (i, j), (12)
S ′tfilt ⊆Sfilt
t
,|S ′tfilt |=K (i,j)∈S ′ t
pT = −(qT B ⊙ pB ). (8) filt

The resulting vector pT expresses the quadrotor’s position The selected pixel coordinates are back-projected into 3D
relative to the target. This signal is tracked over time using points in the camera frame C :
an Unscented Kalman Filter (UKF) [66]. The UKF incorpo- ( )
z = Dtcom (i, j),

rates IMU’s angular velocity as process input, while visual t (i − cx )z (j − cy )z
SC = , ,z t .
measurements provide asynchronous updates. The full target- fx fy (i, j) ∈ Stop- K
centric state consists of position pT , velocity vT , orientation (13)
qT B , and angular velocity ωB . Finally, these 3D points are transformed into the body
This approach eliminates the need for any environment- frame B and then to the target-centric frame T using the
specific information, such as visual texture or persistent known extrinsics:
features, for state estimation. Unlike traditional VIO or
STt = qT B · (qBC ⊙ sC + tBC ) sC ∈ SCt ,

SLAM systems, our method relies solely on the detection (14)
and tracking of the target to estimate the full relative state
of the quadrotor. This enables robust performance even in These target-centric obstacle points STt are published
textureless, dynamic, and unknown environments, making it at each planning cycle and used to construct hard safety
especially well-suited for agile target tracking. constraints within our NMPC formulation.

C. Planning and Control


4) High-Risk Collision Points Selection We adopt an NMPC framework to compute optimal control
To support safe target tracking in cluttered environments, commands that enable fast, smooth, and reactive flight while
we extract a sparse set of high-risk obstacle points from the tracking the target and avoiding obstacles. At each time
completed depth map Dtcom for use in downstream model step, the NMPC solves an optimal control problem over
predictive control. These points are selected based on a a finite horizon, minimizing deviations from a dynamically
Time-To-Collision (TTC) criterion computed with respect to updated target reference while ensuring dynamic feasibility
the robot’s instantaneous velocity vT . and safety.

7
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

1) Object-Centric Modeling position:


We define the object-centric state of the quadrotor relative [pT ,x , pT ,y , 0]⊤
p̄tT = dsafe · q , (19)
to the moving target as p2T ,x + p2T ,y
 
pT
ensuring that the reference remains directionally aligned in
 vT 
 
 ∈ R13 , the plane while enforcing a safe lateral offset.
x=  (15)
qT B 
 To avoid discontinuities during tracking and to facilitate
safe takeoff behavior, this shifted goal is blended with the
ωB
robot’s current position using a time-varying ramp function:
and the control input u ∈ R4 as the vector of motor thrusts.
p̄tT = α(t) · pT + (1 − α(t)) · p̄tT , (20)
Thus, the quadrotor’s continuous-time dynamics evolve as

ṗT
 
vT
 where α(t) is a monotonically decreasing ramp function with
quadratic decay in the horizontal plane and linear decay in
 v̇T   (qT B ⊙ τ )/m + gT 
   
ẋ =  altitude. Initially, α(0) = 1, anchoring the reference to the
=
   , (16)
q̇T B   (qT B ⊙ ωB )/2 
 current position. As α(t) → 0, the reference converges to
ω˙B J−1 (µ − ωB × JωB ) the shifted goal, ensuring smooth takeoff and tracking.
 ⊤ The reference trajectory passed to the NMPC is generated
where m is the quadrotor mass, gT = 0 0 −9.81 is by replicating the same desired state over the prediction
gravity in the target frame, J = diag(Jxx , Jyy , Jzz ) is the horizon of length N :
diagonal moment of inertia matrix, and the collective thrust  t 
p̄T
τ and torque µ of the quadrotor are defined as
 0 
 
kτ l(u20 + u21 − u22 − u23 ) t
 
x̄k =  t   ∀k ∈ [0, N ), (21)

3
X q̄yaw 
τ = kτ u2i , µ = kτ l(−u20 + u21 + u22 − u23 ) . (17)
 
i=0 0
kµ (u20 − u21 + u22 − u23 )
where q̄tyaw is the desired yaw quaternion. The desired yaw
where kτ is the rotor thrust constant, kµ the rotor torque
angle ψ̄ t is obtained by correcting the current vehicle yaw
constant, and l the length of the quadrotor arm.
ψ t , extracted from the onboard quaternion qtT B , with the
To obtain a discrete-time model, we define the dynamics
yaw offset that would center the target in the image:
function f (x, u, δt) as the result of forward integration of  
the continuous-time dynamics over a fixed time step δt. x − cx
ψ̄ t = ψ t − tan−1 . (22)
Specifically, the next state is computed as fx
xt+1 = f (xt , ut , δt), (18) This yaw angle is then converted to a quaternion, assuming
zero roll and pitch:
where, in practice, this integration can be approximated using ⊤
q̄tyaw = cos ψ̄ t /2 0 0 sin ψ̄ t /2
 
numerical methods such as Euler or Runge-Kutta [67]. (23)
This state-action-dynamics formulation is commonly used
The reference control inputs are simply set to zero for
for quadrotors and provides a flexible foundation that sup-
regularizing the control effort, ūk = 0 ∀k ∈ [0, N ).
ports sophisticated dynamic models, including aerodynam-
ics, vibrations, motor interactions, and drag forces [68]–[71].
By adopting this formulation for visual target tracking, we
3) Cost Function
can leverage the extensive prior work on dynamics modeling
The NMPC cost is designed to promote accurate target
developed for traditional trajectory tracking tasks [72].
tracking, smooth control, and stable forward-facing flight.
At each time step k , the stage cost is defined as:
2 2 2
2) Reference Generation J(x, u) = ∥x − x̄k ∥Qx + ∥u − ūk ∥Qu + vyB Q , (24)
Directly commanding the quadrotor to reach the target’s | {z } | {z } | {z }o
State cost Input cost Orbiting cost
estimated position can lead to unsafe behavior, particularly
under noisy observations or occlusion. To mitigate this, we where Qx and Qu are positive-definite weighting matrices,
implement a goal-shifting strategy that maintains a minimum and Qo is a scalar that penalizes lateral motion in the
horizontal safety margin, denoted dsafe , in the xy -plane. This body frame. The term vyB is the lateral component of the
creates a cylindrical exclusion zone around the target, allow- quadrotor’s velocity in its own body frame.
ing for vertical flexibility (e.g., during takeoff or hovering) The first two terms are standard quadratic costs that drive
while maintaining lateral separation. the system toward the reference state while regularizing
The shifted reference is computed along the horizontal actuation effort. The third term addresses a fundamental
line-of-sight vector from the target to the current quadrotor observability limitation inherent to visual target tracking.

8
Specifically, when only a single target is detected, the
quadrotor’s position relative to the target is observable, but
its yaw orientation remains underconstrained. This creates
ambiguity in the rotation about the vertical axis, allowing
orbiting behaviors that preserve the perceived target location
but degrade stability and control.
While previous methods [60] resolve this by tracking mul-
tiple targets, we argue that this assumption is overly restric-
tive in real-world aerial scenarios. Maintaining visibility of
multiple targets while tracking one is extremely challenging
in cluttered, dynamic, or degraded visual environments. Figure 7: Aerial robot platform. Custom quadrotor used in
Instead, we introduce a soft regularization on vyB , exploit- experiments, equipped with an onboard computer, a stereo
ing the fact that orbiting motion correlates with lateral body RGB-D camera, and a PX4 flight controller.
velocity. By penalizing this component, the NMPC naturally
favors forward-facing, stable flight without additional per-
ceptual burden. This approach is lightweight, assumption- and safe control sequence over a finite horizon N . The
free, and effective across diverse tracking tasks. objective is to minimize a cumulative cost that promotes
accurate tracking, smooth control, and orbit-free motion:
N −1
4) Control Barrier Functions for Obstacle Avoidance
X
min J(xt+j , ut+j ) + J(xt+N , 0), (29)
To ensure safety during target tracking, we incorporate xt ,...,xt+N
j=0
second-order CBFs into the NMPC to enforce minimum ut ,...,ut+N −1

distance constraints with respect to perceived obstacles. Each where the terminal cost is evaluated with zero control input.
obstacle point stT ∈ STt is derived from depth perception The optimization is subject to the following constraints for
and expressed in the target frame T , consistent with object- all predictions j ∈ [0, N ] and safety constraints k ∈ [0, K):
centric control formulations proposed in recent works [55]. xt = x̂t , (30a)
For each high-risk point, we define a safety function: x t+1+j
= f (x t+j t+j
,u ), (30b)
t
hk (x ) = ∥stT − ptT 2
∥ − Q2max , (25) xmin ≤ x t+j
≤ xmax , (30c)
t+j
where Qmax = max(Qx , Qy ) is the minimum allowable umin ≤ u ≤ umax , (30d)
distance that accounts for the robot’s physical dimensions. ḧk (x t+j
,ut+j
) + 2λḣk (x t+j 2 t+j
) + λ hk (x ) ≥ 0. (30e)
Safety is enforced by requiring the CBF condition:
This formulation allows the NMPC to reason over future
ḧk (xt , ut ) + 2λḣk (xt ) + λ2 hk (xt ) ≥ 0, (26) trajectories while embedding safety constraints, actuation
which guarantees forward invariance of the safe set defined limits, and regularization strategies into a unified and reactive
by hk (x) ≥ 0. Here, λ > 0 controls the rate of enforcement. control framework.
The first derivative of the safety function is:
IV. EXPERIMENTAL SETUP
ḣk (xt ) = 2(stT − ptT )⊤ vTt , (27) A. Aerial System
and the second derivative, capturing the effect of control Our quadrotor platform, shown in Figure 7, weighs 1.3 kg
inputs, is given by: and has a motor span of 25 cm, with a thrust-to-weight ratio
of approximately 4 to 1. It is powered by a 6S LiPo battery
ḧk (xt , ut ) = 2∥vTt ∥2 + 2(stT − ptT )⊤ v̇T (u), (28)
and uses 1850KV motors paired with a NewBeeDrone
where v̇T (u) is the acceleration of the robot in the target Infinity200 V2 4IN1 ESC 55A to enable agile flight.
frame, obtained from the dynamics model f (xt , ut , δt). The system integrates an Intel RealSense D455 stereo
Each of these constraints is imposed at runtime for every camera for visual and depth sensing and an onboard
selected obstacle point stT , ensuring that the NMPC respects IMU. The camera is mounted front-facing, and its in-
proximity constraints while pursuing the tracking objective. trinsics are (fx , fy , cx , cy ) = (325.2, 430.9, 323.1, 246.9),
This formulation allows us to encode collision avoidance in a obtained from standard calibration [65]. The camera-to-
computationally efficient and differentiable form compatible body transformation is defined by the static quaternion
with real-time optimization. qBC = [−0.5, 0.5, −0.5, 0.5]⊤ and translation tBC =
[0.061, 0.047, −0.065]⊤ m. RGB-D frames are captured at
60 Hz with a resolution of 640 × 480.
5) Optimization Problem Onboard processing is handled by an NVIDIA Jetson
At each planning step, the NMPC solves a constrained Orin NX 16GB, and low-level stabilization is managed by a
optimal control problem to compute a dynamically feasible PixRacer Pro flight controller.

9
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

B. Target Detector
We employ a YOLOv11-small model for object detection,
trained from scratch on the COCO dataset with heavy
augmentation. Images are randomly zoomed within a scale
range of [0.6, 1.4], rotated within ±15◦ in pitch and roll, and
augmented with synthetic motion blur (30% probability) and
brightness shifts of ±20%. Crops are resized to 320×256 for
inference. The model is trained using Adam with a learning Figure 8: Targets used for tracking. All experiments use
rate of 2 × 10−4 , batch size 32, weight decay 10−5 , and for one of two object-mounted targets: a mannequin positioned
300 epochs. The inference runs at 92 Hz using TensorRT on upright on an ATV, or a stop sign mounted at the rear. The
the Orin NX with floating 16 precision. mannequin is approximately 1.8 m tall and wears typical
To initialize tracking, we use a one-time prompting stage clothing to simulate a human presence. The stop sign has
that leverages a large, high-capacity object detector. Specif- a diameter of 0.8 m, is red and reflective, and presents
ically, a YOLOv11x model [73] is run once on the full- challenges due to glare and light reflections. The ATV is used
resolution RGB image to acquire an initial bounding box for to transport the targets dynamically during experiments and
the target. This model is not trained or fine-tuned within our can reach speeds of up to 50 km/h across various terrains.
system. It is used solely for inference. Among all detections
returned, the bounding box with the highest confidence score for the position estimate, and 0.0001 for all four quaternion
is selected. This bounding box seeds the zooming module components in the orientation update.
for subsequent frames. After this initialization, the large The IMU provides linear acceleration and angular velocity
YOLOv11x model is terminated to eliminate unnecessary at 200 Hz. The UKF is also executed at this rate and
computational overhead. From that point onward, tracking handles bias estimation online without requiring separate
continues using only the lightweight onboard detector. Im- pre-calibration. All position and orientation estimates are
portantly, this prompting mechanism is modular. Any object computed in the target-centric frame described in Section 3.
detector capable of producing a coarse bounding box on For outdoor flights, a u-Blox Neo-M9N GPS module was
the initial frame can be substituted without modifying the fused with IMU data using an Extended Kalman Filter
downstream pipeline. The sole requirement is that it provides (EKF), yielding position updates at 100 Hz. This global
an early spatial prior for focusing the zoomed window. localization information was purely used for benchmarking.
The proposed system operates independently of this infor-
C. Depth Completion mation during all the flight experiments.
Depth maps from the RealSense D455 (in HighAccuracy
mode [74]) are completed using DepthAnythingV2 [64] with F. Planning and Control
a ViT-S backbone, producing per-frame estimates in 31.6 ms. We formulate an NMPC problem with a prediction horizon
A histogram bin width of 0.15 m is used for mode filtering. of N = 10 steps over 2 s and a discrete time step of δt =
0.2 s. The dynamics are integrated using a Runge–Kutta 4
D. Obstacle Selection method. The goal reference is shifted using a lateral safety
We divide the image into 17 × 17 non-overlapping cells margin of dsafe chosen based on the experiment, and blended
and select K = 10 obstacle points with the lowest TTC. with a ramp function α(t) with quadratic horizontal decay
The drone’s projected dimensions are Qx = 0.4 m and over 1 s. The desired yaw angle is computed to center the
Qy = 0.2 m. These are used in pixel-space filtering to reject target in the image, and converted to a quaternion assuming
low-risk points. The TTC map uses a forward-projected zero roll and pitch.
velocity in the camera frame and ray directions computed The NMPC cost function includes a state deviation penalty
from calibrated intrinsics. Qx = diag(150, 100, 150, 15, 15, 15, 50, 15, 15, 50, 5, 5, 5),
input cost Qu = diag(1, 1, 1, 1), and an orbiting penalty
E. State Estimation Qo = 20 on lateral body velocity. State constraints are
Target-relative state is tracked using an UKF that fuses enforced with element-wise bounds x ∈ [−999, 999]3 ×
asynchronous visual detections with high-rate inertial mea- [−25, 25]3 × [−10, 10]4 × [−40, 40]3 , and control inputs are
surements. The filter estimates the full 6-DoF state along constrained to u ∈ [0.05, 8.0]4 .
with accelerometer and gyroscope biases. Safety constraints are enforced using second-order control
The UKF process noise standard deviations are defined barrier functions with a safety radius Qmax = 0.3 m and gain
as follows: linear acceleration noise is (0.1, 0.1, 1.0) m/s2 , λ = 2. Up to 10 constraints are active per solve.
angular velocity noise is (0.2, 0.2, 0.2) rad/s, accelerometer The optimization problem is solved with acados [75],
bias noise is (0.1, 0.1, 0.1) m/s2 , and gyroscope bias noise is using SQP with Gauss–Newton Hessian approximation and
(0.1, 0.1, 0.1) rad/s. Measurement noise standard deviations Levenberg–Marquardt regularization equal to 10−2 , with a
for the visual update are 0.01 m in (x, y) and 0.001 m in z max iteration limit of 20 and a feasibility tolerance of 10−4 .

10
sensed color
completed depth

Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
4
100 12
3 1.0
velocity (m/s)

velocity (m/s)
75
distance (m)

distance (m)
10
50 2
8 0.5
25 1
0 0 dsafe
6 0.0
0 20 40 60 0 20 40 60 0 20 40 60 0 20 40 60
time (s) time (s) time (s) time (s)
Figure 9: Tracking performance in a structured urban maze. NOVA follows the target through narrow corridors and
around occlusions, maintaining stable depth estimation and obstacle awareness despite minimal texture. Bottom plots show
consistent velocity regulation and relative distance control, confirming safe and responsive tracking throughout the mission.

V. EXPERIMENTAL RESULTS throughout. No parameter tuning, no map registration, and


We evaluate NOVA in a series of real-world field experiments no external localization is used. All flight decisions are made
designed to test the system’s performance under fast motion, from raw onboard observations in real-time onboard.
degraded perception, and minimal prior knowledge. These
tests are not isolated benchmarks, but full flight missions
A. Container Maze Navigation
where the system must plan, perceive, and control entirely
We begin by evaluating NOVA in a structured yet highly con-
onboard and in real time. Our goal is to understand not only
strained setting: a container maze assembled from stacked
if NOVA can track a moving target, but how it behaves when
shipping units. The layout resembles an urban canyon,
the environment challenges each component of the stack.
with tight corridors, sharp-angle turns, and minimal escape
Specifically, we focus on three central questions: (i) Can
paths. Poles, fencing, and a wrecked vehicle introduce fixed
the system remain visually coupled to fast-moving targets
obstacles, while blind corners disrupt line of sight. The walls
across large distances and terrain types? (ii) Does the system
of the containers are flat and grey, offering little to no texture.
generalize across distinct environments, including urban,
Figure 9 shows representative sequences from one flight.
forested, and indoor–outdoor transitions, without parameter
The RGB images capture moments when the target is par-
retuning or prior knowledge? (iii) Is its behavior robust and
tially occluded or exiting a turn. The depth maps highlight
consistent across repeated trials and changes in geometry,
the system’s ability to extract consistent geometric infor-
such as altered target trajectories or forced viewpoint offsets?
mation despite the lack of strong visual cues. As the UAV
To explore these questions, we deploy NOVA in environ-
advances, it modulates its velocity continuously, slowing in
ments that progressively increase in complexity. Each sce-
tighter sections and accelerating when visibility improves.
nario stresses different components of the perception–control
The bottom plots quantify this behavior. Relative distance
stack: visual tracking under occlusion and blur, depth esti-
remains within a safe range, and the UAV’s speed adjusts
mation under lighting collapse, and planning under geomet-
tightly to match that of the ATV, even as it encounters sudden
ric constraint. The system runs with a fixed configuration
turns and obstacles.

11
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

sensed color
completed depth

Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
15 17.5 1.5
800
15.0
velocity (m/s)

velocity (m/s)
600
distance (m)

distance (m)
10 1.0
400 12.5
5 0.5
200 10.0
dsafe
0 0 7.5 0.0
0 50 100 0 50 100 0 50 100 0 50 100
time (s) time (s) time (s) time (s)
Figure 10: High-speed pursuit in a rugged forest trail. NOVA tracks the target over a 1 km unstructured path with
potholes, dust, and vegetation. Despite motion blur and strong lighting variation, detection and depth estimation remain
stable. Bottom plots confirm that the UAV sustains target speed while maintaining safe distance and minimal relative drift.

B. Forest Trail Pursuit not only keeps up but tracks with high temporal fidelity.
We now move into unstructured natural terrain, where sens- The system maintains a safe following distance with smooth
ing and control are stressed by long-range motion, fast speed regulation, navigating cluttered vegetation without
speeds, and an unpredictable environment. The forest trail hesitation or instability.
is a rugged 1 km path cutting through alternating segments
of dense canopy and open clearings. This terrain exposes
the system to rapid lighting transitions. The ground is C. Hangar Transition with Visual Ambiguity
broken and uneven, filled with potholes that jolt the ATV
We take a step further by designing a scenario where percep-
and introduce constant motion disturbances. Dust kicked
tion is pushed to its breaking point. The mission begins in an
up during traversal adds noise to the image stream, while
open gravel lot under direct sunlight, transitions into a tall
overhanging branches and narrow passageways demand fast
metallic hangar, and then returns to outdoor terrain. This path
reactive obstacle avoidance.
introduces a cascade of perceptual challenges: complete GPS
The ATV reaches speeds of up to 50 km/h, forcing the
dropout inside the hangar, abrupt lighting shifts that drive the
UAV to maintain high-speed flight while keeping visual lock,
camera from overexposure to near-total underexposure, and
estimating depth, and planning safe paths through dynamic
a cluttered interior filled with reflective surfaces, structural
obstacles. There are no structured boundaries or predictable
obstacles, and narrow doorways.
layouts here. NOVA must rely entirely on its onboard sensors
To amplify the difficulty, we place two additional man-
to adapt in real time, safely plan and track the dynamic target.
nequin targets near one of the hangar exits. These decoys
As illustrated in Figure 10, the UAV holds tight formation
are visually similar to the mannequin mounted on the ATV,
throughout the run. Despite motion blur and exposure shifts,
requiring the system to maintain target identity precisely
the system consistently detects the target and completes
during a moment when visibility is degraded and the lighting
dense depth maps that preserve usable geometry. Velocity
is unstable. This setup explicitly tests NOVA’s robustness to
plots show near-zero relative speed, indicating that NOVA
visual ambiguity, occlusion, and target switching.

12
sensed color
completed depth

Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
2.0
300 10.0
12 1.5
7.5
velocity (m/s)

velocity (m/s)
distance (m)

distance (m)
200
5.0 1.0
100 10
2.5 0.5
dsafe
0 0.0 8 0.0
0 25 50 75 0 25 50 75 0 25 50 75 0 25 50 75
time (s) time (s) time (s) time (s)
Figure 11: Robust tracking across indoor–outdoor transitions. The UAV follows the target from a gravel lot through a
tall metallic hangar and back outside, navigating abrupt lighting changes, GPS dropout, and visual sparsity. Despite exposure
shifts and structural occlusions, NOVA maintains safe distance and continuous target lock throughout the entire mission.

Such environments typically break conventional tracking UAV and the target. In this experiment, NOVA is tasked with
pipelines. Feature-based methods struggle with low texture, maintaining a constant 6 m vertical offset from the estimated
stereo matching fails under reflections, and bright-to-dark target position, resulting in a sustained flight altitude of
transitions overwhelm standard exposure control. GPS-based nearly 10 m above ground. This constraint is evaluated end
localization is unavailable inside the structure, and visual to end across a single continuous mission: from the open
odometry is easily corrupted. parking lot, through the metallic hangar, across the gravel
Despite these conditions, NOVA maintains continuous yard, and into the corridors of the simulated container maze.
tracking. As illustrated in Figure 11, RGB and completed This setup introduces a compounded set of perception and
depth frames remain functional even under poor lighting control challenges. At elevated altitudes, stereo parallax is
and low-feature geometry. When stereo cues degrade, NOVA significantly reduced, degrading the quality and density of
falls back on inertial integration and histogram-filtered depth depth estimates. The target appears smaller in the frame,
estimates to maintain a consistent relative pose. It does not limiting the number of visible features and reducing de-
require artificial markers, global maps, or external references. tector robustness. The UAV’s field of view becomes more
The trajectory plots confirm smooth flight. Relative dis- constrained with respect to lateral motion, making it harder
tance remains stable, with no abrupt control shifts or recov- to anticipate turns or sudden accelerations. Precise obstacle
ery maneuvers. The UAV transitions through the building avoidance becomes critical, especially around overhanging
without reinitialization, rejects the decoy mannequins, and structures, hangar entrances, and confined turns in the maze.
maintains accurate lock on the moving target throughout the The environments themselves amplify these difficulties.
most perceptually challenging phase of the mission. The hangar imposes complete GPS dropout and rapid light-
ing changes. The container maze presents frequent occlu-
sions, low-texture surfaces, and sharp-angle geometry. From
D. Elevated Tracking Across Mixed Terrain
a top-down viewpoint, NOVA must maintain tight target
We take an even further step to complicate the mission by
coupling while flying above clutter.
deliberately altering the spatial configuration between the

13
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

sensed color
completed depth

200 Traveled Distance XY (GPS) 8


Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) 1.5 Relative Velocity XYZ (UKF)
14
150 6 1.0
velocity (m/s)

velocity (m/s)
distance (m)

distance (m)
12
100 4
10 0.5
50 2
dsafe
0 0 8 0.0
0 20 40 0 20 40 0 20 40 0 20 40
time (s) time (s) time (s) time (s)
Figure 12: Tracking from elevated viewpoints with forced height offset. The UAV follows the target through a combined
indoor–outdoor and urban maze mission while maintaining a 6 m vertical offset. This configuration introduces degraded
stereo geometry, reduced field-of-view, and tight spatial constraints. NOVA preserves visual lock, avoids obstacles, and
regulates target distance despite elevated flight and complex terrain.

As shown in Figure 12, NOVA completes the full mission and frequent reacceleration. NOVA sustains high detection
without loss of lock or deviation from safe bounds. Despite reliability (98.6%) while maintaining an average separation
increased vertical separation and degraded sensing geometry, of 8.2 m from the target, with a minimum of 6.2 m,
it maintains coherent relative state estimates and reconstructs just above the dsafe = 6.0 m threshold. The high peak
high-fidelity depth maps. Relative distance and velocity re- acceleration of 32.4 m/s2 and max pitch of 15.0◦ reflect
main tightly regulated, and the UAV maneuvers with smooth, sharp control inputs required to navigate tight geometry
responsive control throughout. while maintaining visual lock.
This experiment confirms that NOVA generalizes not The Forest Trail scenario emphasizes sustained high-
only across different environments, but also across spatial speed tracking over unstructured terrain. The target exceeds
configurations that significantly alter perception and planning 53 km/h, and the UAV maintains a mean relative distance
dynamics. No changes are made to the underlying system; of 10.0 m, dipping to 8.4 m at its closest—well within the
all modules run as configured in prior tests, reinforcing the dsafe = 8.0 m constraint. Control demands are moderate,
adaptability of the stack under shifted tracking regimes. with limited acceleration spikes and a max pitch of 19.6◦ ,
indicating smooth flight despite blur, shadows, and environ-
mental noise. NOVA’s object detector remains effective under
E. Measured Performance across Tested Environments
these fast, noisy conditions, with 94.5% detection accuracy.
To complement the qualitative results, we report quantitative
In the Building Transition trial, the system experiences
metrics for the previous four representative experiments.
full GPS dropout and sharp illumination shifts when passing
Table 1 summarizes performance across these trials, covering
through a metallic hangar. NOVA responds with increased
GPS-based velocity, control effort, pitch dynamics, target
control activity, reaching 22.6 m/s2 peak acceleration and
distance regulation, and object detection consistency.
maintaining an average pitch of 5.3◦ . The UAV holds an
In the Urban Maze, the UAV operates within narrow corri-
average target distance of 9.9 m, with a minimum of 8.9 m,
dors and sharp turns, demanding precise obstacle avoidance

14
Table 1: Quantitative performance across the four outdoor tracking missions. We report results from representative trials
in diverse environments, each introducing unique challenges in geometry, speed, sensing, and elevation. NOVA consistently
maintains safe separation, stable flight, and high detection rates without environment-specific tuning. Reported metrics
include UAV velocity (mean and peak), linear acceleration, pitch angle, mean and minimum distance to the target compared
to the safety threshold dsafe , and overall detection rate.

GPS Velocity (km/h) Control Effort (m/s²) Pitch (°) Rel. Distance (m)
Scenario Detections (%)
Mean Max Mean Max Mean Max Mean Min dsafe

Urban Maze 5.7 14.0 10.2 32.4 2.3 15.0 8.2 6.2 6.0 98.6
Forest Trail 35.0 53.3 10.2 13.9 10.1 19.6 10.0 8.4 8.0 94.5
Building Transition 17.6 37.6 10.3 22.6 5.3 16.6 9.9 8.9 8.0 96.2
Elevated Tracking 14.5 27.5 10.3 22.8 7.7 28.8 9.4 8.8 8.0 93.1

Table 2: Tracking performance consistency across mul-


tiple indoor–outdoor trials. Metrics include mean and
minimum relative distance, user-defined safety threshold, and
relative velocity statistics. NOVA maintains stable perfor-
mance across all trials without tuning or adaptation.

Relative Distance (m) Rel. Velocity (m/s)


Trial Direction
Mean Min dsafe Mean Max

1 Forward 9.9 8.9 8.0 0.1 0.5


2 Forward 9.5 8.7 8.0 0.2 0.4
3 Reverse 9.1 8.9 8.0 0.1 0.5
4 Reverse 9.3 8.6 8.0 0.2 0.5

both within the dsafe = 8.0 m threshold, while sustaining


Figure 13: Robust target tracking across repeated trials.
visual target lock at 96.2%.
Sample RGB frames from four separate indoor–outdoor tran-
The Elevated Tracking experiment introduces an artificial
sition experiments. NOVA maintains consistent target lock
vertical offset, forcing the UAV to maintain a higher flight
despite variations in lighting, entry angle, and background.
altitude. At these distances, stereo depth estimation becomes
less reliable and target scale diminishes in the image. NOVA
compensates with stronger dynamic control, producing a
maximum pitch of 28.8◦ and acceleration up to 22.8 m/s2 .
The UAV sustains a mean distance of 9.4 m and minimum
of 8.8 m, again within the dsafe = 8.0 m limit, and direction, while two others are conducted in reverse, intro-
maintains 93.1% detection despite reduced resolution and ducing mirrored geometry, altered lighting transitions, and
more cluttered backgrounds. different entry angles.
In all cases, NOVA avoids collisions, keeps the target Despite these differences, NOVA maintains stable and
visibility, and respects user-defined separation thresholds. consistent behavior. Figure 13 presents sample onboard
These results highlight the system’s robustness across varied images from each run, illustrating reliable visual tracking
real-world conditions. under changes in viewpoint, shadowing, and background. No
manual intervention is required between trials.
F. Repeatability and Robustness Quantitative results in Table 2 confirm this consistency.
Robust operation in the real world requires more than Across all runs, the UAV maintains safe separation well be-
handling isolated challenges. A tracking system must per- low the user-defined threshold of dsafe = 8.0 m, while keep-
form consistently across repeated trials, despite minor vari- ing relative velocity and average spacing tightly bounded.
ations in conditions. To evaluate this, we repeat the full Importantly, the reversed-direction flights perform on par
indoor–outdoor transition experiment four times under nom- with the forward cases, indicating that NOVA does not
inally identical setups. Two trials proceed in the forward depend on environment-specific priors.

15
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

Urban Maze Forest Trail Building Transition Elevated Tracking


0 500 50 10
pos x (m)

50 0 0
0 10
0 25 50 0 50 100 0 50 0 20 40 3.5
0 0 0 0
pos y (m)

3.0
50 500 100
200

eph noise (m)


0 25 50 0 50 100 0 50 0 20 40 2.5
13 15 17.5
pos z (m)

12 10.0
10 15.0 2.0
11 5 7.5
12.5
0 25 50 0 50 100 0 50 0 20 40 1.5
27
25
satellites

25 20
20 1.0
20 10
26
0 25 50 0 50 100 0 50 0 20 40
time (s) time (s) time (s) time (s)
Figure 14: GPS signal degradation across outdoor tracking scenarios. The top rows show fused GPS+IMU global
position traces for each of the four outdoor missions, color-coded by estimated horizontal positional error (eph). Lighter
colors indicate higher uncertainty. The bottom plot shows the number of GPS satellites tracked over time. In the Urban
Maze and Forest Trail scenarios, signal quality remains high, with low positional error and stable satellite lock. In contrast,
the Building Transition and Elevated Tracking missions exhibit significant degradation, including extended satellite dropout
and increased eph. This effect is especially pronounced during indoor segments and when the UAV flies near structural
ceilings, highlighting the limitations of GPS-based localization in partially enclosed or cluttered environments.

G. GPS Signal Quality and Its Limitations The Elevated Tracking scenario presents even more severe
Although NOVA operates without GPS during flight, we GPS degradation. Although the trajectory includes both
use fused GPS+IMU estimates post hoc to visualize global indoor and outdoor segments, the required vertical offset
trajectories and assess how environmental geometry affects places the UAV closer to the ceiling during the hangar phase.
localization quality. This provides a comparative baseline This reduces the already limited visibility to the sky and
for understanding where GPS-based methods remain reliable exacerbates signal dropout. In several segments, the number
and where they degrade. of visible satellites falls to near zero, and position estimates
Figure 14 summarizes GPS performance across all four become highly uncertain.
representative scenarios. The top rows show the global po- These observations highlight the limits of GPS-based
sition trace, color-coded by estimated horizontal error (eph). localization in built environments. Even in outdoor set-
The bottom plot reports the number of visible satellites over tings, partial occlusion or reflective surfaces can compromise
time. Together, these metrics reflect how surrounding struc- satellite visibility and introduce substantial pose uncertainty.
tures influence satellite visibility and positioning accuracy. NOVA’s design avoids these failure modes by relying exclu-
In the Urban Maze and Forest Trail scenarios, satellite sively on onboard visual and inertial sensing, which remains
visibility remains high, typically above 20 satellites, and the operational across all tested scenarios regardless of external
estimated horizontal error stays below 1.0 m. These envi- infrastructure availability.
ronments, although partially obstructed, allow for relatively
consistent signal reception. The GPS data in these trials H. Ablation Studies
remains smooth and usable throughout the mission. To evaluate the contribution of individual components within
In contrast, the Building Transition and Elevated Offset the NOVA framework, we conduct a series of ablation
experiments exhibit significant signal degradation. During studies. Each study isolates one module and examines its
the Building Transition mission, satellite count drops sharply effect on overall system performance, while keeping the
as the UAV enters the hangar, with corresponding increases remainder of the stack fixed. The goal is to assess how spe-
in estimated error, often exceeding 3.5 m. These effects are cific design elements contribute to robustness, accuracy, and
due to occlusion, multipath reflections, and direct signal loss. safety during target tracking in unstructured environments.

16
Figure 15: Impact of Adaptive Zoom in Multi-Target Scenes. In cluttered scenes with multiple visually similar targets
(black bounding boxes), tracking can become unstable when relying on full-frame detection, often resulting in identity
switches or drift. The adaptive zoom module addresses this by cropping tightly around the prompted target, suppressing
distractors and maintaining detection focus. Shown here are drone trajectories (blue) for three separate trials, each prompted
to track a different mannequin. NOVA successfully adheres to the correct target in all cases.

Human jects. The setup includes three mannequins placed at different


positions, all resembling each other in size and appearance.
0.8 The UAV is instructed to track one specific mannequin while
confidence

the others remain in view.


0.6 Figure 15 shows representative trials. Without zoom, the
0.4 W/ Zoom detector occasionally confuses the target with a nearby
W/out Zoom distractor, leading to identity switches or temporary tracking
1.0 Stop Sign failure. With adaptive zooming enabled, the system con-
sistently maintains the correct target across all runs. This
0.8
confidence

supports the hypothesis that reducing visual clutter at the


0.6 detector input helps identify the target in ambiguous scenes.
Long-Range Detection. We further evaluate the zoom
0.4 module’s impact on detection reliability at extended dis-
tances. Two static targets, a human and a stop sign, are
2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0
distance (m) placed in an open indoor space. The robot is handheld and
moved toward and away from each target, while the onboard
Figure 16: Effect of Adaptive Zooming on Detection Con- perception stack is run twice: once with zooming enabled
fidence vs. Distance. Detection confidence for a human and and once with full-resolution images passed directly to the
a stop sign is compared with and without the adaptive zoom detector. Detection confidence is recorded for each case.
module. The robot was handheld toward and away from Figure 16 summarizes the results. Without zoom, the
fixed targets in an open indoor space, and the perception detector’s confidence drops rapidly beyond 10–12 m, often
pipeline was run in both modes. Without zooming, detection failing to register the target altogether. In contrast, the zoom
confidence degrades sharply beyond 10–12 m due to visual module enables consistent detections up to the camera’s
clutter and scale compression. The zooming strategy, by maximum effective range of 20–25 m. By cropping out
focusing the input on the target and cropping out irrelevant irrelevant regions and rescaling the image around the target,
context, maintains high confidence and extends detection zooming mitigates scale compression and distractor interfer-
range to the camera’s depth limit (> 20 m). ence, boosting confidence and range.
Together, these findings highlight that adaptive zoom
contributes to robust tracking in two key regimes: it improves
1) Adaptive Zoom Strategy resilience to distractors in multi-object environments and
The adaptive zoom module improves detection robustness extends detection range under scale-constrained conditions.
by dynamically cropping the image around the expected
target location before passing it to the detector. This strategy
is designed to address two key challenges: (i) suppressing 2) Depth Completion
visually similar distractors that can trigger identity switches, The raw stereo depth output from the robot’s onboard sensor
and (ii) preserving the detector’s ability to recognize small, is often sparse and unreliable, particularly in low-texture
distant targets by maintaining spatial resolution over the regions, near thin structures, or under degraded lighting.
region of interest. These limitations reduce the effectiveness of downstream
Multi-Target Robustness. We first assess the role of obstacle avoidance and planning components, especially in
adaptive zoom in scenes with multiple visually similar ob- fast or cluttered environments.

17
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

Figure 18: Robustness of Depth Estimation Methods


During Occlusion. A cart with an obstacle occludes the
Figure 17: Qualitative Results of Depth Completion. Top: target mid-experiment, inducing background intrusion in
RGB inputs. Middle: Raw stereo depth from the onboard the bounding box. The plot shows depth estimates over
sensor, which suffers from missing or noisy regions, es- time using three strategies: center-pixel, mean-pixel, and our
pecially around thin structures, glass, and foliage. Bottom: histogram-based mode filtering. Ground truth is manually
Completed depth maps produced by our fusion module. The measured. The mean estimator exhibits large spikes dur-
system recovers fine details, such as railings and branches, ing occlusion, and the center-pixel fails intermittently. The
that are critical for collision avoidance but often missed histogram-based method remains stable and close to ground
by stereo matching alone. This enables safe navigation in truth throughout.
visually complex environments.

The ground truth distance to the target is measured


To address this challenge, NOVA incorporates a disparity- manually for reference. As shown in Figure 18, the center-
based depth completion module. It leverages monocular pixel approach produces unstable estimates during occlusion,
priors and disparity cues to fill in missing or noisy depth while the mean-pixel method exhibits a consistent bias
regions, producing denser and smoother maps. Figure 17 toward the background. In contrast, the histogram-based
presents qualitative comparisons between the raw stereo strategy maintains a stable and accurate estimate throughout
maps and the completed depth output across several rep- the occlusion event, closely matching the ground truth.
resentative scenarios. The completed maps better preserve This improved robustness in depth estimation directly
obstacle geometry, enable earlier obstacle detection, and enhances control performance, enabling more stable tracking
improve motion planning and control safety. under visual uncertainty and in cluttered environments.

VI. LIMITATIONS AND FUTURE WORKS


3) Histogram-Based Mode Filtering NOVA operates under two key assumptions that currently
Accurate target localization requires reliable estimation of define its functional scope and deployment regime.
depth within the predicted bounding box. However, raw First, the system assumes that the target is initially visible
depth values are often noisy and may include background within the field of view of the onboard camera and belongs
clutter, particularly in dynamic scenes with partial occlu- to a known object category supported by the detector. This
sion. To address this, we introduce a histogram-based mode assumption is valid in prompted missions or pre-configured
filtering approach that selects the most frequent depth bin tracking scenarios, but limits applicability in open-world set-
within the bounding box, offering a robust estimate that is tings where targets may enter the scene later, from arbitrary
resilient to outliers and background interference. directions, or lack a pre-defined semantic label. Addressing
To evaluate this method, we conduct an experiment where this constraint would require extending the system to han-
the robot hovers in front of a stationary target. A cart carrying dle open-set detection and tracking, where targets are not
a vertical pole is then moved horizontally between the robot assumed to belong to a fixed category and may be specified
and the target, creating a temporary occlusion. During the through visual prompts, user feedback, or learned embed-
sequence, we compare three depth estimation strategies: dings [76], [77]. Recent approaches in category-agnostic
reading the depth value at the center pixel of the bounding tracking and self-supervised object discovery offer promising
box, averaging all valid depth pixels within the box, and directions for relaxing this assumption and enabling more
applying our proposed histogram-mode filtering. flexible engagement with previously unseen targets [47].

18
Second, the system relies on the ability to estimate the viewpoint offsets, repeatability over multiple trials, and ro-
target’s depth using stereo-based depth completion. While bustness under degraded GPS and perceptual ambiguity.
this approach is sufficient at close to mid-range (typically Ablation studies confirmed the contribution of individual
up to 30–40 m), it breaks down when the target is detected components, such as adaptive zooming for long-range detec-
at longer distances and stereo cues become unreliable. In tion and identity preservation. The system’s limitations were
these conditions, detection may still succeed due to the zoom also analyzed, including assumptions about initial visibility,
module, but the depth estimates become noisy or flat, making semantic category constraints, and the reliance on stereo-
the x-axis velocity (change in relative depth) unobservable. based depth at range.
As a result, the UAV may accelerate rapidly toward the Together, these results support the conclusion that robust,
target with limited feedback. In flight experiments, the robot real-time target tracking can be achieved using only onboard
eventually stabilizes once the depth becomes consistent, but sensing, even in the absence of external infrastructure or
the initial approach is often jerky and fast, reflecting the lack structured environments. Future works will focus on relaxing
of velocity feedback in the unobservable regime. current assumptions, extending to open-set target represen-
To relax the reliance on stereo depth, several alternative tations, and enabling perception-driven control under long-
strategies could be explored. One approach is to augment range uncertainty.
the sensing stack with additional depth-aware sensors, such
as radar or lightweight range finders, which are less sensi- References
tive to texture and lighting and can extend the operational [1] S. Wu, R. Li, Y. Shi, and Q. Liu, “Vision-based target detection and
range [78]–[80] . Another is to infer short-term relative tracking system for a quadcopter,” IEEE Access, vol. 9, pp. 62 043–
motion using optical flow, either from dense image features 62 054, 2021.
[2] M. Xu, A. Hu, and H. Wang, “Visual-impedance-based human–robot
or from bounding box displacements of detected objects, cotransportation with a tethered aerial vehicle,” IEEE Transactions on
including in open-set configurations [81]. A third direction Industrial Informatics, vol. 19, no. 10, pp. 10 356–10 365, 2023.
involves registering successive point clouds generated from [3] A. Hu, M. Xu, H. Wang, and H. Castañeda, “Vision-based impedance
control of an aerial manipulator using a nonlinear observer,” IEEE
depth completion, even if sparse or noisy, to recover frame- Transactions on Automation Science and Engineering, vol. 20, no. 2,
to-frame translation [82]. These approaches would allow for pp. 1441–1451, 2022.
estimation of instantaneous robot velocity without relying on [4] J. Gu, T. Su, Q. Wang, X. Du, and M. Guizani, “Multiple moving
targets surveillance based on a cooperative network for multi-uav,”
world-frame position or persistent maps, and could remain IEEE Communications Magazine, vol. 56, no. 4, pp. 82–89, 2018.
robust under partial feature disruption due to their minimal [5] N. Bashir, S. Boudjit, and S. Zeadally, “A closed-loop control archi-
or stateless memory requirements. tecture of uav and wsn for traffic surveillance on highways,” Computer
These assumptions are not intrinsic limitations of the Communications, vol. 190, pp. 78–86, 2022.
[6] H. Huang, A. V. Savkin, and W. Ni, “Online uav trajectory planning
overall framework, but they define the boundaries of the cur- for covert video surveillance of mobile targets,” IEEE Transactions
rent implementation. Future work targeting open-set object on Automation Science and Engineering, vol. 19, no. 2, pp. 735–746,
understanding and depth-agnostic motion estimation could 2021.
[7] L. Quan, L. Yin, T. Zhang, M. Wang, R. Wang, S. Zhong, Y. Cao,
expand NOVA’s applicability to more uncertain, long-range, C. Xu, and F. Gao, “Formation flight in dense environments,” CoRR,
or semantically unconstrained tracking scenarios. 2022.
[8] G. A. Di Caro and A. W. Z. Yousaf, “Multi-robot informative path
planning using a leader-follower architecture,” in IEEE International
Conference on Robotics and Automation, 2021, pp. 10 045–10 051.
[9] T. Miki, P. Khrapchenkov, and K. Hori, “Uav/ugv autonomous co-
VII. CONCLUSION operation: Uav assists ugv to climb a cliff by attaching a tether,” in
We presented NOVA, an onboard visual-inertial framework IEEE International Conference on Robotics and Automation, 2019, pp.
for agile target tracking in unstructured and GPS-denied 8041–8047.
[10] G. Niu, Q. Yang, Y. Gao, and M.-O. Pun, “Vision-based autonomous
environments. The system avoids reliance on external local- landing for unmanned aerial and ground vehicles cooperative systems,”
ization, global maps, or precomputed scene priors. Instead, IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 6234–6241,
it formulates perception, estimation, and control directly in 2021.
[11] M. Demirhan and C. Premachandra, “Development of an auto-
the target’s reference frame, using only stereo vision and in- mated camera-based drone landing system,” IEEE Access, vol. 8, pp.
ertial sensing. The approach combines a lightweight detector 202 111–202 121, 2020.
with adaptive zooming, depth completion, and visual-inertial [12] P. Vlantis, P. Marantos, C. P. Bechlioulis, and K. J. Kyriakopoulos,
“Quadrotor landing on an inclined platform of a moving ground vehi-
state estimation, followed by a NMPC that operates under cle,” in IEEE International Conference on Robotics and Automation,
collision-aware constraints derived from onboard sensing. 2015, pp. 2202–2207.
We validated NOVA across a series of real-world trials [13] F. Chaumette and S. Hutchinson, “Visual servo control. ii. ad-
that stress the system across different terrain types, motion vanced approaches [tutorial],” IEEE Robotics & Automation Magazine,
vol. 14, no. 1, pp. 109–118, 2007.
regimes, and sensory conditions. The system demonstrated [14] S. Hutchinson, G. D. Hager, and P. I. Corke, “A tutorial on visual servo
consistent performance in forested, urban, and mixed in- control,” IEEE transactions on robotics and automation, vol. 12, no. 5,
door–outdoor environments, maintaining visual lock, re- pp. 651–670, 1996.
[15] S. Cho and D. H. Shim, “Sampling-based visual path planning
specting safety distances, and operating without manual tun- framework for a multirotor uav,” International Journal of Aeronautical
ing. Additional experiments evaluated generalization across and Space Sciences, vol. 20, pp. 732–760, 2019.

19
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments

[16] A. A. Oliva, E. Aertbeliën, J. De Schutter, P. R. Giordano, and ing,” Journal of Marine Science and Engineering, vol. 10, no. 3, p.
F. Chaumette, “Towards dynamic visual servoing for interaction con- 383, 2022.
trol and moving targets,” in IEEE International conference on robotics [37] W. Zhang, K. Song, X. Rong, and Y. Li, “Coarse-to-fine uav target
and automation, 2022, pp. 150–156. tracking with deep reinforcement learning,” IEEE Transactions on
[17] S. Raj, P. R. Giordano, and F. Chaumette, “Appearance-based indoor Automation Science and Engineering, vol. 16, no. 4, pp. 1522–1530,
navigation by ibvs using mutual information,” in 14th International 2018.
Conference on Control, Automation, Robotics and Vision, 2016, pp. [38] C. Min, S. Si, X. Wang, H. Xue, W. Jiang, Y. Liu, J. Wang, Q. Zhu,
1–6. Q. Zhu, L. Luo, et al., “Autonomous driving in unstructured environ-
[18] P. Geneva, K. Eckenhoff, W. Lee, Y. Yang, and G. Huang, “Openvins: ments: How far have we come?” arXiv preprint arXiv:2410.07701,
A research platform for visual-inertial estimation,” in IEEE Interna- 2024.
tional Conference on Robotics and Automation, 2020, pp. 4666–4672. [39] R. Khanam and M. Hussain, “Yolov11: An overview of the key
[19] T. Qin, P. Li, and S. Shen, “Vins-mono: A robust and versatile monoc- architectural enhancements,” arXiv preprint arXiv:2410.17725, 2024.
ular visual-inertial state estimator,” IEEE transactions on robotics, [40] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and
vol. 34, no. 4, pp. 1004–1020, 2018. J. Chen, “Detrs beat yolos on real-time object detection,” in IEEE/CVF
[20] D. Scaramuzza and Z. Zhang, “Visual-inertial odometry of aerial conference on computer vision and pattern recognition, 2024, pp.
robots,” arXiv preprint arXiv:1906.03289, 2019. 16 965–16 974.
[21] M. Labbé and F. Michaud, “Rtab-map as an open-source lidar and [41] Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu,
visual simultaneous localization and mapping library for large-scale and X. Wang, “Bytetrack: Multi-object tracking by associating every
and long-term online operation,” Journal of field robotics, vol. 36, detection box,” in European conference on computer vision. Springer,
no. 2, pp. 416–446, 2019. 2022, pp. 1–21.
[22] W. G. Aguilar, G. A. Rodrı́guez, L. Álvarez, S. Sandoval, F. Quis- [42] N. Aharon, R. Orfaig, and B.-Z. Bobrovsky, “Bot-sort: Robust asso-
aguano, and A. Limaico, “Visual slam with a rgb-d camera on a ciations multi-pedestrian tracking,” arXiv preprint arXiv:2206.14651,
quadrotor uav using on-board processing,” in Advances in Computa- 2022.
tional Intelligence: 14th International Work-Conference on Artificial [43] T. Cheng, L. Song, Y. Ge, W. Liu, X. Wang, and Y. Shan, “Yolo-world:
Neural Networks, IWANN 2017, Cadiz, Spain, June 14-16, 2017, Real-time open-vocabulary object detection,” in IEEE/CVF Conference
Proceedings, Part II 14. Springer, 2017, pp. 596–606. on Computer Vision and Pattern Recognition, 2024, pp. 16 901–16 911.
[23] R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: A [44] A. Wang, L. Liu, H. Chen, Z. Lin, J. Han, and G. Ding, “Yoloe:
versatile and accurate monocular slam system,” IEEE transactions on Real-time seeing anything,” arXiv preprint arXiv:2503.07465, 2025.
robotics, vol. 31, no. 5, pp. 1147–1163, 2015. [45] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr,
[24] Y. Wang and A. Zell, “Improving feature-based visual slam by seman- R. Rädle, C. Rolland, L. Gustafson, et al., “Sam 2: Segment anything
tics,” in International Conference on Image Processing, Applications in images and videos,” in The Thirteenth International Conference on
and Systems (IPAS), 2018, pp. 7–12. Learning Representations, 2025.
[25] T. Tzanetos, M. Aung, J. Balaram, H. F. Grip, J. T. Karras, T. K.
[46] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li,
Canham, G. Kubiak, J. Anderson, G. Merewether, M. Starch, et al.,
J. Yang, H. Su, et al., “Grounding dino: Marrying dino with grounded
“Ingenuity mars helicopter: From technology demonstration to ex-
pre-training for open-set object detection,” in European Conference on
traterrestrial scout,” in aerospace conference (AERO), 2022, pp. 01–19.
Computer Vision. Springer, 2024, pp. 38–55.
[26] J. Balaram, M. Aung, and M. P. Golombek, “The ingenuity helicopter
[47] A. Saviolo, P. Rao, V. Radhakrishnan, J. Xiao, and G. Loianno,
on the perseverance rover,” Space Science Reviews, vol. 217, no. 4,
“Unifying foundation models with quadrotor control for visual track-
p. 56, 2021.
ing beyond object categories,” in IEEE International Conference on
[27] S. Withrow, W. Johnson, L. A. Young, H. Cummings, J. Balaram, and
Robotics and Automation, 2024, pp. 7389–7396.
T. Tzanetos, “An advanced mars helicopter design,” in ASCEND 2020,
2020, p. 4028. [48] S. Hu, Q. Wang, F. Wang, and Y. Li, “Finite-time dynamic visual servo
[28] H. F. Grip, D. Conway, J. Lam, N. Williams, M. P. Golombek, control for quadrotor tracking unknown motion target,” Nonlinear
R. Brockers, M. Mischna, and M. R. Cacan, “Flying a helicopter on Dynamics, vol. 113, no. 7, pp. 6959–6977, 2025.
mars: How ingenuity’s flights were planned, executed, and analyzed,” [49] Y. Kumar, S. B. Roy, et al., “Adaptive ibvs based planar non-
in Aerospace Conference, 2022, pp. 1–17. holonomic target tracking for quadrotors,” in International Conference
[29] K. Zhang, Y. Shi, and H. Sheng, “Robust nonlinear model predictive on Unmanned Aircraft Systems, 2024, pp. 201–208.
control based visual servoing of quadrotor uavs,” IEEE/ASME Trans- [50] M. Leomanni, F. Ferrante, A. Dionigi, G. Costante, P. Valigi, and M. L.
actions on Mechatronics, vol. 26, no. 2, pp. 700–708, 2021. Fravolini, “Quadrotor control system design for robust monocular
[30] D. Guo and K. K. Leang, “Image-based estimation, planning, and visual tracking,” IEEE Transactions on Control Systems Technology,
control for high-speed flying through multiple openings,” The Inter- 2024.
national Journal of Robotics Research, vol. 39, no. 9, pp. 1122–1137, [51] Y. Jiang, H. Wang, and W. Yu, “Perception-aware model predictive
2020. control for target tracking with uavs,” in 14th Asian Control Confer-
[31] P. Serra, R. Cunha, T. Hamel, D. Cabecinhas, and C. Silvestre, ence, 2024, pp. 998–1003.
“Landing of a quadrotor on a moving target using dynamic image- [52] A. Altan and R. Hacıoğlu, “Model predictive control of three-axis
based visual servo control,” IEEE Transactions on Robotics, vol. 32, gimbal system mounted on uav for real-time target tracking under
no. 6, pp. 1524–1535, 2016. external disturbances,” Mechanical Systems and Signal Processing,
[32] D. Zheng, H. Wang, J. Wang, S. Chen, W. Chen, and X. Liang, “Image- vol. 138, p. 106548, 2020.
based visual servoing of a quadrotor using virtual camera approach,” [53] A. H. González and D. Odloak, “Robust model predictive controller
IEEE/ASME Transactions on Mechatronics, vol. 22, no. 2, pp. 972– with output feedback and target tracking,” IET Control Theory &
982, 2016. Applications, vol. 4, no. 8, pp. 1377–1390, 2010.
[33] J. Thomas, G. Loianno, K. Daniilidis, and V. Kumar, “Visual servoing [54] N. Sugie, “A model of predictive control in visual target tracking,”
of quadrotors for perching by hanging from cylindrical objects,” IEEE IEEE Transactions on Systems, Man, and Cybernetics, no. 1, pp. 2–7,
robotics and automation letters, vol. 1, no. 1, pp. 57–64, 2015. 2010.
[34] B. Ma, Z. Liu, W. Zhao, J. Yuan, H. Long, X. Wang, and Z. Yuan, [55] A. Saviolo, N. Picello, J. Mao, R. Verma, and G. Loianno, “Re-
“Target tracking control of uav through deep reinforcement learning,” active collision avoidance for safe agile navigation,” arXiv preprint
IEEE Transactions on Intelligent Transportation Systems, vol. 24, arXiv:2409.11962, 2024.
no. 6, pp. 5983–6000, 2023. [56] T. Do, L. C. Carrillo-Arce, and S. I. Roumeliotis, “High-speed
[35] A. Dionigi, M. Leomanni, A. Saviolo, G. Loianno, and G. Costante, autonomous quadrotor navigation through visual and inertial paths,”
“Exploring deep reinforcement learning for robust target tracking using The International Journal of Robotics Research, vol. 38, no. 4, pp.
micro aerial vehicles,” in 21st International Conference on Advanced 486–504, 2019.
Robotics, 2023, pp. 506–513. [57] A. Loquercio, A. Saviolo, and D. Scaramuzza, “Autotune: Controller
[36] Y. Mao, F. Gao, Q. Zhang, and Z. Yang, “An auv target-tracking tuning for high-speed flight,” IEEE Robotics and Automation Letters,
method combining imitation learning and deep reinforcement learn- vol. 7, no. 2, pp. 4432–4439, 2022.

20
[58] M. Kulkarni, B. Moon, K. Alexis, and S. Scherer, “Aerial field [82] M. Lyu, J. Yang, Z. Qi, R. Xu, and J. Liu, “Rigid pairwise 3d point
robotics,” arXiv preprint arXiv:2401.10837, 2024. cloud registration: A survey,” Pattern Recognition, p. 110408, 2024.
[59] B. Lindqvist, S. S. Mansouri, A.-a. Agha-mohammadi, and G. Niko-
lakopoulos, “Nonlinear mpc for collision avoidance and control of
uavs with dynamic obstacles,” IEEE Robotics and Automation Letters,
vol. 5, no. 4, pp. 6001–6008, 2020.
[60] Y. Li, G. Lu, D. He, and F. Zhang, “Robocentric model-based
visual servoing for quadrotor flights,” IEEE/ASME Transactions on
Mechatronics, vol. 28, no. 4, pp. 2155–2166, 2023.
[61] B. Zhang, X. Chen, Z. Li, G. Beltrame, C. Xu, F. Gao, and Y. Cao,
“Coni-mpc: Cooperative non-inertial frame based model predictive
control,” IEEE Robotics and Automation Letters, vol. 8, no. 12, pp.
8082–8089, 2023.
[62] B. Zhang, X. Chen, Q. Chen, C. Xu, F. Gao, and Y. Cao, “Global-
state-free obstacle avoidance for quadrotor control in air-ground co-
operation,” IEEE Robotics and Automation Letters, 2025.
[63] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan,
P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in
context,” in European Conference on Computer Vision. Springer,
2014, pp. 740–755.
[64] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao,
“Depth anything v2,” arXiv preprint arXiv:2406.09414, 2024.
[65] L. Oth, P. Furgale, L. Kneip, and R. Siegwart, “Rolling shutter camera
calibration,” in IEEE Conference on Computer Vision and Pattern
Recognition, 2013, pp. 1360–1367.
[66] G. Loianno, M. Watterson, and V. Kumar, “Visual inertial odometry for
quadrotors on se (3),” in IEEE International Conference on Robotics
and Automation, 2016, pp. 1544–1551.
[67] J. H. Cartwright and O. Piro, “The dynamics of runge–kutta methods,”
International Journal of Bifurcation and Chaos, vol. 2, no. 03, pp.
427–449, 1992.
[68] A. Saviolo, J. Frey, A. Rathod, M. Diehl, and G. Loianno, “Active
learning of discrete-time dynamics for uncertainty-aware model pre-
dictive control,” IEEE Transactions on Robotics, 2023.
[69] A. Saviolo, G. Li, and G. Loianno, “Physics-inspired temporal learn-
ing of quadrotor dynamics for accurate model predictive trajectory
tracking,” IEEE Robotics and Automation Letters, 2022.
[70] L. Bauersfeld, E. Kaufmann, P. Foehn, S. Sun, and D. Scaramuzza,
“NeuroBEM: Hybrid Aerodynamic Quadrotor Model,” Robotics: Sci-
ence and Systems Foundation, 2021.
[71] F. Crocetti, J. Mao, A. Saviolo, G. Costante, and G. Loianno, “Gapt:
Gaussian process toolkit for online regression with application to
learning quadrotor dynamics,” arXiv preprint arXiv:2303.08181, 2023.
[72] A. Saviolo and G. Loianno, “Learning quadrotor dynamics for precise,
safe, and agile flight control,” Annual Reviews in Control, 2023.
[73] G. Jocher, J. Qiu, and A. Chaurasia, “Ultralytics YOLO,” Jan. 2023.
[Online]. Available: [Link]
[74] A. Grunnet-Jepsen, J. N. Sweetser, and J. Woodfill, “Best-known-
methods for tuning intel® realsense™ d400 depth cameras for best
performance,” Intel Corporation: Satan Clara, CA, USA, vol. 1, 2018.
[75] R. Verschueren, G. Frison, D. Kouzoupis, J. Frey, N. van Duijkeren,
A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, and M. Diehl,
“acados – a modular open-source framework for fast embedded
optimal control,” Mathematical Programming Computation, 2021.
[76] A. Maalouf, N. Jadhav, K. M. Jatavallabhula, M. Chahine, D. M. Vogt,
R. J. Wood, A. Torralba, and D. Rus, “Follow anything: Open-set
detection, tracking, and following in real-time,” IEEE Robotics and
Automation Letters, vol. 9, no. 4, pp. 3283–3290, 2024.
[77] Y. Liu, I. E. Zulfikar, J. Luiten, A. Dave, D. Ramanan, B. Leibe,
A. Ošep, and L. Leal-Taixé, “Opening up open world tracking,” in
IEEE/CVF conference on computer vision and pattern recognition,
2022, pp. 19 045–19 055.
[78] M. Nissov, N. Khedekar, and K. Alexis, “Degradation resilient
lidar-radar-inertial odometry,” in IEEE International Conference on
Robotics and Automation, 2024, pp. 8587–8594.
[79] B. Kim, M. B. Azhari, J. Park, and D. H. Shim, “An autonomous
uav system based on adaptive lidar inertial odometry for practical
exploration in complex environments,” Journal of Field Robotics,
vol. 41, no. 3, pp. 669–698, 2024.
[80] J. Zhang and S. Singh, “Low-drift and real-time lidar odometry and
mapping,” Autonomous robots, vol. 41, pp. 401–416, 2017.
[81] A. Alfarano, L. Maiano, L. Papa, and I. Amerini, “Estimating optical
flow: A comprehensive review of the state of the art,” Computer Vision
and Image Understanding, p. 104160, 2024.

21

You might also like