NOVA: High-Speed UAV Target Tracking
NOVA: High-Speed UAV Target Tracking
ABSTRACT Autonomous aerial target tracking in unstructured and GPS-denied environments remains a
fundamental challenge in robotics. Many existing methods rely on motion capture systems, pre-mapped
scenes, or feature-based localization to ensure safety and control, limiting their deployment in real-world
conditions. We introduce NOVA, a fully onboard, object-centric framework that enables robust target
tracking and collision-aware navigation using only a stereo camera and an IMU. Rather than constructing
a global map or relying on absolute localization, NOVA formulates perception, estimation, and control
entirely in the target’s reference frame. A tightly integrated stack combines a lightweight object detector
with stereo depth completion, followed by histogram-based filtering to infer robust target distances under
occlusion and noise. These measurements feed a visual-inertial state estimator that recovers the full 6-DoF
pose of the robot relative to the target. A nonlinear model predictive controller (NMPC) plans dynamically
feasible trajectories in the target frame. To ensure safety, high-order control barrier functions (CBFs) are
constructed online from a compact set of high-risk collision points extracted from depth, enabling real-
time obstacle avoidance without maps or dense representations. We validate NOVA across challenging
real-world scenarios, including urban mazes, forest trails, and repeated transitions through buildings with
intermittent GPS loss and severe lighting changes that disrupt feature-based localization. Each experiment
is repeated multiple times under similar conditions to assess resilience, showing consistent and reliable
performance. NOVA achieves agile target following at speeds exceeding 50 km/h. These results show that
high-speed, vision-based tracking is possible in the wild using only onboard sensing, with no reliance on
external localization or assumptions on the environment structure. Video: [Link]
INDEX TERMS Target tracking, Aerial robotics, Vision-based navigation, Model predictive control, Control
barrier functions, Visual-inertial odometry, Depth completion, Object detection, GPS-denied environments.
1
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
Figure 1: Satellite imagery of representative NOVA flight missions in real-world environments. Each overlaid UAV
trajectory illustrates a distinct experimental scenario designed to test specific failure modes in visual-inertial target tracking
and control. Red: High-speed tracking in a forest trail over 1 km, with target speeds exceeding 50 km/h and motion blur.
Blue: Indoor–outdoor transition with complete GPS loss, exposure collapse, and perceptual ambiguity due to structural
occlusions and distractors. Orange: Elevated tracking from a 10 m offset, challenging stereo depth perception, feature
consistency, and planning geometry. Green: Urban container maze with narrow corridors, cluttered turns, and minimal
texture. Insets show raw onboard RGB frames during each flight. NOVA maintains real-time target lock, depth estimation,
obstacle avoidance, and closed-loop control using only a stereo camera and IMU. GPS is not used by NOVA and appears
here only for visualization of flight paths.
is essential. UAVs in leader–follower formations or aerial ambiguous semantics, and degraded visibility. As a result,
swarms must track teammates to maintain coordinated mo- feature-based localization quickly becomes unreliable [24].
tion [7]–[9]. In autonomous landing, a UAV must descend A notable example of this failure was NASA’s Inge-
onto a moving platform, requiring precise tracking [10]–[12]. nuity Mars Helicopter [25]–[27], which suffered a navi-
Visual tracking methods for UAVs traditionally fall into gation failure while flying over smooth, featureless sand
two categories: Image-Based Visual Servoing (IBVS) and ripples [28]. Without distinctive landmarks, its visual-inertial
Position-Based Visual Servoing (PBVS) [13]–[17]. IBVS op- system produced erroneous velocity estimates, leading to a
erates directly on image-plane measurements, allowing low- crash 20 seconds after takeoff. Had Ingenuity been able to
latency control and robustness to calibration errors. However, localize relative to a known object in the scene, such as the
its effectiveness is limited by the camera’s field of view and Perseverance rover, it might have maintained stable flight and
sensitivity to target occlusion or rapid motion. PBVS instead avoided the crash. This example highlights a broader insight:
reconstructs the target’s 3D position, typically in a globally when a known and observable target is available, it can serve
referenced frame, and uses it to guide control. While PBVS as a reliable frame of reference than the surrounding terrain.
mitigates many of IBVS’s limitations, it introduces a new This paper introduces NOVA, a unified framework for
dependency: accurate consistent global localization. Navigation via Object-centric Visual Autonomy. NOVA en-
In GPS-denied settings such as forests, urban mazes, ables real-time aerial target tracking and obstacle-aware
indoor spaces, or extraterrestrial terrain, global localization navigation using only onboard sensing, with no dependence
typically relies on Visual-Inertial Odometry (VIO) [18]–[20] on GPS, external maps, or motion capture. Unlike traditional
or Simultaneous Localization and Mapping (SLAM) [21]– VIO or SLAM-based methods that require globally consis-
[23], which estimate the robot’s pose by tracking visual tent localization, NOVA formulates perception, estimation,
features across image sequences. These methods assume and control directly in the target’s frame of reference. This
sufficient texture, consistent lighting, and a relatively static object-centric approach enables tracking in environments
environment to maintain reliable feature correspondences. where localization is unreliable or unavailable.
However, unstructured environments often violate these as- At the core of NOVA is a tightly integrated perception-
sumptions due to cluttered geometry, dynamic elements, to-control stack that couples a custom object detector, depth
completion module, and visual-inertial state estimator. The
2
object detector operates in real time on low-resolution inputs to movement. This method is efficient and effective for tasks
using an adaptive zoom strategy to maintain tracking at long where the target remains in view and image-space motion
ranges. Depth estimates are computed using a fusion of corresponds smoothly to inputs.
stereo and monocular cues, then processed via histogram- PBVS, on the other hand, estimates the full 3D pose of
based filtering to produce robust target distance estimates, the target using stereo vision, depth sensors, or geometric
even under occlusion or noise. From this, a full 6-DoF state priors, and computes control actions in Euclidean space. It
of the robot relative to the target is inferred in real time. aims to minimize the spatial error between the current and
Crucially, this perception system is designed to operate under desired poses of the robot relative to the target. This spatial
realistic sensing conditions, such as degraded lighting, partial reasoning improves robustness when the target leaves the
occlusions, and unstructured terrain. field of view. However, PBVS depends on accurate global
These target-centric estimates are passed to a Nonlinear pose estimates of both the robot and the target, requiring
Model Predictive Controller (NMPC) that optimizes dynam- consistent global localization across frames. In GPS-denied
ically feasible trajectories directly in the target’s reference or visually degraded environments, this dependency often
frame. To ensure safe operation in cluttered environments, becomes a critical point of failure.
obstacle avoidance is enforced online using high-order Con- While both paradigms are effective in structured, static en-
trol Barrier Functions (CBFs), derived from a compact set vironments [29]–[33], their assumptions often break down in
of high-risk collision points identified in the depth map. aerial target tracking. UAVs operate under rapid motion and
This eliminates the need for dense mapping or environmental wide viewpoint changes, which distort target appearance and
priors, enabling real-time, map-free collision avoidance. invalidate interaction models. During aggressive maneuvers,
We validate NOVA across a suite of challenging real-world UAVs experience frequent occlusions, motion blur, and target
experiments, illustrated in Figure 1. Each scenario targets a scale variation, causing visual features to become unreliable.
specific failure mode common to traditional visual tracking IBVS fails when the target exits the frame, undergoes large
pipelines. In the forest trail mission, the UAV follows a appearance shifts, or when image-space motion no longer
moving target at speeds exceeding 50 km/h over more than correlates predictably with control. PBVS struggles in GPS-
1 km, inducing severe motion blur and visual degradation denied environments or texture-sparse scenes where global
from airborne dust. The building transition scenario involves pose estimation is not available or stable.
abrupt lighting changes and full GPS loss as the target moves These challenges are further exacerbated by the limited
from an open parking lot into an enclosed space and back computational resources available onboard UAVs and the
outdoors. In the height-offset trial, the UAV maintains a stringent real-time requirements for flight control. In prac-
vertical height of 10 m, challenging feature matching due tice, neither classical IBVS nor PBVS alone can reliably
to reduced parallax and resolution. handle the demands of high-speed, long-range target tracking
Across all experiments, NOVA maintains stable target in unstructured, dynamic environments.
lock, plans feasible trajectories, and avoids obstacles in real
time. To evaluate repeatability, each mission is conducted
B. Data-driven Aerial Target Tracking
multiple times under similar conditions, with consistent per-
formance observed throughout. All sensing and computation Recent approaches in aerial target tracking have leveraged
are performed fully onboard using only a stereo camera the power of deep learning to enhance perception and
and an IMU. Unlike prior methods that rely on motion decision-making under challenging conditions. One promi-
capture, artificial landmarks, or pre-built maps, NOVA op- nent direction is the use of end-to-end learning, where a
erates autonomously in unstructured environments with no neural network directly maps raw sensory inputs to low-
external infrastructure. To the best of our knowledge, this level control actions. These models, often trained via rein-
is the first demonstration of real-time, high-speed object- forcement learning or imitation learning, have demonstrated
centric navigation that operates fully onboard and in-the- fast reaction times and adaptability to perceptual noise [34]–
wild, without external infrastructure, explicit mapping, or [37]. However, this tight coupling of perception and control
dependency on structured priors. sacrifices interpretability, and poses significant challenges in
generalization to unseen environments. Moreover, training
such models requires extensive data collected in diverse,
II. RELATED WORKS unstructured settings, an effort that is not only costly but
A. Traditional Visual Servoing for Target Tracking also difficult to scale. Unstructured outdoor environments,
Visual servoing provides a well-established framework for in particular, exhibit extreme variability in appearance, ge-
controlling robotic motion with respect to a target, typically ometry, and semantics, making it impractical to fully cover
categorized into IBVS and PBVS [13]–[17]. In IBVS, control their distribution during training [38].
commands are computed by minimizing the error between To address these limitations, a more modular paradigm
observed and desired image-space features. These features has emerged where perception and control are treated as
are extracted from the camera view and mapped to robot distinct but tightly integrated subsystems. In this setting, the
motion using an interaction matrix that links image variation perception module is responsible for detecting and localizing
3
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
the target, while the control module translates this infor- More onboard-focused designs have emerged, such as
mation into safe and efficient motion commands. Extensive [60], which uses a monocular camera and a blob detector
research has been devoted to each side of the pipeline. On to track a single spherical obstacle. The framework inte-
the perception front, the computer vision community has grates perception and control onboard, but the simplicity of
produced a range of approaches for object detection and the sensing pipeline imposes its own constraints. Obstacle
tracking, from compact architectures optimized for real-time avoidance is implemented as a soft cost, tuned manually
inference on embedded platforms [39]–[42], to large-scale against the tracking objective, and the system assumes only
foundation models capable of open-set recognition and zero- one obstacle is visible at a time. The drone must be launched
shot generalization [43]–[46]. Techniques including temporal manually, and no mechanism is provided for adapting to
attention, depth prediction, and visual transformers have clutter or dynamic changes. While a step forward, the system
improved robustness to occlusion and visual degradation. is difficult to generalize to real-world scenarios involving
Alongside advances in perception, control design has dense geometry or multiple obstacles.
emerged as a key focus in enabling robust aerial target Multi-agent frameworks like CoNi-MPC [61] take a dif-
following. While model-free approaches such as reinforce- ferent approach, where a UAV follows a ground robot by
ment learning have demonstrated flexibility and general- optimizing a relative objective. This eliminates the need for
ization [47]–[50], they often struggle to provide consistent global SLAM, but again depends on full-state telemetry from
safety or performance guarantees, particularly under distri- the target and is tested under motion capture. The absence
bution shifts or in dynamic settings. In contrast, model- of onboard perception and low operational speeds (under
based methods, especially those based on NMPC [51]–[54], 0.4 m/s) further limit real-world relevance.
offer a structured way to embed physical constraints, safety Later extensions add LiDAR-based obstacle sensing [62],
margins, and task-specific objectives directly into the control improving situational awareness. However, core limitations
policy. By explicitly optimizing trajectories over a finite persist. The system still relies on external tracking for the
horizon, NMPC enables predictive and context-aware behav- target and operates in structured, static environments at low
iors, which are critical for maintaining target visibility and speeds (below 1.5 m/s). Visual perception remains absent,
avoiding obstacles during high-speed, reactive flight [55]. and the setup assumes favorable sensing conditions.
This separation of perception and control not only im- Together, these efforts reflect growing interest in object-
proves system modularity and interpretability, but also allows relative control, but also reveal a gap: current systems are not
for leveraging the complementary strengths of data-driven designed for fast, reactive tracking in complex environments
learning and analytical planning. However, the efficacy of using only onboard sensing. Our framework is built to
such hybrid systems depends critically on how well the two address exactly that gap. By combining real-time object
modules are coupled. detection, depth-completed state estimation, and constraint-
aware planning in a fully onboard pipeline, we enable agile,
target-relative control in GPS-denied, cluttered environment,
C. Object-Centric Visual Autonomy
without reliance on external infrastructure or prior maps.
Traditional navigation frameworks rely on global localization
through GPS, VIO, or SLAM, followed by waypoint-based
III. METHODOLOGY
trajectory generation [56]–[58]. While effective in struc-
A. System Overview
tured environments, these methods often fail in GPS-denied
or visually degraded settings, where feature tracking and NOVA is a fully onboard framework for vision-based aerial
map consistency degrade. Object-centric navigation offers target tracking and collision avoidance in GPS-denied and
a promising alternative: instead of referencing a global visually degraded environments. It requires no maps, infras-
frame, control objectives are defined relative to specific scene tructure, or prior tuning, enabling deployment in unknown
objects, such as people or vehicles. This allows robots to and cluttered settings. The system is structured around two
maintain situational awareness and track targets even when tightly integrated modules: perception and control (Figure 2).
global localization is unreliable or unavailable. The perception stack processes visual and inertial input to
Several recent works have pursued this idea, though typi- estimate the target’s relative state and detect potential colli-
cally under restrictive assumptions that limit applicability in sion risks in real time. These estimates feed into a NMPC
high-speed or unstructured environments. that dynamically tracks the target while enforcing safety
An early example is [59], which formulates control di- through system dynamics and real-time obstacle constraints.
rectly in the target’s frame of reference. While conceptually The entire pipeline runs onboard, supporting reactive agile
aligned with object-centric reasoning, the system relies en- flight with no external dependency.
tirely on motion capture to provide perfect state estimates
for both robot and obstacle. No onboard sensing is used, B. Perception
and full target velocity is assumed known. This removes The perception pipeline estimates the target’s position and
perception from the equation, making the setup impractical surrounding obstacles in real time using only onboard visual
outside controlled environments. and inertial sensors. At each frame, the system detects the
4
Figure 2: System architecture overview. The framework consists of two tightly integrated modules: perception and control.
The perception stack detects the target using a lightweight object detector, estimates its depth via stereo-monocular disparity
alignment, and processes obstacle information for collision avoidance. These outputs are fused with inertial data (omitted
from figure for simplicity) to produce a smooth, high-frequency target-centric odometry. The control module uses NMPC,
augmented with high-order CBFs, to generate safe and agile thrust and angular rate commands for tracking the target.
1) Target Detection
A central challenge in target tracking lies in balancing de- Figure 3: Adaptive zoom strategy. The system dynamically
tection range with computational efficiency. High-resolution crops and rescales the input image to center the target while
inputs, such as 640 × 480 pixels, support target detection suppressing irrelevant context. This targeted zoom improves
at distances beyond 40 m but exceed the processing limits detection confidence and robustness, particularly at long
of embedded platforms. In contrast, lower-resolution inputs range or in the presence of visually similar objects, by
like 320 × 240 pixels are more computationally efficient but reducing ambiguity and limiting false positives.
constrain detection to shorter ranges, often under 10 m. To
resolve this trade-off, we introduce an adaptive zooming
mechanism that dynamically crops and rescales the input
image around a region of interest (Figure 3). We employ a lightweight, custom-trained object detector
The zoom module takes the full-resolution RGB frame based on the YOLOv11-small architecture to enable onboard
It ∈ RH×W ×3 and a bounding box bt−1 = [x, y, w, h], real-time detection with limited computational resources.
where (x, y) denotes the center coordinate and (w, h) the The input to the detector is the zoomed-in crop Itcrop , and
width and height, from the previous detection. This box the output is a 2D bounding box bt in crop coordinates,
is enlarged by a factor α > 1 to introduce a margin which is then projected back into the full-resolution frame.
that accounts for motion and detection uncertainty. A crop The object detector is trained from scratch on the COCO
window is then centered on the enlarged box and shaped dataset [63] with domain-specific augmentations designed to
to match the detector’s aspect ratio. The zoomed-in region mirror the visual effects introduced by our zooming strategy.
is extracted from It and resized to the detector’s input In particular, randomized zooming and resizing simulate the
resolution, yielding Itcrop . pixelation and scale variations that occur during aggressive
If detection fails at time t, the crop size is progressively cropping, ensuring the detector remains effective even when
increased to widen the search region. Once the target is targets appear coarse or low-resolution. Additional augmen-
reacquired, the zoom readjusts to tightly frame the detection. tations include pitch and roll rotations and lighting variations
This mechanism preserves long-range detection capabilities to improve robustness under aerial tracking conditions.
at lower resolutions while improving robustness. By nar-
rowing the detector’s receptive field to a localized region,
the system avoids distractions from background clutter and 2) Depth Estimation
allows the network to focus on the target, increasing confi- Accurate depth perception is essential for estimating the
dence and accuracy under occlusion and motion. target distance as well as ensuring safe navigation. While
5
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
6
We first generate a TTC map Tt ∈ RH×W by projecting
robot’s velocity vector vCt , expressed in the camera frame,
along the viewing ray direction of each pixel. These ray
directions Rt (i, j) are unit-length vectors pointing from the
camera center through each pixel (i, j), and are computed
from the intrinsic calibration parameters. The TTC map is
then computed as:
Dtcom (i, j)
Tt (i, j) = . (9)
∥vCt · Rt (i, j)∥
Figure 6: Coordinate frame convention. The object-centric
To reduce the dimensionality of Tt , we divide it into non-
frame T is defined at the target location, with its orientation
overlapping grid cells Pu,v of size P ×P , and select the pixel
fixed and aligned with gravity using the IMU. A 2D detection
in each cell with the minimum TTC:
(x, y) in the camera frame C is back-projected into 3D and
transformed into the body frame B using known extrinsics
( )
t
(qBC , tBC ). The resulting position is then rotated into the Sttc = (i, j) ∈ Tt | (i, j) = arg min Tt (i, j), ∀ u, v .
(i,j)∈Pu,v
target frame T using the IMU-derived attitude qT B . This
(10)
defines the relative pose pT , which serves as the reference
We then apply a filtering step to retain only those points
for target-centric perception and control.
likely to intersect the quadrotor’s projected dimensions at
their corresponding depths:
( )
where qBC is the camera-to-body rotation (as a quaternion), t
t t |i − c x | ≤ Qx f x /[2D com (i, j)],
tBC the translation, and ⊙ the quaternion-vector product. Sfilt = (i, j) ∈ Sttc ,
|j − cy | ≤ Qy fy /[2Dtcom (i, j)]
We define a moving target-centric frame T , centered on (11)
the target and aligned with the ground plane. Although this where Qx , Qy represent the robot’s width and height.
frame may translate over time, we assume it remains parallel From this filtered set, we extract the top-K most critical
to the ground, and we model its orientation using the IMU- collision points:
derived attitude of the robot qT B . The relative position of X
the robot in this target frame is given by: t
Stop- K = arg min Tt (i, j), (12)
S ′tfilt ⊆Sfilt
t
,|S ′tfilt |=K (i,j)∈S ′ t
pT = −(qT B ⊙ pB ). (8) filt
The resulting vector pT expresses the quadrotor’s position The selected pixel coordinates are back-projected into 3D
relative to the target. This signal is tracked over time using points in the camera frame C :
an Unscented Kalman Filter (UKF) [66]. The UKF incorpo- ( )
z = Dtcom (i, j),
rates IMU’s angular velocity as process input, while visual t (i − cx )z (j − cy )z
SC = , ,z t .
measurements provide asynchronous updates. The full target- fx fy (i, j) ∈ Stop- K
centric state consists of position pT , velocity vT , orientation (13)
qT B , and angular velocity ωB . Finally, these 3D points are transformed into the body
This approach eliminates the need for any environment- frame B and then to the target-centric frame T using the
specific information, such as visual texture or persistent known extrinsics:
features, for state estimation. Unlike traditional VIO or
STt = qT B · (qBC ⊙ sC + tBC ) sC ∈ SCt ,
SLAM systems, our method relies solely on the detection (14)
and tracking of the target to estimate the full relative state
of the quadrotor. This enables robust performance even in These target-centric obstacle points STt are published
textureless, dynamic, and unknown environments, making it at each planning cycle and used to construct hard safety
especially well-suited for agile target tracking. constraints within our NMPC formulation.
7
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
8
Specifically, when only a single target is detected, the
quadrotor’s position relative to the target is observable, but
its yaw orientation remains underconstrained. This creates
ambiguity in the rotation about the vertical axis, allowing
orbiting behaviors that preserve the perceived target location
but degrade stability and control.
While previous methods [60] resolve this by tracking mul-
tiple targets, we argue that this assumption is overly restric-
tive in real-world aerial scenarios. Maintaining visibility of
multiple targets while tracking one is extremely challenging
in cluttered, dynamic, or degraded visual environments. Figure 7: Aerial robot platform. Custom quadrotor used in
Instead, we introduce a soft regularization on vyB , exploit- experiments, equipped with an onboard computer, a stereo
ing the fact that orbiting motion correlates with lateral body RGB-D camera, and a PX4 flight controller.
velocity. By penalizing this component, the NMPC naturally
favors forward-facing, stable flight without additional per-
ceptual burden. This approach is lightweight, assumption- and safe control sequence over a finite horizon N . The
free, and effective across diverse tracking tasks. objective is to minimize a cumulative cost that promotes
accurate tracking, smooth control, and orbit-free motion:
N −1
4) Control Barrier Functions for Obstacle Avoidance
X
min J(xt+j , ut+j ) + J(xt+N , 0), (29)
To ensure safety during target tracking, we incorporate xt ,...,xt+N
j=0
second-order CBFs into the NMPC to enforce minimum ut ,...,ut+N −1
distance constraints with respect to perceived obstacles. Each where the terminal cost is evaluated with zero control input.
obstacle point stT ∈ STt is derived from depth perception The optimization is subject to the following constraints for
and expressed in the target frame T , consistent with object- all predictions j ∈ [0, N ] and safety constraints k ∈ [0, K):
centric control formulations proposed in recent works [55]. xt = x̂t , (30a)
For each high-risk point, we define a safety function: x t+1+j
= f (x t+j t+j
,u ), (30b)
t
hk (x ) = ∥stT − ptT 2
∥ − Q2max , (25) xmin ≤ x t+j
≤ xmax , (30c)
t+j
where Qmax = max(Qx , Qy ) is the minimum allowable umin ≤ u ≤ umax , (30d)
distance that accounts for the robot’s physical dimensions. ḧk (x t+j
,ut+j
) + 2λḣk (x t+j 2 t+j
) + λ hk (x ) ≥ 0. (30e)
Safety is enforced by requiring the CBF condition:
This formulation allows the NMPC to reason over future
ḧk (xt , ut ) + 2λḣk (xt ) + λ2 hk (xt ) ≥ 0, (26) trajectories while embedding safety constraints, actuation
which guarantees forward invariance of the safe set defined limits, and regularization strategies into a unified and reactive
by hk (x) ≥ 0. Here, λ > 0 controls the rate of enforcement. control framework.
The first derivative of the safety function is:
IV. EXPERIMENTAL SETUP
ḣk (xt ) = 2(stT − ptT )⊤ vTt , (27) A. Aerial System
and the second derivative, capturing the effect of control Our quadrotor platform, shown in Figure 7, weighs 1.3 kg
inputs, is given by: and has a motor span of 25 cm, with a thrust-to-weight ratio
of approximately 4 to 1. It is powered by a 6S LiPo battery
ḧk (xt , ut ) = 2∥vTt ∥2 + 2(stT − ptT )⊤ v̇T (u), (28)
and uses 1850KV motors paired with a NewBeeDrone
where v̇T (u) is the acceleration of the robot in the target Infinity200 V2 4IN1 ESC 55A to enable agile flight.
frame, obtained from the dynamics model f (xt , ut , δt). The system integrates an Intel RealSense D455 stereo
Each of these constraints is imposed at runtime for every camera for visual and depth sensing and an onboard
selected obstacle point stT , ensuring that the NMPC respects IMU. The camera is mounted front-facing, and its in-
proximity constraints while pursuing the tracking objective. trinsics are (fx , fy , cx , cy ) = (325.2, 430.9, 323.1, 246.9),
This formulation allows us to encode collision avoidance in a obtained from standard calibration [65]. The camera-to-
computationally efficient and differentiable form compatible body transformation is defined by the static quaternion
with real-time optimization. qBC = [−0.5, 0.5, −0.5, 0.5]⊤ and translation tBC =
[0.061, 0.047, −0.065]⊤ m. RGB-D frames are captured at
60 Hz with a resolution of 640 × 480.
5) Optimization Problem Onboard processing is handled by an NVIDIA Jetson
At each planning step, the NMPC solves a constrained Orin NX 16GB, and low-level stabilization is managed by a
optimal control problem to compute a dynamically feasible PixRacer Pro flight controller.
9
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
B. Target Detector
We employ a YOLOv11-small model for object detection,
trained from scratch on the COCO dataset with heavy
augmentation. Images are randomly zoomed within a scale
range of [0.6, 1.4], rotated within ±15◦ in pitch and roll, and
augmented with synthetic motion blur (30% probability) and
brightness shifts of ±20%. Crops are resized to 320×256 for
inference. The model is trained using Adam with a learning Figure 8: Targets used for tracking. All experiments use
rate of 2 × 10−4 , batch size 32, weight decay 10−5 , and for one of two object-mounted targets: a mannequin positioned
300 epochs. The inference runs at 92 Hz using TensorRT on upright on an ATV, or a stop sign mounted at the rear. The
the Orin NX with floating 16 precision. mannequin is approximately 1.8 m tall and wears typical
To initialize tracking, we use a one-time prompting stage clothing to simulate a human presence. The stop sign has
that leverages a large, high-capacity object detector. Specif- a diameter of 0.8 m, is red and reflective, and presents
ically, a YOLOv11x model [73] is run once on the full- challenges due to glare and light reflections. The ATV is used
resolution RGB image to acquire an initial bounding box for to transport the targets dynamically during experiments and
the target. This model is not trained or fine-tuned within our can reach speeds of up to 50 km/h across various terrains.
system. It is used solely for inference. Among all detections
returned, the bounding box with the highest confidence score for the position estimate, and 0.0001 for all four quaternion
is selected. This bounding box seeds the zooming module components in the orientation update.
for subsequent frames. After this initialization, the large The IMU provides linear acceleration and angular velocity
YOLOv11x model is terminated to eliminate unnecessary at 200 Hz. The UKF is also executed at this rate and
computational overhead. From that point onward, tracking handles bias estimation online without requiring separate
continues using only the lightweight onboard detector. Im- pre-calibration. All position and orientation estimates are
portantly, this prompting mechanism is modular. Any object computed in the target-centric frame described in Section 3.
detector capable of producing a coarse bounding box on For outdoor flights, a u-Blox Neo-M9N GPS module was
the initial frame can be substituted without modifying the fused with IMU data using an Extended Kalman Filter
downstream pipeline. The sole requirement is that it provides (EKF), yielding position updates at 100 Hz. This global
an early spatial prior for focusing the zoomed window. localization information was purely used for benchmarking.
The proposed system operates independently of this infor-
C. Depth Completion mation during all the flight experiments.
Depth maps from the RealSense D455 (in HighAccuracy
mode [74]) are completed using DepthAnythingV2 [64] with F. Planning and Control
a ViT-S backbone, producing per-frame estimates in 31.6 ms. We formulate an NMPC problem with a prediction horizon
A histogram bin width of 0.15 m is used for mode filtering. of N = 10 steps over 2 s and a discrete time step of δt =
0.2 s. The dynamics are integrated using a Runge–Kutta 4
D. Obstacle Selection method. The goal reference is shifted using a lateral safety
We divide the image into 17 × 17 non-overlapping cells margin of dsafe chosen based on the experiment, and blended
and select K = 10 obstacle points with the lowest TTC. with a ramp function α(t) with quadratic horizontal decay
The drone’s projected dimensions are Qx = 0.4 m and over 1 s. The desired yaw angle is computed to center the
Qy = 0.2 m. These are used in pixel-space filtering to reject target in the image, and converted to a quaternion assuming
low-risk points. The TTC map uses a forward-projected zero roll and pitch.
velocity in the camera frame and ray directions computed The NMPC cost function includes a state deviation penalty
from calibrated intrinsics. Qx = diag(150, 100, 150, 15, 15, 15, 50, 15, 15, 50, 5, 5, 5),
input cost Qu = diag(1, 1, 1, 1), and an orbiting penalty
E. State Estimation Qo = 20 on lateral body velocity. State constraints are
Target-relative state is tracked using an UKF that fuses enforced with element-wise bounds x ∈ [−999, 999]3 ×
asynchronous visual detections with high-rate inertial mea- [−25, 25]3 × [−10, 10]4 × [−40, 40]3 , and control inputs are
surements. The filter estimates the full 6-DoF state along constrained to u ∈ [0.05, 8.0]4 .
with accelerometer and gyroscope biases. Safety constraints are enforced using second-order control
The UKF process noise standard deviations are defined barrier functions with a safety radius Qmax = 0.3 m and gain
as follows: linear acceleration noise is (0.1, 0.1, 1.0) m/s2 , λ = 2. Up to 10 constraints are active per solve.
angular velocity noise is (0.2, 0.2, 0.2) rad/s, accelerometer The optimization problem is solved with acados [75],
bias noise is (0.1, 0.1, 0.1) m/s2 , and gyroscope bias noise is using SQP with Gauss–Newton Hessian approximation and
(0.1, 0.1, 0.1) rad/s. Measurement noise standard deviations Levenberg–Marquardt regularization equal to 10−2 , with a
for the visual update are 0.01 m in (x, y) and 0.001 m in z max iteration limit of 20 and a feasibility tolerance of 10−4 .
10
sensed color
completed depth
Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
4
100 12
3 1.0
velocity (m/s)
velocity (m/s)
75
distance (m)
distance (m)
10
50 2
8 0.5
25 1
0 0 dsafe
6 0.0
0 20 40 60 0 20 40 60 0 20 40 60 0 20 40 60
time (s) time (s) time (s) time (s)
Figure 9: Tracking performance in a structured urban maze. NOVA follows the target through narrow corridors and
around occlusions, maintaining stable depth estimation and obstacle awareness despite minimal texture. Bottom plots show
consistent velocity regulation and relative distance control, confirming safe and responsive tracking throughout the mission.
11
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
sensed color
completed depth
Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
15 17.5 1.5
800
15.0
velocity (m/s)
velocity (m/s)
600
distance (m)
distance (m)
10 1.0
400 12.5
5 0.5
200 10.0
dsafe
0 0 7.5 0.0
0 50 100 0 50 100 0 50 100 0 50 100
time (s) time (s) time (s) time (s)
Figure 10: High-speed pursuit in a rugged forest trail. NOVA tracks the target over a 1 km unstructured path with
potholes, dust, and vegetation. Despite motion blur and strong lighting variation, detection and depth estimation remain
stable. Bottom plots confirm that the UAV sustains target speed while maintaining safe distance and minimal relative drift.
B. Forest Trail Pursuit not only keeps up but tracks with high temporal fidelity.
We now move into unstructured natural terrain, where sens- The system maintains a safe following distance with smooth
ing and control are stressed by long-range motion, fast speed regulation, navigating cluttered vegetation without
speeds, and an unpredictable environment. The forest trail hesitation or instability.
is a rugged 1 km path cutting through alternating segments
of dense canopy and open clearings. This terrain exposes
the system to rapid lighting transitions. The ground is C. Hangar Transition with Visual Ambiguity
broken and uneven, filled with potholes that jolt the ATV
We take a step further by designing a scenario where percep-
and introduce constant motion disturbances. Dust kicked
tion is pushed to its breaking point. The mission begins in an
up during traversal adds noise to the image stream, while
open gravel lot under direct sunlight, transitions into a tall
overhanging branches and narrow passageways demand fast
metallic hangar, and then returns to outdoor terrain. This path
reactive obstacle avoidance.
introduces a cascade of perceptual challenges: complete GPS
The ATV reaches speeds of up to 50 km/h, forcing the
dropout inside the hangar, abrupt lighting shifts that drive the
UAV to maintain high-speed flight while keeping visual lock,
camera from overexposure to near-total underexposure, and
estimating depth, and planning safe paths through dynamic
a cluttered interior filled with reflective surfaces, structural
obstacles. There are no structured boundaries or predictable
obstacles, and narrow doorways.
layouts here. NOVA must rely entirely on its onboard sensors
To amplify the difficulty, we place two additional man-
to adapt in real time, safely plan and track the dynamic target.
nequin targets near one of the hangar exits. These decoys
As illustrated in Figure 10, the UAV holds tight formation
are visually similar to the mannequin mounted on the ATV,
throughout the run. Despite motion blur and exposure shifts,
requiring the system to maintain target identity precisely
the system consistently detects the target and completes
during a moment when visibility is degraded and the lighting
dense depth maps that preserve usable geometry. Velocity
is unstable. This setup explicitly tests NOVA’s robustness to
plots show near-zero relative speed, indicating that NOVA
visual ambiguity, occlusion, and target switching.
12
sensed color
completed depth
Traveled Distance XY (GPS) Absolute Velocity XY (GPS) Relative Distance XYZ (UKF) Relative Velocity XYZ (UKF)
2.0
300 10.0
12 1.5
7.5
velocity (m/s)
velocity (m/s)
distance (m)
distance (m)
200
5.0 1.0
100 10
2.5 0.5
dsafe
0 0.0 8 0.0
0 25 50 75 0 25 50 75 0 25 50 75 0 25 50 75
time (s) time (s) time (s) time (s)
Figure 11: Robust tracking across indoor–outdoor transitions. The UAV follows the target from a gravel lot through a
tall metallic hangar and back outside, navigating abrupt lighting changes, GPS dropout, and visual sparsity. Despite exposure
shifts and structural occlusions, NOVA maintains safe distance and continuous target lock throughout the entire mission.
Such environments typically break conventional tracking UAV and the target. In this experiment, NOVA is tasked with
pipelines. Feature-based methods struggle with low texture, maintaining a constant 6 m vertical offset from the estimated
stereo matching fails under reflections, and bright-to-dark target position, resulting in a sustained flight altitude of
transitions overwhelm standard exposure control. GPS-based nearly 10 m above ground. This constraint is evaluated end
localization is unavailable inside the structure, and visual to end across a single continuous mission: from the open
odometry is easily corrupted. parking lot, through the metallic hangar, across the gravel
Despite these conditions, NOVA maintains continuous yard, and into the corridors of the simulated container maze.
tracking. As illustrated in Figure 11, RGB and completed This setup introduces a compounded set of perception and
depth frames remain functional even under poor lighting control challenges. At elevated altitudes, stereo parallax is
and low-feature geometry. When stereo cues degrade, NOVA significantly reduced, degrading the quality and density of
falls back on inertial integration and histogram-filtered depth depth estimates. The target appears smaller in the frame,
estimates to maintain a consistent relative pose. It does not limiting the number of visible features and reducing de-
require artificial markers, global maps, or external references. tector robustness. The UAV’s field of view becomes more
The trajectory plots confirm smooth flight. Relative dis- constrained with respect to lateral motion, making it harder
tance remains stable, with no abrupt control shifts or recov- to anticipate turns or sudden accelerations. Precise obstacle
ery maneuvers. The UAV transitions through the building avoidance becomes critical, especially around overhanging
without reinitialization, rejects the decoy mannequins, and structures, hangar entrances, and confined turns in the maze.
maintains accurate lock on the moving target throughout the The environments themselves amplify these difficulties.
most perceptually challenging phase of the mission. The hangar imposes complete GPS dropout and rapid light-
ing changes. The container maze presents frequent occlu-
sions, low-texture surfaces, and sharp-angle geometry. From
D. Elevated Tracking Across Mixed Terrain
a top-down viewpoint, NOVA must maintain tight target
We take an even further step to complicate the mission by
coupling while flying above clutter.
deliberately altering the spatial configuration between the
13
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
sensed color
completed depth
velocity (m/s)
distance (m)
distance (m)
12
100 4
10 0.5
50 2
dsafe
0 0 8 0.0
0 20 40 0 20 40 0 20 40 0 20 40
time (s) time (s) time (s) time (s)
Figure 12: Tracking from elevated viewpoints with forced height offset. The UAV follows the target through a combined
indoor–outdoor and urban maze mission while maintaining a 6 m vertical offset. This configuration introduces degraded
stereo geometry, reduced field-of-view, and tight spatial constraints. NOVA preserves visual lock, avoids obstacles, and
regulates target distance despite elevated flight and complex terrain.
As shown in Figure 12, NOVA completes the full mission and frequent reacceleration. NOVA sustains high detection
without loss of lock or deviation from safe bounds. Despite reliability (98.6%) while maintaining an average separation
increased vertical separation and degraded sensing geometry, of 8.2 m from the target, with a minimum of 6.2 m,
it maintains coherent relative state estimates and reconstructs just above the dsafe = 6.0 m threshold. The high peak
high-fidelity depth maps. Relative distance and velocity re- acceleration of 32.4 m/s2 and max pitch of 15.0◦ reflect
main tightly regulated, and the UAV maneuvers with smooth, sharp control inputs required to navigate tight geometry
responsive control throughout. while maintaining visual lock.
This experiment confirms that NOVA generalizes not The Forest Trail scenario emphasizes sustained high-
only across different environments, but also across spatial speed tracking over unstructured terrain. The target exceeds
configurations that significantly alter perception and planning 53 km/h, and the UAV maintains a mean relative distance
dynamics. No changes are made to the underlying system; of 10.0 m, dipping to 8.4 m at its closest—well within the
all modules run as configured in prior tests, reinforcing the dsafe = 8.0 m constraint. Control demands are moderate,
adaptability of the stack under shifted tracking regimes. with limited acceleration spikes and a max pitch of 19.6◦ ,
indicating smooth flight despite blur, shadows, and environ-
mental noise. NOVA’s object detector remains effective under
E. Measured Performance across Tested Environments
these fast, noisy conditions, with 94.5% detection accuracy.
To complement the qualitative results, we report quantitative
In the Building Transition trial, the system experiences
metrics for the previous four representative experiments.
full GPS dropout and sharp illumination shifts when passing
Table 1 summarizes performance across these trials, covering
through a metallic hangar. NOVA responds with increased
GPS-based velocity, control effort, pitch dynamics, target
control activity, reaching 22.6 m/s2 peak acceleration and
distance regulation, and object detection consistency.
maintaining an average pitch of 5.3◦ . The UAV holds an
In the Urban Maze, the UAV operates within narrow corri-
average target distance of 9.9 m, with a minimum of 8.9 m,
dors and sharp turns, demanding precise obstacle avoidance
14
Table 1: Quantitative performance across the four outdoor tracking missions. We report results from representative trials
in diverse environments, each introducing unique challenges in geometry, speed, sensing, and elevation. NOVA consistently
maintains safe separation, stable flight, and high detection rates without environment-specific tuning. Reported metrics
include UAV velocity (mean and peak), linear acceleration, pitch angle, mean and minimum distance to the target compared
to the safety threshold dsafe , and overall detection rate.
GPS Velocity (km/h) Control Effort (m/s²) Pitch (°) Rel. Distance (m)
Scenario Detections (%)
Mean Max Mean Max Mean Max Mean Min dsafe
Urban Maze 5.7 14.0 10.2 32.4 2.3 15.0 8.2 6.2 6.0 98.6
Forest Trail 35.0 53.3 10.2 13.9 10.1 19.6 10.0 8.4 8.0 94.5
Building Transition 17.6 37.6 10.3 22.6 5.3 16.6 9.9 8.9 8.0 96.2
Elevated Tracking 14.5 27.5 10.3 22.8 7.7 28.8 9.4 8.8 8.0 93.1
15
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
50 0 0
0 10
0 25 50 0 50 100 0 50 0 20 40 3.5
0 0 0 0
pos y (m)
3.0
50 500 100
200
12 10.0
10 15.0 2.0
11 5 7.5
12.5
0 25 50 0 50 100 0 50 0 20 40 1.5
27
25
satellites
25 20
20 1.0
20 10
26
0 25 50 0 50 100 0 50 0 20 40
time (s) time (s) time (s) time (s)
Figure 14: GPS signal degradation across outdoor tracking scenarios. The top rows show fused GPS+IMU global
position traces for each of the four outdoor missions, color-coded by estimated horizontal positional error (eph). Lighter
colors indicate higher uncertainty. The bottom plot shows the number of GPS satellites tracked over time. In the Urban
Maze and Forest Trail scenarios, signal quality remains high, with low positional error and stable satellite lock. In contrast,
the Building Transition and Elevated Tracking missions exhibit significant degradation, including extended satellite dropout
and increased eph. This effect is especially pronounced during indoor segments and when the UAV flies near structural
ceilings, highlighting the limitations of GPS-based localization in partially enclosed or cluttered environments.
G. GPS Signal Quality and Its Limitations The Elevated Tracking scenario presents even more severe
Although NOVA operates without GPS during flight, we GPS degradation. Although the trajectory includes both
use fused GPS+IMU estimates post hoc to visualize global indoor and outdoor segments, the required vertical offset
trajectories and assess how environmental geometry affects places the UAV closer to the ceiling during the hangar phase.
localization quality. This provides a comparative baseline This reduces the already limited visibility to the sky and
for understanding where GPS-based methods remain reliable exacerbates signal dropout. In several segments, the number
and where they degrade. of visible satellites falls to near zero, and position estimates
Figure 14 summarizes GPS performance across all four become highly uncertain.
representative scenarios. The top rows show the global po- These observations highlight the limits of GPS-based
sition trace, color-coded by estimated horizontal error (eph). localization in built environments. Even in outdoor set-
The bottom plot reports the number of visible satellites over tings, partial occlusion or reflective surfaces can compromise
time. Together, these metrics reflect how surrounding struc- satellite visibility and introduce substantial pose uncertainty.
tures influence satellite visibility and positioning accuracy. NOVA’s design avoids these failure modes by relying exclu-
In the Urban Maze and Forest Trail scenarios, satellite sively on onboard visual and inertial sensing, which remains
visibility remains high, typically above 20 satellites, and the operational across all tested scenarios regardless of external
estimated horizontal error stays below 1.0 m. These envi- infrastructure availability.
ronments, although partially obstructed, allow for relatively
consistent signal reception. The GPS data in these trials H. Ablation Studies
remains smooth and usable throughout the mission. To evaluate the contribution of individual components within
In contrast, the Building Transition and Elevated Offset the NOVA framework, we conduct a series of ablation
experiments exhibit significant signal degradation. During studies. Each study isolates one module and examines its
the Building Transition mission, satellite count drops sharply effect on overall system performance, while keeping the
as the UAV enters the hangar, with corresponding increases remainder of the stack fixed. The goal is to assess how spe-
in estimated error, often exceeding 3.5 m. These effects are cific design elements contribute to robustness, accuracy, and
due to occlusion, multipath reflections, and direct signal loss. safety during target tracking in unstructured environments.
16
Figure 15: Impact of Adaptive Zoom in Multi-Target Scenes. In cluttered scenes with multiple visually similar targets
(black bounding boxes), tracking can become unstable when relying on full-frame detection, often resulting in identity
switches or drift. The adaptive zoom module addresses this by cropping tightly around the prompted target, suppressing
distractors and maintaining detection focus. Shown here are drone trajectories (blue) for three separate trials, each prompted
to track a different mannequin. NOVA successfully adheres to the correct target in all cases.
17
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
18
Second, the system relies on the ability to estimate the viewpoint offsets, repeatability over multiple trials, and ro-
target’s depth using stereo-based depth completion. While bustness under degraded GPS and perceptual ambiguity.
this approach is sufficient at close to mid-range (typically Ablation studies confirmed the contribution of individual
up to 30–40 m), it breaks down when the target is detected components, such as adaptive zooming for long-range detec-
at longer distances and stereo cues become unreliable. In tion and identity preservation. The system’s limitations were
these conditions, detection may still succeed due to the zoom also analyzed, including assumptions about initial visibility,
module, but the depth estimates become noisy or flat, making semantic category constraints, and the reliance on stereo-
the x-axis velocity (change in relative depth) unobservable. based depth at range.
As a result, the UAV may accelerate rapidly toward the Together, these results support the conclusion that robust,
target with limited feedback. In flight experiments, the robot real-time target tracking can be achieved using only onboard
eventually stabilizes once the depth becomes consistent, but sensing, even in the absence of external infrastructure or
the initial approach is often jerky and fast, reflecting the lack structured environments. Future works will focus on relaxing
of velocity feedback in the unobservable regime. current assumptions, extending to open-set target represen-
To relax the reliance on stereo depth, several alternative tations, and enabling perception-driven control under long-
strategies could be explored. One approach is to augment range uncertainty.
the sensing stack with additional depth-aware sensors, such
as radar or lightweight range finders, which are less sensi- References
tive to texture and lighting and can extend the operational [1] S. Wu, R. Li, Y. Shi, and Q. Liu, “Vision-based target detection and
range [78]–[80] . Another is to infer short-term relative tracking system for a quadcopter,” IEEE Access, vol. 9, pp. 62 043–
motion using optical flow, either from dense image features 62 054, 2021.
[2] M. Xu, A. Hu, and H. Wang, “Visual-impedance-based human–robot
or from bounding box displacements of detected objects, cotransportation with a tethered aerial vehicle,” IEEE Transactions on
including in open-set configurations [81]. A third direction Industrial Informatics, vol. 19, no. 10, pp. 10 356–10 365, 2023.
involves registering successive point clouds generated from [3] A. Hu, M. Xu, H. Wang, and H. Castañeda, “Vision-based impedance
control of an aerial manipulator using a nonlinear observer,” IEEE
depth completion, even if sparse or noisy, to recover frame- Transactions on Automation Science and Engineering, vol. 20, no. 2,
to-frame translation [82]. These approaches would allow for pp. 1441–1451, 2022.
estimation of instantaneous robot velocity without relying on [4] J. Gu, T. Su, Q. Wang, X. Du, and M. Guizani, “Multiple moving
targets surveillance based on a cooperative network for multi-uav,”
world-frame position or persistent maps, and could remain IEEE Communications Magazine, vol. 56, no. 4, pp. 82–89, 2018.
robust under partial feature disruption due to their minimal [5] N. Bashir, S. Boudjit, and S. Zeadally, “A closed-loop control archi-
or stateless memory requirements. tecture of uav and wsn for traffic surveillance on highways,” Computer
These assumptions are not intrinsic limitations of the Communications, vol. 190, pp. 78–86, 2022.
[6] H. Huang, A. V. Savkin, and W. Ni, “Online uav trajectory planning
overall framework, but they define the boundaries of the cur- for covert video surveillance of mobile targets,” IEEE Transactions
rent implementation. Future work targeting open-set object on Automation Science and Engineering, vol. 19, no. 2, pp. 735–746,
understanding and depth-agnostic motion estimation could 2021.
[7] L. Quan, L. Yin, T. Zhang, M. Wang, R. Wang, S. Zhong, Y. Cao,
expand NOVA’s applicability to more uncertain, long-range, C. Xu, and F. Gao, “Formation flight in dense environments,” CoRR,
or semantically unconstrained tracking scenarios. 2022.
[8] G. A. Di Caro and A. W. Z. Yousaf, “Multi-robot informative path
planning using a leader-follower architecture,” in IEEE International
Conference on Robotics and Automation, 2021, pp. 10 045–10 051.
[9] T. Miki, P. Khrapchenkov, and K. Hori, “Uav/ugv autonomous co-
VII. CONCLUSION operation: Uav assists ugv to climb a cliff by attaching a tether,” in
We presented NOVA, an onboard visual-inertial framework IEEE International Conference on Robotics and Automation, 2019, pp.
for agile target tracking in unstructured and GPS-denied 8041–8047.
[10] G. Niu, Q. Yang, Y. Gao, and M.-O. Pun, “Vision-based autonomous
environments. The system avoids reliance on external local- landing for unmanned aerial and ground vehicles cooperative systems,”
ization, global maps, or precomputed scene priors. Instead, IEEE Robotics and Automation Letters, vol. 7, no. 3, pp. 6234–6241,
it formulates perception, estimation, and control directly in 2021.
[11] M. Demirhan and C. Premachandra, “Development of an auto-
the target’s reference frame, using only stereo vision and in- mated camera-based drone landing system,” IEEE Access, vol. 8, pp.
ertial sensing. The approach combines a lightweight detector 202 111–202 121, 2020.
with adaptive zooming, depth completion, and visual-inertial [12] P. Vlantis, P. Marantos, C. P. Bechlioulis, and K. J. Kyriakopoulos,
“Quadrotor landing on an inclined platform of a moving ground vehi-
state estimation, followed by a NMPC that operates under cle,” in IEEE International Conference on Robotics and Automation,
collision-aware constraints derived from onboard sensing. 2015, pp. 2202–2207.
We validated NOVA across a series of real-world trials [13] F. Chaumette and S. Hutchinson, “Visual servo control. ii. ad-
that stress the system across different terrain types, motion vanced approaches [tutorial],” IEEE Robotics & Automation Magazine,
vol. 14, no. 1, pp. 109–118, 2007.
regimes, and sensory conditions. The system demonstrated [14] S. Hutchinson, G. D. Hager, and P. I. Corke, “A tutorial on visual servo
consistent performance in forested, urban, and mixed in- control,” IEEE transactions on robotics and automation, vol. 12, no. 5,
door–outdoor environments, maintaining visual lock, re- pp. 651–670, 1996.
[15] S. Cho and D. H. Shim, “Sampling-based visual path planning
specting safety distances, and operating without manual tun- framework for a multirotor uav,” International Journal of Aeronautical
ing. Additional experiments evaluated generalization across and Space Sciences, vol. 20, pp. 732–760, 2019.
19
Saviolo et al.: NOVA: Navigation via Object-Centric Visual Autonomy for High-Speed Target Tracking in Unstructured GPS-Denied Environments
[16] A. A. Oliva, E. Aertbeliën, J. De Schutter, P. R. Giordano, and ing,” Journal of Marine Science and Engineering, vol. 10, no. 3, p.
F. Chaumette, “Towards dynamic visual servoing for interaction con- 383, 2022.
trol and moving targets,” in IEEE International conference on robotics [37] W. Zhang, K. Song, X. Rong, and Y. Li, “Coarse-to-fine uav target
and automation, 2022, pp. 150–156. tracking with deep reinforcement learning,” IEEE Transactions on
[17] S. Raj, P. R. Giordano, and F. Chaumette, “Appearance-based indoor Automation Science and Engineering, vol. 16, no. 4, pp. 1522–1530,
navigation by ibvs using mutual information,” in 14th International 2018.
Conference on Control, Automation, Robotics and Vision, 2016, pp. [38] C. Min, S. Si, X. Wang, H. Xue, W. Jiang, Y. Liu, J. Wang, Q. Zhu,
1–6. Q. Zhu, L. Luo, et al., “Autonomous driving in unstructured environ-
[18] P. Geneva, K. Eckenhoff, W. Lee, Y. Yang, and G. Huang, “Openvins: ments: How far have we come?” arXiv preprint arXiv:2410.07701,
A research platform for visual-inertial estimation,” in IEEE Interna- 2024.
tional Conference on Robotics and Automation, 2020, pp. 4666–4672. [39] R. Khanam and M. Hussain, “Yolov11: An overview of the key
[19] T. Qin, P. Li, and S. Shen, “Vins-mono: A robust and versatile monoc- architectural enhancements,” arXiv preprint arXiv:2410.17725, 2024.
ular visual-inertial state estimator,” IEEE transactions on robotics, [40] Y. Zhao, W. Lv, S. Xu, J. Wei, G. Wang, Q. Dang, Y. Liu, and
vol. 34, no. 4, pp. 1004–1020, 2018. J. Chen, “Detrs beat yolos on real-time object detection,” in IEEE/CVF
[20] D. Scaramuzza and Z. Zhang, “Visual-inertial odometry of aerial conference on computer vision and pattern recognition, 2024, pp.
robots,” arXiv preprint arXiv:1906.03289, 2019. 16 965–16 974.
[21] M. Labbé and F. Michaud, “Rtab-map as an open-source lidar and [41] Y. Zhang, P. Sun, Y. Jiang, D. Yu, F. Weng, Z. Yuan, P. Luo, W. Liu,
visual simultaneous localization and mapping library for large-scale and X. Wang, “Bytetrack: Multi-object tracking by associating every
and long-term online operation,” Journal of field robotics, vol. 36, detection box,” in European conference on computer vision. Springer,
no. 2, pp. 416–446, 2019. 2022, pp. 1–21.
[22] W. G. Aguilar, G. A. Rodrı́guez, L. Álvarez, S. Sandoval, F. Quis- [42] N. Aharon, R. Orfaig, and B.-Z. Bobrovsky, “Bot-sort: Robust asso-
aguano, and A. Limaico, “Visual slam with a rgb-d camera on a ciations multi-pedestrian tracking,” arXiv preprint arXiv:2206.14651,
quadrotor uav using on-board processing,” in Advances in Computa- 2022.
tional Intelligence: 14th International Work-Conference on Artificial [43] T. Cheng, L. Song, Y. Ge, W. Liu, X. Wang, and Y. Shan, “Yolo-world:
Neural Networks, IWANN 2017, Cadiz, Spain, June 14-16, 2017, Real-time open-vocabulary object detection,” in IEEE/CVF Conference
Proceedings, Part II 14. Springer, 2017, pp. 596–606. on Computer Vision and Pattern Recognition, 2024, pp. 16 901–16 911.
[23] R. Mur-Artal, J. M. M. Montiel, and J. D. Tardos, “Orb-slam: A [44] A. Wang, L. Liu, H. Chen, Z. Lin, J. Han, and G. Ding, “Yoloe:
versatile and accurate monocular slam system,” IEEE transactions on Real-time seeing anything,” arXiv preprint arXiv:2503.07465, 2025.
robotics, vol. 31, no. 5, pp. 1147–1163, 2015. [45] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr,
[24] Y. Wang and A. Zell, “Improving feature-based visual slam by seman- R. Rädle, C. Rolland, L. Gustafson, et al., “Sam 2: Segment anything
tics,” in International Conference on Image Processing, Applications in images and videos,” in The Thirteenth International Conference on
and Systems (IPAS), 2018, pp. 7–12. Learning Representations, 2025.
[25] T. Tzanetos, M. Aung, J. Balaram, H. F. Grip, J. T. Karras, T. K.
[46] S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li,
Canham, G. Kubiak, J. Anderson, G. Merewether, M. Starch, et al.,
J. Yang, H. Su, et al., “Grounding dino: Marrying dino with grounded
“Ingenuity mars helicopter: From technology demonstration to ex-
pre-training for open-set object detection,” in European Conference on
traterrestrial scout,” in aerospace conference (AERO), 2022, pp. 01–19.
Computer Vision. Springer, 2024, pp. 38–55.
[26] J. Balaram, M. Aung, and M. P. Golombek, “The ingenuity helicopter
[47] A. Saviolo, P. Rao, V. Radhakrishnan, J. Xiao, and G. Loianno,
on the perseverance rover,” Space Science Reviews, vol. 217, no. 4,
“Unifying foundation models with quadrotor control for visual track-
p. 56, 2021.
ing beyond object categories,” in IEEE International Conference on
[27] S. Withrow, W. Johnson, L. A. Young, H. Cummings, J. Balaram, and
Robotics and Automation, 2024, pp. 7389–7396.
T. Tzanetos, “An advanced mars helicopter design,” in ASCEND 2020,
2020, p. 4028. [48] S. Hu, Q. Wang, F. Wang, and Y. Li, “Finite-time dynamic visual servo
[28] H. F. Grip, D. Conway, J. Lam, N. Williams, M. P. Golombek, control for quadrotor tracking unknown motion target,” Nonlinear
R. Brockers, M. Mischna, and M. R. Cacan, “Flying a helicopter on Dynamics, vol. 113, no. 7, pp. 6959–6977, 2025.
mars: How ingenuity’s flights were planned, executed, and analyzed,” [49] Y. Kumar, S. B. Roy, et al., “Adaptive ibvs based planar non-
in Aerospace Conference, 2022, pp. 1–17. holonomic target tracking for quadrotors,” in International Conference
[29] K. Zhang, Y. Shi, and H. Sheng, “Robust nonlinear model predictive on Unmanned Aircraft Systems, 2024, pp. 201–208.
control based visual servoing of quadrotor uavs,” IEEE/ASME Trans- [50] M. Leomanni, F. Ferrante, A. Dionigi, G. Costante, P. Valigi, and M. L.
actions on Mechatronics, vol. 26, no. 2, pp. 700–708, 2021. Fravolini, “Quadrotor control system design for robust monocular
[30] D. Guo and K. K. Leang, “Image-based estimation, planning, and visual tracking,” IEEE Transactions on Control Systems Technology,
control for high-speed flying through multiple openings,” The Inter- 2024.
national Journal of Robotics Research, vol. 39, no. 9, pp. 1122–1137, [51] Y. Jiang, H. Wang, and W. Yu, “Perception-aware model predictive
2020. control for target tracking with uavs,” in 14th Asian Control Confer-
[31] P. Serra, R. Cunha, T. Hamel, D. Cabecinhas, and C. Silvestre, ence, 2024, pp. 998–1003.
“Landing of a quadrotor on a moving target using dynamic image- [52] A. Altan and R. Hacıoğlu, “Model predictive control of three-axis
based visual servo control,” IEEE Transactions on Robotics, vol. 32, gimbal system mounted on uav for real-time target tracking under
no. 6, pp. 1524–1535, 2016. external disturbances,” Mechanical Systems and Signal Processing,
[32] D. Zheng, H. Wang, J. Wang, S. Chen, W. Chen, and X. Liang, “Image- vol. 138, p. 106548, 2020.
based visual servoing of a quadrotor using virtual camera approach,” [53] A. H. González and D. Odloak, “Robust model predictive controller
IEEE/ASME Transactions on Mechatronics, vol. 22, no. 2, pp. 972– with output feedback and target tracking,” IET Control Theory &
982, 2016. Applications, vol. 4, no. 8, pp. 1377–1390, 2010.
[33] J. Thomas, G. Loianno, K. Daniilidis, and V. Kumar, “Visual servoing [54] N. Sugie, “A model of predictive control in visual target tracking,”
of quadrotors for perching by hanging from cylindrical objects,” IEEE IEEE Transactions on Systems, Man, and Cybernetics, no. 1, pp. 2–7,
robotics and automation letters, vol. 1, no. 1, pp. 57–64, 2015. 2010.
[34] B. Ma, Z. Liu, W. Zhao, J. Yuan, H. Long, X. Wang, and Z. Yuan, [55] A. Saviolo, N. Picello, J. Mao, R. Verma, and G. Loianno, “Re-
“Target tracking control of uav through deep reinforcement learning,” active collision avoidance for safe agile navigation,” arXiv preprint
IEEE Transactions on Intelligent Transportation Systems, vol. 24, arXiv:2409.11962, 2024.
no. 6, pp. 5983–6000, 2023. [56] T. Do, L. C. Carrillo-Arce, and S. I. Roumeliotis, “High-speed
[35] A. Dionigi, M. Leomanni, A. Saviolo, G. Loianno, and G. Costante, autonomous quadrotor navigation through visual and inertial paths,”
“Exploring deep reinforcement learning for robust target tracking using The International Journal of Robotics Research, vol. 38, no. 4, pp.
micro aerial vehicles,” in 21st International Conference on Advanced 486–504, 2019.
Robotics, 2023, pp. 506–513. [57] A. Loquercio, A. Saviolo, and D. Scaramuzza, “Autotune: Controller
[36] Y. Mao, F. Gao, Q. Zhang, and Z. Yang, “An auv target-tracking tuning for high-speed flight,” IEEE Robotics and Automation Letters,
method combining imitation learning and deep reinforcement learn- vol. 7, no. 2, pp. 4432–4439, 2022.
20
[58] M. Kulkarni, B. Moon, K. Alexis, and S. Scherer, “Aerial field [82] M. Lyu, J. Yang, Z. Qi, R. Xu, and J. Liu, “Rigid pairwise 3d point
robotics,” arXiv preprint arXiv:2401.10837, 2024. cloud registration: A survey,” Pattern Recognition, p. 110408, 2024.
[59] B. Lindqvist, S. S. Mansouri, A.-a. Agha-mohammadi, and G. Niko-
lakopoulos, “Nonlinear mpc for collision avoidance and control of
uavs with dynamic obstacles,” IEEE Robotics and Automation Letters,
vol. 5, no. 4, pp. 6001–6008, 2020.
[60] Y. Li, G. Lu, D. He, and F. Zhang, “Robocentric model-based
visual servoing for quadrotor flights,” IEEE/ASME Transactions on
Mechatronics, vol. 28, no. 4, pp. 2155–2166, 2023.
[61] B. Zhang, X. Chen, Z. Li, G. Beltrame, C. Xu, F. Gao, and Y. Cao,
“Coni-mpc: Cooperative non-inertial frame based model predictive
control,” IEEE Robotics and Automation Letters, vol. 8, no. 12, pp.
8082–8089, 2023.
[62] B. Zhang, X. Chen, Q. Chen, C. Xu, F. Gao, and Y. Cao, “Global-
state-free obstacle avoidance for quadrotor control in air-ground co-
operation,” IEEE Robotics and Automation Letters, 2025.
[63] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan,
P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in
context,” in European Conference on Computer Vision. Springer,
2014, pp. 740–755.
[64] L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao,
“Depth anything v2,” arXiv preprint arXiv:2406.09414, 2024.
[65] L. Oth, P. Furgale, L. Kneip, and R. Siegwart, “Rolling shutter camera
calibration,” in IEEE Conference on Computer Vision and Pattern
Recognition, 2013, pp. 1360–1367.
[66] G. Loianno, M. Watterson, and V. Kumar, “Visual inertial odometry for
quadrotors on se (3),” in IEEE International Conference on Robotics
and Automation, 2016, pp. 1544–1551.
[67] J. H. Cartwright and O. Piro, “The dynamics of runge–kutta methods,”
International Journal of Bifurcation and Chaos, vol. 2, no. 03, pp.
427–449, 1992.
[68] A. Saviolo, J. Frey, A. Rathod, M. Diehl, and G. Loianno, “Active
learning of discrete-time dynamics for uncertainty-aware model pre-
dictive control,” IEEE Transactions on Robotics, 2023.
[69] A. Saviolo, G. Li, and G. Loianno, “Physics-inspired temporal learn-
ing of quadrotor dynamics for accurate model predictive trajectory
tracking,” IEEE Robotics and Automation Letters, 2022.
[70] L. Bauersfeld, E. Kaufmann, P. Foehn, S. Sun, and D. Scaramuzza,
“NeuroBEM: Hybrid Aerodynamic Quadrotor Model,” Robotics: Sci-
ence and Systems Foundation, 2021.
[71] F. Crocetti, J. Mao, A. Saviolo, G. Costante, and G. Loianno, “Gapt:
Gaussian process toolkit for online regression with application to
learning quadrotor dynamics,” arXiv preprint arXiv:2303.08181, 2023.
[72] A. Saviolo and G. Loianno, “Learning quadrotor dynamics for precise,
safe, and agile flight control,” Annual Reviews in Control, 2023.
[73] G. Jocher, J. Qiu, and A. Chaurasia, “Ultralytics YOLO,” Jan. 2023.
[Online]. Available: [Link]
[74] A. Grunnet-Jepsen, J. N. Sweetser, and J. Woodfill, “Best-known-
methods for tuning intel® realsense™ d400 depth cameras for best
performance,” Intel Corporation: Satan Clara, CA, USA, vol. 1, 2018.
[75] R. Verschueren, G. Frison, D. Kouzoupis, J. Frey, N. van Duijkeren,
A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, and M. Diehl,
“acados – a modular open-source framework for fast embedded
optimal control,” Mathematical Programming Computation, 2021.
[76] A. Maalouf, N. Jadhav, K. M. Jatavallabhula, M. Chahine, D. M. Vogt,
R. J. Wood, A. Torralba, and D. Rus, “Follow anything: Open-set
detection, tracking, and following in real-time,” IEEE Robotics and
Automation Letters, vol. 9, no. 4, pp. 3283–3290, 2024.
[77] Y. Liu, I. E. Zulfikar, J. Luiten, A. Dave, D. Ramanan, B. Leibe,
A. Ošep, and L. Leal-Taixé, “Opening up open world tracking,” in
IEEE/CVF conference on computer vision and pattern recognition,
2022, pp. 19 045–19 055.
[78] M. Nissov, N. Khedekar, and K. Alexis, “Degradation resilient
lidar-radar-inertial odometry,” in IEEE International Conference on
Robotics and Automation, 2024, pp. 8587–8594.
[79] B. Kim, M. B. Azhari, J. Park, and D. H. Shim, “An autonomous
uav system based on adaptive lidar inertial odometry for practical
exploration in complex environments,” Journal of Field Robotics,
vol. 41, no. 3, pp. 669–698, 2024.
[80] J. Zhang and S. Singh, “Low-drift and real-time lidar odometry and
mapping,” Autonomous robots, vol. 41, pp. 401–416, 2017.
[81] A. Alfarano, L. Maiano, L. Papa, and I. Amerini, “Estimating optical
flow: A comprehensive review of the state of the art,” Computer Vision
and Image Understanding, p. 104160, 2024.
21