InstantAvatar: Learning Avatars from Monocular Video in 60 Seconds
Tianjian Jiang1 *, Xu Chen1,2 *, Jie Song1 , Otmar Hilliges1
1
ETH Zürich 2 Max Planck Institute for Intelligent Systems, Tübingen
[Link]
arXiv:2212.10550v1 [[Link]] 20 Dec 2022
1 min 15 FPS
Training Rendering
Figure 1. InstantAvatar: we propose a system that can reconstruct animatable high-fidelity human avatars from monocular video within
60 seconds, providing poses and masks, and can animate and render the model at 15 FPS at 540 × 540 resolution. To achieve this we
integrate accelerated neural radiance fields, originally designed for rigid scenes, with a fast correspondence search module for articulation.
An efficient empty-space skipping strategy further speeds up training and inference, enabling near-instant avatar learning.
Abstract 1. Introduction
Creating high-fidelity digital humans is important
In this paper, we take a significant step towards real- for many applications including immersive tele-presence,
world applicability of monocular neural avatar reconstruc- AR/VR, 3D graphics, and the emerging metaverse. Cur-
tion by contributing InstantAvatar, a system that can re- rently acquiring personalized avatars is an involved process
construct human avatars from a monocular video within that typically requires the use of calibrated multi-camera
seconds, and these avatars can be animated and rendered systems and incurs significant computational cost. In this
at an interactive rate. To achieve this efficiency we pro- paper, we embark on the quest to build a system for the
pose a carefully designed and engineered system, that lever- learning of 3D virtual humans from monocular video alone
ages emerging acceleration structures for neural fields, in that is lightweight enough to be widely deployable and fast
combination with an efficient empty space-skipping strat- enough to allow for walk-up and use scenarios.
egy for dynamic scenes. We also contribute an efficient im- The emergence of powerful neural fields has enabled a
plementation that we will make available for research pur- number of methods for the reconstruction of animatable
poses. Compared to existing methods, InstantAvatar con- avatars from monocular videos of moving humans [1, 2,
verges 130× faster and can be trained in minutes instead 5, 47, 59]. These methods typically model human shape
of hours. It achieves comparable or even better reconstruc- and appearance in a pose-independent canonical space. To
tion quality and novel pose synthesis results. When given reconstruct the model from images that depict humans in
the same time budget, our method significantly outperforms different poses, such methods must use animation (e.g.
SoTA methods. InstantAvatar can yield acceptable visual skinning) and rendering algorithms, to deform and ren-
quality in as little as 10 seconds training time. der the model into posed space in a differentiable way.
This mapping between posed and canonical space allows
optimization of network weights by minimizing the dif-
ference between the generated pixel values and real im-
ages. Especially methods that leverage neural radiance
* Equal Contribution. fields (NeRFs) [38] as the canonical model have demon-
1
strated high-fidelity avatar reconstruction results. However, updated every few training iterations with the densities of
due to the dual need for differentiable deformation mod- randomly sampled points, in the posed space of randomly
ules and for volume rendering, these models require hours sampled frames. This scheme balances computational effi-
of training time and cannot be rendered at interactive rates, ciency and rendering quality.
prohibiting their broader application. We evaluate our method on both synthetic and real
In this paper, we aim to take a significant step towards monocular videos of moving humans, and compare it with
real-world applicability of monocular neural avatar recon- state-of-the-art methods on monocular avatar reconstruc-
struction by contributing a method that takes no longer tion. Our method achieves on-par reconstruction quality
for reconstruction, than it takes to capture the input video. and better animation quality in comparison to SoTA meth-
To this end, we propose InstantAvatar a system that re- ods, while only requiring minutes of training time instead
constructs high-fidelity avatars within 60 seconds, instead of more than 10 hours. When given the same time budget,
of hours, given a monocular video, pose parameters and our method significantly outperforms SoTA methods. We
masks. Once learned the avatar can be animated and ren- also provide an ablation study to demonstrate the effect of
dered at interactive rates. Achieving such a speed-up is our system’s components on speed and accuracy.
clearly a challenging task that requires careful method de-
sign, requires fast differentiable algorithms for rendering 2. Related Work
and articulation, and requires efficient implementation.
Our simple yet highly efficient pipeline combines sev- 3D Human Reconstruction Reconstructing 3D human
eral key components. First, to learn the canonical shape and appearance and shape is a long-standing problem. High-
appearance we leverage a recently proposed efficient neural quality reconstruction has been achieved in [10, 14, 18, 35]
radiance field variant [40]. Instant-NGP accelerates neu- by fusing observations from a dense array of cameras or
ral volume rendering by replacing multi-layer perceptrons depth sensors. The expensive hardware requirement limits
(MLP) with a more efficient hash table as data structure. such methods to professional settings. Recent work [1,2,17,
However, because the spatial features are represented ex- 19, 20, 27, 62] demonstrates 3D human reconstruction from
plicitly, Instant-NGP is limited to rigid objects. Second, to a monocular video by leveraging personalized or generic
enable learning from posed observations and to be able to template mesh models such as SMPL [34]. These meth-
animate the avatar, we interface the canonical NeRF with ods reconstruct 3D humans by deforming the template to fit
an efficient articulation module, Fast-SNARF [6], which ef- 2D joints and silhouettes. However, personalized template
ficiently derives a continuous deformation field to warp the mesh might not be available in many scenarios and generic
canonical radiance field into the posed space. Fast-SNARF template mesh cannot model high-fidelity details and differ-
is orders of magnitude faster compared to its much slower ent clothing topologies.
predecessor [8]. Recently, neural representations [36, 39, 43, 45] have
emerged as a powerful tool to model 3D humans [3, 5, 7–
Finally, simply integrating existing acceleration tech-
9, 11–13, 15, 17, 21–26, 29, 30, 33, 37, 41, 42, 44, 46, 47, 50,
niques is not sufficient to yield the desired efficiency (see
51, 55, 57–60, 60, 61, 64, 66]. Using neural representations,
Tab. 3). With acceleration structures for the canonical space
many works [5, 25, 26, 29, 33, 41, 46, 47, 58, 59, 61] can di-
and a fast articulation module in place, rendering the actual
rectly reconstruct high fidelity neural human avatars from
volume becomes the computational bottleneck. To compute
a sparse set of views or a monocular video without pre-
the color of a pixel, standard volume rendering needs to
scanning personalized template. These methods model 3D
query and accumulate densities of hundreds of points along
human shape and appearance via neural radiance field [38]
the ray. A common approach to accelerating this is to main-
or signed distance and texture field in a pose-independent
tain an occupancy grid to skip samples in the empty space.
canonical space and then deform and render the model into
However, such an approach assumes rigid scenes and can-
various body poses in order to learn from posed observa-
not be applied to dynamic scenes such as humans in motion.
tions. While achieving impressive quality and can learn
We propose an empty space skipping scheme that is de- avatars from a monocular video, these methods suffer from
signed for dynamic scenes with known articulation patterns. slow training and rendering speed due to the slow speed of
At inference time, for each input body pose, we sample the canonical representation as well as deformation algo-
points on a regular grid in posed space and map them back rithms. Our method addresses this issue and enables learn-
to the canonical model to query densities. Thresholding ing avatars within minutes.
these densities yields an occupancy grid in canonical space,
which can then be used to skip empty space during volume
rendering. For training, we maintain a shared occupancy Accelerating Neural Radiance Field Several methods
grid over all training frames, recording the union of occu- have been proposed to improve the training and inference
pied regions over individual frames. This occupancy grid is speed of neural representations [4, 16, 28, 31, 32, 40, 49,
2
52–54, 63]. The core idea is to replace MLPs in neu- its neighboring grid points and then concatenate the interpo-
ral representations with more efficient representations. A lated features at different levels. The concatenated features
few works [32, 52, 63] propose to use voxel grids to rep- are finally decoded with a shallow MLP.
resent neural fields and achieve fast training and inference
speed. Instant-NGP [40] further replaces dense voxels with Articulating Radiance Fields To create animations and
a multi-resolution hash table, which is more memory ef- to learn from posed images, we need to generate deformed
ficient and hence can record high-frequency details. Be- radiance fields in target poses fσ0 f . The posed radiance field
sides improving the efficiency of the representation, several is defined as
works [28, 31, 40] also improve the rendering efficiency by
skipping empty space via an occupancy grid to further in- fσ0 f : R3 → R+ , R3 (3)
crease training and inference speed. 0
x 7→ σ, c, (4)
While achieving impressive quality and training effi-
ciency, these methods are specifically designed for rigid ob- which outputs color and density for each point in posed
jects. Generalizing these methods to non-rigid objects is not space. We use a skinning weight field w in canonical space
straightforward. We combine Instant-NGP with a recent ar- to model articulation, with σw being its parameters:
ticulation algorithm to enable animation and learning from
posed observations. In addition, we propose an empty space wσw : R3 → Rnb , (5)
skinning scheme for dynamic articulated humans. x 7→ w1 , ..., wnb . (6)
3. Method where nb is the number of bones in the skeleton. To avoid
the computational cost of [8], [6] represents this skinning
Given a monocular video of a moving human, our pri- weight field as a low-resolution voxel grid. The value of
mary goal is to reconstruct a 3D human avatar within a each grid point is determined as the skinning weights of
tight computational budget. In this section, we first describe its nearest vertex on the SMPL [34] model. With this the
the preliminaries that our method is based on (Sec. 3.1), canonical skinning weight field and target bone transforma-
which include an accelerated neural radiance field that we tions B = {B1 , ..., Bnb }, a point x in canonical space is
use to model the appearance and shape in canonical space transformed to deformed space x0 via linear blend skinning
and an efficient articulation module to deform the canonical as following:
radiance field into posed space. We then describe our im- Pnb
plementation of the volumetric renderer to produce images x0 = i=1 wi Bi x (7)
from the radiance fields in an efficient manner (Sec. 3.2).
The canonical correspondences x∗ of a deformed point x0
To avoid inefficient sampling of empty space, we leverage
are defined by the inverse mapping of Equation. 7. The key
the observation that the 3D bounding box around the hu-
is to establish the mapping from points in posed space x0
man body is dominated by empty space. We then propose
to their correspondences in the canonical space x∗ . This
an empty space skipping scheme specifically designed for
is efficiently derived by root-finding in Fast-SNARF [6].
humans (Sec. 3.3). Finally, we discuss training objectives
The posed radiance field fσ0 f can then be determined as
and regularization strategies (Sec. 3.4).
fσ0 f (x0 ) = fσf (x∗ ).
3.1. Preliminaries 3.2. Rendering Radiance Fields
Efficient Canonical Neural Radiance Field We model The articulated radiance field fσ0 f can be rendered into
human shape and appearance in a canonical space using a novel views via volume rendering. Given a pixel, we cast a
radiance field fσf , which predicts the density σ and color c ray r = o + td with o being the camera center and d be-
of each 3D point x in the canonical space: ing the ray direction. We sample N points {x0i }N along the
ray between the near and far bound, and query the color and
fσf : R3 → R+ , R3 (1) density of each point from the articulated radiance field fσ0 f
x 7→ σ, c (2) by mapping {x0i }N back to the canonical space and query-
ing from the canonical NeRF model fσf , as illustrate in
where σf are the parameters of the radiance field. Fig. 2. We then accumulate queried radiance and density
We use Instant-NGP [40] to parameterize fσf , which along the ray to get the pixel color C
achieves fast training and inference speed by using a hash
table to store feature grids at different coarseness scales. To N
X Y
predict the texture and geometry properties of a query point C= αi (1 − αj )ci , with αi = 1 − exp(σi δi ) (8)
in space, they read and tri-linearly interpolate the features at i=1 j<i
3
! = {!! , "" , … , "#! }
Rotation Body
Translation Pose
Articulation
Rigid
(′ c, σ
Tansform (
Camera Posed Space Normalized Space Empty Space Skipping Canonical Space Instant-NGP %
Figure 2. Method Overview. For each frame, we sample points along the rays in posed space. We then transform these points into a
normalized space where we remove the global orientation and translation of the person. In this normalized space, we filter points in empty
space using our occupancy grid. The remaining points are deformed to canonical space using an articulation module and then fed into the
canonical neural radiance field to evaluate the color and density.
where δi = kx0i+1 − x0i k is the distance between samples. Training Stage During training, however, the overhead to
While the acceleration modules of Sec. 3.1 already construct such an occupancy grid at each training iteration
achieve significant speed-up over the vanilla variants is no longer negligible. To avoid this overhead, we con-
(NeRF [38], SNARF [8]), the rendering itself now becomes struct a single occupancy grid for the entire sequence by
the bottleneck. In this paper, we optimize the process of recording the union of occupied regions in each of the in-
neural rendering, specifically for the use-case of dynamic dividual frames. Specifically, we build an occupancy grid
humans. at the start of training and update it every k iterations, by
taking the moving average of the current occupancy values
3.3. Empty Space Skipping for Dynamic Objects and the densities queried from the posed radiance field fσ0 f
We note that the 3D bounding box surrounding the hu- at the current iteration. Note that this occupancy grid is de-
man body is dominated by empty space due to the articu- fined in a normalized space where the global orientation and
lated structure of 3D human limbs. This results in a large translation are factored out so that the union of the occupied
amount of redundant sample queries during rendering and space is as tight as possible and hence unnecessary queries
hence significantly slows down rendering. For rigid objects, are further reduced.
this problem is eliminated by caching a coarse occupancy
grid and skipping samples within non-occupied grid cells. 3.4. Training Losses
However, for dynamic objects, the exact location of empty We train our model by minimizing the robust Huber loss
space varies across different frames, depending on the pose. ρ between the predicted color of the pixels C and the corre-
sponding ground-truth color Cgt :
Inference Stage At inference time, for each input body
pose, we sample points on a 64 × 64 × 64 grid in posed Lrgb = ρ(kC − Cgt k) (9)
space and query their densities from the posed radiance field
fσ0 f . We then threshold these densities into binary occu- In addition, we assume an estimate of the human mask
pancy values. To remove cells that have been falsely la- available and apply a loss on the rendered 2D alpha values,
beled as empty, due to the low spatial resolution, we dilate in order to reduce floating artifacts in space.
the occupied region to fully cover the subject. Due to the
low resolution of this grid and the large amount of queries Lalpha = kα − αgt k1 (10)
required to render an image, the overhead to construct such
an occupancy grid is negligible.
During volumetric rendering, for point samples inside Hard Surface Regularization Following [48], we add
the non-occupied cells, we directly set their density to zero further regularization to encourage the NeRF model to pre-
without querying the posed radiance field fσ0 f . This re- dict solid surfaces:
duces unnecessary computation to a minimum and hence
improves the inference speed. Lhard = − log(exp−|α| + exp−|α−1| ) + const. (11)
4
where const. is a constant to ensure loss value to be non- Baselines
negative. Encouraging solid surfaces helps to speed up ren-
We consider the following methods as our baselines:
dering because we can terminate rays early once the accu-
mulated opacity reaches 1.
Anim-NeRF [5] This baseline models human shapes and
appearance in a canonical space with an MLP-based NeRF.
Occupancy-based regularization Previous methods for Given a pose, they first generate a SMPL body in the tar-
the learning of human avatars [5, 26] often encourage mod- get pose. Then for each query point in deformed space, its
els to predict zero density for points outside of the surface corresponding skinning weights are defined as the weighted
and solid density for points inside the surface by leverag- average of skinning weights of its K nearest vertices on the
ing the SMPL body model as regularizer. This is done to posed SMPL mesh. Finally, with the skinning weights, the
reduce artifacts near the body surface. However such regu- query point can be transformed back to the canonical space
larization makes heavy assumptions about the shape of the based on inverse LBS.
body and does not generalize well for loose clothing. More-
over, we empirically found this regularization is not effec-
tive in removing artifacts near the body. This can be seen Neural Body [47] This baseline learns a set of latent
in Fig. 3. Instead of using SMPL for regularization, we use codes anchored to a deformable SMPL mesh. These latent
our occupancy grid which is a more conservative estimate codes deform with the SMPL mesh and are decoded into
of the shape of the subject and the clothing, and define an radiance fields in different poses.
additional loss Lreg which encourages the points inside the 4.2. Comparison with SoTA
empty cells of the occupancy grid to have zero density:
( Reconstruction Quality To measure the appearance
|σ(x)| if x is in the empty space quality of the reconstructed avatar, we animate and ren-
Ldensity = (12) der the reconstructed model with the poses of test frames
0 otherwise
in PeopleSnapshot, and measure the difference between the
generated images and the real images. When training all
4. Experiments methods to convergence, our generated images are signif-
We evaluate the accuracy and speed of our method on icantly better than Neural Body [47] and achieve on-par
monocular videos and compare it with other SoTA methods. quality as SoTA method Anim-NeRF [5], as indicated by
In addition, we provide an ablation study to investigate the the image quality metrics in Tab. 1 and the qualitative re-
effect of individual technical contributions. sults in Fig. 3.
4.1. Evaluation Setting Speed Our method requires much less training time and
Datasets computation resources than SoTA methods. We only re-
quire 1 minute training time on a single RTX 3090 while
PeopleSnapshot We conduct experiments on the Peo- Anim-NeRF [5] requires 13 hours on 2× RTX 3090 and
pleSnapshot [1] dataset, which contains videos of humans Neural Body [47] requires 14 hours on 4× RTX 2080. Our
rotating in front of a camera. We follow the evaluation pro- method also achieves superior rendering speed - we can ren-
tocol defined in Anim-NeRF [5]. Because the pose param- der images at 540 × 540 resolution on a single RTX 3090 at
eters provided in this dataset are not perfect and do not al- 15 FPS, which is orders of magnitude faster than baselines.
ways align with the image, Anim-NeRF optimizes the poses Given the same training time budget, our method
of training and test frames. We train our model with the achieves significantly better image quality than Anim-NeRF
pose parameters optimized by Anim-NeRF and keep them as shown in Tab. 1. Comparing our training progression
frozen throughout training for a fair comparison. with Anim-NeRF in Fig. 4, we note that our method already
learns meaningful appearance and moderate details within
SURREAL The PeopleSnapshot dataset has limited pose 5s and acceptable visual quality at 10s. After only 1 minute
variations. To evaluate the performance on more challeng- of training time, our method already achieves high-fidelity
ing test poses, we also generate synthetic monocular se- reconstruction quality. In contrast, Anim-NeRF does not
quences by rendering SMPL with texture maps from the produce meaningful results this early in training and only
SURREAL [56] dataset. For training, we drive the textured learns the rough shape after 3 minutes.
SMPL model with the same SMPL parameters from Peo-
pleSnapshot, and for test, we generate challenging out-of- Novel Pose Synthesis Quality The previous evaluation
distribution poses. This allows us to evaluate the perfor- does not reflect the performance of novel pose synthesis,
mance of methods on novel pose synthesis. because the pose variation in the PeopleSnapshot dataset
5
male-3-casual male-4-casual female-3-casual female-4-casual
PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓
Neural Body [47] (∼ 14 hours) 24.94 0.9428 0.0326 24.71 0.9469 0.0423 23.87 0.9504 0.0346 24.37 0.9451 0.0382
Anim-NeRF [5] (∼ 13 hours) 29.37 0.9703 0.0168 28.37 0.9605 0.0268 28.91 0.9743 0.0215 28.90 0.9678 0.0174
Ours (1 minute) 29.65 0.9730 0.0192 27.97 0.9649 0.0346 27.90 0.9722 0.0249 28.92 0.9692 0.0180
Anim-NeRF [5] (5 minutes) 23.17 0.9266 0.0784 22.30 0.9235 0.0911 22.37 0.9311 0.0784 23.18 0.9292 0.0687
Ours (5 minutes) 29.53 0.9716 0.0155 27.67 0.9626 0.0307 27.66 0.9709 0.0210 29.11 0.9683 0.0167
Anim-NeRF [5] (3 minutes) 19.75 0.8927 0.1286 20.66 0.8986 0.1414 19.77 0.9003 0.1255 20.20 0.9044 0.1109
Ours (3 minutes) 29.58 0.9719 0.0157 27.83 0.9640 0.0342 27.68 0.9708 0.0217 29.05 0.9689 0.0263
Anim-NeRF [5] (1 minute) 12.39 0.7929 0.3393 13.10 0.7705 0.3460 11.71 0.7797 0.3321 12.31 0.8089 0.3344
Ours (1 minute) 29.65 0.9730 0.0192 27.97 0.9649 0.0346 27.90 0.9722 0.0249 28.92 0.9692 0.0180
Table 1. Qualitative Comparison with SoTA on the PeopleSnapshot [1] dataset. We report PSNR, SSIM and LPIPS [65] between real
images and the images generated by our method and two SoTA methods, Neural Body [47] and Anim-NeRF [5]. We compare all three
methods at their convergence, and also compare ours with Anim-NeRF at 5 minutes, 3 minutes and 1 minute training time.
Input View 1 View 2 Pose 1 Pose 2 Pose 3
Anim-NeRF
Ours
Anim-NeRF
Ours
Figure 3. Qualitative Results on SURREAL [56] and PeopleSnapshot dataset [1]. We show reconstructed avatars on SURREAL (top)
and PeopleSnapshot (bottom) from different viewpoints (column 2-3) and in various poses (column 4-6).
is limited (self-rotating). Due to the lack of ground truth In contrast, Anim-NeRF suffers from artifacts under arms
images in novel poses, we resort to evaluating novel pose and between legs, because their methods cannot correctly
synthesis qualitatively. We generate images in novel chal- disambiguate body parts that are close to each other in the
lenging poses with our method and Anim-NeRF. As shown posed space. Our method outperforms our baseline espe-
in Fig. 3, our method can faithfully generate images even cially for loose clothing as shown in the bottom example
in challenging body poses while preserving high fidelity. in Fig. 3. This is because we don’t rely on the SMPL body
6
5s 10s 1 min 3 min 5 min
Anim-NeRF
Ours
Figure 4. Training Progression. We show the image quality at different training iterations. Our method converges significantly faster than
SoTA Anim-NeRF [5].
Anim-NeRF Ours
PSNR SSIM LPIPS PSNR SSIM LPIPS
S1 21.66 0.9450 0.07615 24.48 0.9353 0.0304
S2 20.00 0.9483 0.09693 23.94 0.9354 0.0343
S3 20.06 0.9326 0.07948 25.08 0.9494 0.0275
Table 2. Qualitative Results on the SURREAL Dataset. We
evaluate novel pose synthesis quality of our method and Anim-
NeRF [5] on 3 synthetic subjects. No Reg Ours Global Sparsity
Figure 5. Effect of Occupancy-based Regularization. With-
Training Rendering out our regularization loss, the model suffers from floating arti-
facts. Our occupancy-based regularization loss successfully re-
w/o empty space skipping 3m 10s ∼ 1 FPS
moves such artifacts. While a global sparsity prior biasing all den-
w/ empty space skipping 1m 47s ∼ 15 FPS
sities towards 0 can also reduce such artifacts, it leads to degener-
ated image quality (semi-transparent).
Table 3. Empty Space Skipping. We compare the training and
rendering speed with and without empty space skipping. For the
training time we report the average training time of 100 epochs
Fig. 3 verify the superiority of our method in terms of novel
among 4 sequences in PeopleSnapshot.
pose synthesis quality.
PSNR SSIM LPIPS 4.3. Ablation Study
w/o occupancy-based regularizer 28.22 0.9680 0.0301 Empty Space Skipping We study the effect of our pro-
w/ occupancy-based regularizer 28.64 0.9700 0.0240 posed empty space skipping scheme for dynamic objects.
As shown in Tab. 3, skipping empty space significantly im-
Table 4. Occupancy-based Regularizer. We evaluate image qual- proves the training and rendering speed.
ity averaged over the 4 PeopleSnapshot sequences. For both cases
we train our model for 100 epochs.
Occupancy-based Regularization Lreg The occupancy
grid for empty space skipping can also help regularize the
model for regularization and hence can better deal with sub- radiance field to reduce noise via our regularization loss
jects and clothing that differ from SMPL. To quantitatively Lreg described in Section. 3.4. As shown in Fig. 5, this loss
evaluate novel pose synthesis, we generate synthetic data in effectively reduces floating artifacts and consequently helps
challenging poses as ground-truth. The results in Tab. 2 and to improve the overall image quality as evidenced by the
7
Input View 1 View 2 View 3 Novel Pose
Figure 6. More Qualitative Results of Our Method.
PSNR improvement in Tab. 4. Another common approach at 15 FPS. To achieve this, we combine an efficient neu-
to reducing floating noise is to encourage zero density for ral representation, Instant-NGP [40], and an efficient artic-
every point in space. We compare our solution with this ulation module Fast-SNARF [6]. This naive combination
strategy and find that this strategy (Global Sparsity) leads does not yield optimal speed. We propose an empty space
to degenerated image quality as shown in Fig. 5. skipping scheme to improve our rendering speed, and an
occupancy-aware regularization loss to reduce floating ar-
4.4. Additional Qualitative Samples tifacts in space. In comparison with SoTA methods, our
We show additional qualitative samples in Fig. 6. method achieves on-par image quality while being signifi-
cantly faster during training and inference.
5. Conclusion
Limitations and Future Work Our method reconstructs
In this paper, we propose a method that can reconstruct avatars purely based on image observations and therefore
animatable human avatars from monocular videos within 60 cannot infer unseen regions. For instance, if the input video
seconds and can animate and render the model afterward only captures the front side of the subject, our method can-
8
not reconstruct the back side. This limitation could poten-
tially be addressed by leveraging learning-based methods
to predict the texture and geometry of unobserved regions.
While this paper focuses on full-body human reconstruc-
tion, the idea could be applied to other objects. An interest-
ing next step is to extend our method to reconstruct general
articulated objects or animals from images efficiently.
Acknowledgements Xu Chen was supported by the Max
Planck ETH Center for Learning Systems.
9
References Escolano, Christoph Rhemann, David Kim, Jonathan Taylor,
et al. Fusion4d: Real-time performance capture of challeng-
[1] Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian ing scenes. ACM Trans. on Graphics, 35, 2016. 2
Theobalt, and Gerard Pons-Moll. Detailed human avatars [15] Yao Feng, Jinlong Yang, Marc Pollefeys, Michael J. Black,
from monocular video. In Proc. of the International Conf. and Timo Bolkart. Capturing and animation of body and
on 3D Vision (3DV), 2018. 1, 2, 5, 6 clothing from monocular video. In SIGGRAPH, 2022. 2
[2] Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian [16] Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie
Theobalt, and Gerard Pons-Moll. Video based reconstruc- Shotton, and Julien Valentin. Fastnerf: High-fidelity neural
tion of 3d people models. In Proc. IEEE Conf. on Computer rendering at 200fps. In Proc. IEEE Conf. on Computer Vision
Vision and Pattern Recognition (CVPR), 2018. 1, 2 and Pattern Recognition (CVPR), 2021. 2
[3] Alexander W. Bergman, Petr Kellnhofer, Wang Yifan, [17] Chen Guo, Xu Chen, Jie Song, and Otmar Hilliges. Human
Eric R. Chan, David B. Lindell, and Gordon Wetzstein. Gen- performance capture from monocular video in the wild. In
erative neural articulated radiance fields. Arxiv, 2022. 2 2021 International Conference on 3D Vision (3DV), 2021. 2
[4] Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and [18] Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch,
Hao Su. Tensorf: Tensorial radiance fields. In Proc. of the Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts-
European Conf. on Computer Vision (ECCV), 2022. 2 Escolano, Rohit Pandey, Jason Dourgarian, et al. The re-
[5] Jianchuan Chen, Ying Zhang, Di Kang, Xuefei Zhe, Linchao lightables: Volumetric performance capture of humans with
Bao, Xu Jia, and Huchuan Lu. Animatable neural radiance realistic relighting. ACM Trans. on Graphics, 38, 2019. 2
fields from monocular rgb videos. [Link], 2021. 1, 2, 5, [19] Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger-
6, 7 ard Pons-Moll, and Christian Theobalt. Livecap: Real-time
[6] Xu Chen, Tianjian Jiang, Jie Song, Max Rietmann, An- human performance capture from monocular video. ACM
dreas Geiger, Michael J. Black, and Otmar Hilliges. Fast- Trans. on Graphics, 38, 2019. 2
SNARF: A fast deformer for articulated neural fields. arXiv, [20] Marc Habermann, Weipeng Xu, Michael Zollhoefer, Ger-
abs/2211.15601, 2022. 2, 3, 8 ard Pons-Moll, and Christian Theobalt. DeepCap: Monocu-
[7] Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael J lar human performance capture using weak supervision. In
Black, Andreas Geiger, and Otmar Hilliges. gdna: Towards Proc. IEEE Conf. on Computer Vision and Pattern Recogni-
generative detailed neural avatars. In Proc. IEEE Conf. on tion (CVPR), 2020. 2
Computer Vision and Pattern Recognition (CVPR), 2022. 2 [21] Tong He, John Collomosse, Hailin Jin, and Stefano Soatto.
[8] Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, Geo-PIFu: Geometry and pixel aligned implicit functions for
and Andreas Geiger. Snarf: Differentiable forward skinning single-view human reconstruction. In Advances in Neural
for animating non-rigid neural implicit shapes. In Proc. of Information Processing Systems (NeurIPS), 2020. 2
the IEEE International Conf. on Computer Vision (ICCV), [22] Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, and
2021. 2, 3, 4 Tony Tung. ARCH++: Animation-ready clothed human re-
[9] Julian Chibane, Thiemo Alldieck, and Gerard Pons-Moll. construction revisited. In Proc. IEEE Conf. on Computer
Implicit functions in feature space for 3D shape reconstruc- Vision and Pattern Recognition (CVPR), 2021. 2
tion and completion. In Proc. IEEE Conf. on Computer Vi- [23] Fangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang
sion and Pattern Recognition (CVPR), 2020. 2 Cai, Lei Yang, and Ziwei Liu. Avatarclip: Zero-shot text-
[10] Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Den- driven generation and animation of 3d avatars. ACM Trans.
nis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, on Graphics, 2022. 2
and Steve Sullivan. High-quality streamable free-viewpoint [24] Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and
video. ACM Trans. on Graphics, 34, 2015. 2 Tony Tung. Arch: Animatable reconstruction of clothed hu-
[11] Enric Corona, Albert Pumarola, Guillem Alenyà, Ger- mans. In Proc. IEEE Conf. on Computer Vision and Pattern
ard Pons-Moll, and Francesc Moreno-Noguer. SMPLicit: Recognition (CVPR), 2020. 2
Topology-aware generative model for clothed people. In [25] Boyi Jiang, Yang Hong, Hujun Bao, and Juyong Zhang. Sel-
Proc. IEEE Conf. on Computer Vision and Pattern Recog- frecon: Self reconstruction your digital avatar from monoc-
nition (CVPR), 2021. 2 ular video. In Proc. IEEE Conf. on Computer Vision and
[12] Boyang Deng, JP Lewis, Timothy Jeruzalski, Gerard Pons- Pattern Recognition (CVPR), 2022. 2
Moll, Geoffrey Hinton, Mohammad Norouzi, and Andrea [26] Wei Jiang, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel,
Tagliasacchi. Neural articulated shape approximation. In and Anurag Ranjan. Neuman: Neural human radiance field
Proc. of the European Conf. on Computer Vision (ECCV), from a single video. In Proc. of the European Conf. on Com-
2020. 2 puter Vision (ECCV), 2022. 2, 5
[13] Zijian Dong, Chen Guo, Jie Song, Xu Chen, Andreas Geiger, [27] Yue Jiang, Marc Habermann, Vladislav Golyanik, and Chris-
and Otmar Hilliges. PINA: Learning a personalized implicit tian Theobalt. Hifecap: Monocular high-fidelity and expres-
neural avatar from a single RGB-D video sequence. In Proc. sive capture of human performances. In Proc. of the British
IEEE Conf. on Computer Vision and Pattern Recognition Machine Vision Conf. (BMVC), 2022. 2
(CVPR), 2022. 2 [28] Ruilong Li, Matthew Tancik, and Angjoo Kanazawa. Ner-
[14] Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip facc: A general nerf accleration toolbox. arXiv preprint
Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts arXiv:2210.04847, 2022. 2, 3
10
[29] Ruilong Li, Julian Tanke, Minh Vo, Michael Zollhoefer, [43] Michael Oechsle, Lars Mescheder, Michael Niemeyer, Thilo
Jürgen Gall, Angjoo Kanazawa, and Christoph Lassner. Strauss, and Andreas Geiger. Texture fields: Learning tex-
Tava: Template-free animatable volumetric actors. In Proc. ture representations in function space. In Proc. of the IEEE
of the European Conf. on Computer Vision (ECCV), 2022. 2 International Conf. on Computer Vision (ICCV), 2019. 2
[30] Siyou Lin, Hongwen Zhang, Zerong Zheng, Ruizhi Shao, [44] Pablo Palafox, Aljaž Božič, Justus Thies, Matthias Nießner,
and Yebin Liu. Learning implicit templates for point-based and Angela Dai. Neural parametric models for 3D de-
clothed human modeling. In Proc. of the European Conf. on formable shapes. In Proc. of the IEEE International Conf.
Computer Vision (ECCV), 2022. 2 on Computer Vision (ICCV), 2021. 2
[31] Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and [45] Jeong Joon Park, Peter Florence, Julian Straub, Richard
Christian Theobalt. Neural sparse voxel fields. In Advances Newcombe, and Steven Lovegrove. DeepSDF: Learning
in Neural Information Processing Systems (NeurIPS), 2020. continuous signed distance functions for shape representa-
2, 3 tion. In Proc. IEEE Conf. on Computer Vision and Pattern
[32] Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Recognition (CVPR), 2019. 2
Christian Theobalt. Neural sparse voxel fields. Advances in [46] Sida Peng, Junting Dong, Qianqian Wang, Shangzhan
Neural Information Processing Systems (NeurIPS), 2020. 2, Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Ani-
3 matable neural radiance fields for modeling dynamic human
[33] Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu bodies. In Proc. of the IEEE International Conf. on Com-
Sarkar, Jiatao Gu, and Christian Theobalt. Neural Actor: puter Vision (ICCV), 2021. 2
Neural free-view synthesis of human actors with pose con- [47] Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang,
trol. ACM Trans. on Graphics, 2021. 2 Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body:
[34] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Implicit neural representations with structured latent codes
Pons-Moll, and Michael J. Black. SMPL: A skinned multi- for novel view synthesis of dynamic humans. In Proc. IEEE
person linear model. ACM Trans. on Graphics, 2015. 2, 3 Conf. on Computer Vision and Pattern Recognition (CVPR),
[35] Wojciech Matusik, Chris Buehler, Ramesh Raskar, Steven J 2021. 1, 2, 5, 6
Gortler, and Leonard McMillan. Image-based visual hulls. [48] Daniel Rebain, Mark Matthews, Kwang Moo Yi, Dmitry La-
In Proceedings of the 27th annual conference on Computer gun, and Andrea Tagliasacchi. LOLNeRF: Learn from one
graphics and interactive techniques, 2000. 2 look. In Proc. IEEE Conf. on Computer Vision and Pattern
[36] Lars Mescheder, Michael Oechsle, Michael Niemeyer, Se- Recognition (CVPR), 2022. 4
bastian Nowozin, and Andreas Geiger. Occupancy networks: [49] Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas
Learning 3D reconstruction in function space. In Proc. IEEE Geiger. Kilonerf: Speeding up neural radiance fields with
Conf. on Computer Vision and Pattern Recognition (CVPR), thousands of tiny mlps. In Proc. of the IEEE International
2019. 2 Conf. on Computer Vision (ICCV), 2021. 2
[37] Marko Mihajlovic, Yan Zhang, Michael J Black, and Siyu [50] Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Mor-
Tang. LEAP: Learning articulated occupancy of people. In ishima, Angjoo Kanazawa, and Hao Li. PIFu: Pixel-aligned
Proc. IEEE Conf. on Computer Vision and Pattern Recogni- implicit function for high-resolution clothed human digitiza-
tion (CVPR), 2021. 2 tion. In Proc. of the IEEE International Conf. on Computer
[38] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Vision (ICCV), 2019. 2
Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: [51] Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul
Representing scenes as neural radiance fields for view syn- Joo. PIFuHD: Multi-level pixel-aligned implicit function for
thesis. In Proc. of the European Conf. on Computer Vision high-resolution 3D human digitization. In Proc. IEEE Conf.
(ECCV), 2020. 1, 2, 4 on Computer Vision and Pattern Recognition (CVPR), 2020.
[39] Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, 2
Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: [52] Sara Fridovich-Keil and Alex Yu, Matthew Tancik, Qinhong
Representing scenes as neural radiance fields for view syn- Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels:
thesis. In Proc. of the European Conf. on Computer Vision Radiance fields without neural networks. In Proc. IEEE
(ECCV), 2020. 2 Conf. on Computer Vision and Pattern Recognition (CVPR),
[40] Thomas Müller, Alex Evans, Christoph Schied, and Alexan- 2022. 2, 3
der Keller. Instant neural graphics primitives with a multires- [53] Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel
olution hash encoding. SIGGRAPH, 41, 2022. 2, 3, 8 grid optimization: Super-fast convergence for radiance fields
[41] Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya reconstruction. In Proc. IEEE Conf. on Computer Vision and
Harada. Neural articulated radiance field. In Proc. of the Pattern Recognition (CVPR), 2022. 2
IEEE International Conf. on Computer Vision (ICCV), 2021. [54] Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten
2 Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson,
[42] Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Morgan McGuire, and Sanja Fidler. Neural geometric level
Harada. Unsupervised learning of efficient geometry-aware of detail: Real-time rendering with implicit 3D shapes. In
neural articulated representations. In Proc. of the European Proc. IEEE Conf. on Computer Vision and Pattern Recogni-
Conf. on Computer Vision (ECCV), 2022. 2 tion (CVPR), 2021. 2
11
[55] Garvita Tiwari, Nikolaos Sarafianos, Tony Tung, and Gerard
Pons-Moll. Neural-GIF: Neural generalized implicit func-
tions for animating people in clothing. In Proc. of the IEEE
International Conf. on Computer Vision (ICCV), 2021. 2
[56] Gül Varol, Javier Romero, Xavier Martin, Naureen Mah-
mood, Michael J. Black, Ivan Laptev, and Cordelia Schmid.
Learning from synthetic humans. In Proc. IEEE Conf. on
Computer Vision and Pattern Recognition (CVPR), 2017. 5,
6
[57] Shaofei Wang, Marko Mihajlovic, Qianli Ma, Andreas
Geiger, and Siyu Tang. MetaAvatar: Learning animatable
clothed human models from few depth images. In Advances
in Neural Information Processing Systems (NeurIPS), 2021.
2
[58] Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu
Tang. Arah: Animatable volume rendering of articulated
human sdfs. In Proc. of the European Conf. on Computer
Vision (ECCV), 2022. 2
[59] Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan,
Jonathan T. Barron, and Ira Kemelmacher-Shlizerman. Hu-
manNeRF: Free-viewpoint rendering of moving people from
monocular video. In Proc. IEEE Conf. on Computer Vision
and Pattern Recognition (CVPR), 2022. 1, 2
[60] Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J
Black. ICON: Implicit Clothed humans Obtained from Nor-
mals. In Proc. IEEE Conf. on Computer Vision and Pattern
Recognition (CVPR), 2022. 2
[61] Hongyi Xu, Thiemo Alldieck, and Cristian Sminchisescu. H-
NeRF: Neural radiance fields for rendering and temporal re-
construction of humans in motion. In Advances in Neural
Information Processing Systems (NeurIPS), 2021. 2
[62] Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer, Helge
Rhodin, Dushyant Mehta, Hans-Peter Seidel, and Christian
Theobalt. Monoperfcap: Human performance capture from
monocular video. SIGGRAPH, 37, 2018. 2
[63] Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and
Angjoo Kanazawa. Plenoctrees for real-time rendering of
neural radiance fields. In Proc. of the IEEE International
Conf. on Computer Vision (ICCV), 2021. 2, 3
[64] Jianfeng Zhang, Zihang Jiang, Dingdong Yang, Hongyi Xu,
Yichun Shi, Guoxian Song, Zhongcong Xu, Xinchao Wang,
and Jiashi Feng. Avatargen: A 3d generative model for ani-
matable human avatars. Arxiv, 2022. 2
[65] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht-
man, and Oliver Wang. The unreasonable effectiveness of
deep features as a perceptual metric. In Proc. IEEE Conf. on
Computer Vision and Pattern Recognition (CVPR), 2018. 6
[66] Zerong Zheng, Tao Yu, Yebin Liu, and Qionghai Dai.
PaMIR: Parametric model-conditioned implicit representa-
tion for image-based human reconstruction. IEEE Trans. on
Pattern Analysis and Machine Intelligence (PAMI), 2021. 2
12