0% found this document useful (0 votes)
8 views10 pages

Physically-Based Face Model Learning

This document presents a framework for learning physically-based face models using a dataset of 4000 high-resolution facial scans, enabling the generation of diverse face geometries and material attributes. The model integrates deep learning techniques to ensure anatomical correctness and high fidelity in rendering, with applications in various fields including VFX and biometrics. Key contributions include the creation of a generative face model that allows for real-time rendering and manipulation of identity and expressions, while addressing limitations of existing morphable face models.

Uploaded by

diego silva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views10 pages

Physically-Based Face Model Learning

This document presents a framework for learning physically-based face models using a dataset of 4000 high-resolution facial scans, enabling the generation of diverse face geometries and material attributes. The model integrates deep learning techniques to ensure anatomical correctness and high fidelity in rendering, with applications in various fields including VFX and biometrics. Key contributions include the creation of a generative face model that allows for real-time rendering and manipulation of identity and expressions, while addressing limitations of existing morphable face models.

Uploaded by

diego silva
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Learning Formation of Physically-Based Face Attributes

Ruilong Li1,2∗ Karl Bladin1∗ Yajie Zhao1∗ Chinmay Chinara1 Owen Ingraham1
Pengda Xiang1,2 Xinglei Ren Pratusha Prasad1
1
Bipin Kishore Jun Xing1 Hao Li 1,2,3
1

1 2 3
USC Institute for Creative Technologies University of Southern California Pinscreen
Identity
Expression

IdNet TexNet

Identity latent

ExpNet Geometry
Map
Expression latent Final Maps

(a) (b) (c) (d)

We introduce a comprehensive framework for learning physically based face models from highly constrained facial scan data. Our deep
learning based approach for 3D morphable face modeling seizes the fidelity of nearly 4000 high resolution face scans encompassing
expression and identity separation (a). The model (b) combines a multitude of anatomical and physically based face attributes to generate
an infinite number of digitized faces (c). Our model generates faces at pore level geometry resolution (d).

Abstract 1. Introduction
Based on a combined data set of 4000 high resolution
Graphical virtual representations of humans are at the
facial scans, we introduce a non-linear morphable face
center of many endeavors in the fields of computer vision
model, capable of producing multifarious face geometry of
and graphics, with applications ranging from cultural me-
pore-level resolution, coupled with material attributes for
dia such as video games, film, and telecommunication to
use in physically-based rendering. We aim to maximize the
medical, biometric modeling, and forensics [6].
variety of the participant’s face identities, while increasing
the robustness of correspondence between unique compo- Designing, modeling, and acquiring high fidelity data for
nents, including middle-frequency geometry, albedo maps, face models of virtual characters is costly and requires spe-
specular intensity maps and high-frequency displacement cialized scanning equipment and a team of skilled artists
details. Our deep learning based generative model learns to and engineers [17, 5, 37]. Due to limiting and restrictive
correlate albedo and geometry, which ensures the anatom- data policies of VFX studios, in conjunction with the ab-
ical correctness of the generated assets. We demonstrate sence of a shared platform that regards the sovereignty of,
potential use of our generative model for novel identity gen- and incentives for the individuals’ data contributions, there
eration, model fitting, interpolation, animation, high fidelity is a large discrepancy in the fidelity of models trained on
data visualization, and low-to-high resolution data domain publicly available data, and those used in large budget game
transferring. We hope the release of this generative model and film production. A single, unified model would democ-
will encourage further cooperation between all graphics, ratize the use of generated assets, shorten production cycles
vision, and data focused professionals, while demonstrating and boost quality and consistency, while incentivizing inno-
the cumulative value of every individual’s complete biomet- vative applications in many markets and fields of research.
ric profile. The unification of a facial scan data set in a 3D mor-
phable face model (3DMM) [7, 12, 41, 6] promotes the fa-
vorable property of representing facial scan data in a com-
∗ Joint first authors pact form, retaining the statistical properties of the source
without exposing the characteristics of any individual data Our main contributions are:
point in the original data set.
• The first published upscaling of a database of high res-
Previous methods, including traditional methods [7, 12,
olution (4K) physically based face model assets.
27, 34, 16, 9], or deep learning [42, 38] to represent 3D face
shapes; lack high resolution (sub-millimeter, < 1mm) geo- • A cascading generative face model, enabling control
metric detail, use limited representations of facial anatomy, of identity and expressions, as well as physically based
or forgo the physically based material properties required surface materials modeled in a low dimensional feature
by modern visual effects (VFX) production pipelines. Phys- space.
ically based material intrinsics have proven difficult to es-
timate through the optimization of unconstrained image • The first morphable face model built for full 3D real
data due to ambiguities and local minima in analisys-by- time and offline rendering applications, with more rel-
synthesis problems, while highly constrained data capture evant anatomical face parts than previously seen.
remains percise but expensive [6]. Although variations
occur due to different applications, most face representa- 2. Related Work
tions used in VFX employ a set of texture maps of at least Facial Capture Systems Physical object scanning de-
4096 × 4096 (4K) pixels resolution. At a minimum, this vices span a wide range of categories; from single RGB
set encorporates diffuse albedo, specular intensity, and dis- cameras [14, 39], to active [3, 17], and passive [4]
placement (or surface normals). light stereo capture setups, and depth sensors based on
Our goal is to build a physically-based, high-resolution time-of-flight or stereo re-projection. Multi-view stereo-
generative face model to begin bridging these parallel, but photogrammetry (MVS) [4] is the most readily available
in some ways divergent, visualization fields; aligning the method for 3D face capturing. However, due to its many
efforts of vision and graphics researchers. Building such advantages over other methods (capture speed, physically-
a model requires high-resolution facial geometry, material based material capturing, resolution), polarized spherical
capturing and automatic registration of multiple assets. The gradient illumination scanning [17] remains state-of-the-art
handling of said data has traditionally required extensive for high-resolution facial scanning. A mesoscopic geome-
manual work, thus scaling such a database is non-trivial. try reconstruction is bootstrapped using an MVS prior, uti-
For the model to be light weight these data need to be com- lizing omni-directional illumination, and progressively fi-
pressed into a compact form that enables controlled recon- nalized using a process known as photometric stereo [17].
struction based on novel input. Traditional methods such as The algorithm promotes the physical reflectance properties
PCA [7] and bi-linear models [12] − which are limited by of dielectric materials such as skin; specifically the separa-
memory size, computing power, and smoothing due to in- ble nature of specular and subsurface light reflections [29].
herent linearity − are not suitable for high-resolution data. This enables accurate estimation of diffuse albedo and spec-
By leveraging state-of-the-art physically-based facial ular intensity as well as pore-level detailed geometry.
scanning [17, 25], in a Light Stage setting, we enable acqui-
sition of diffuse albedo and specular intensity texture maps 3D Morphable Face Models The first published work
in addition to 4K displacement. All scans are registered on morphable face models by Blanz and Vetter [7] repre-
using an automated pipeline that considers pose, geome- sented faces as dense surface geometry and texture, and
try, anatomical morphometrics, and dense correspondence modeled both variations as separate PCA models learned
of 26 expressions per subject. A shared 2D UV param- from around 200 subject scans. To allow intuitive con-
eterization data format [15, 43, 38], enables training of a trol; attributes, such as gender and fullness of faces, were
non-linear 3DMM, while the head, eyes, and teeth are rep- mapped to components of the PCA parameter space. This
resented using a linear PCA model. Hence, we propose a model, known as the Basel Face Model [33] was released
hybrid approach to enable a wide set of head geometry as- for use in the research community, and was later extended to
sets as well as avoiding the assumption of linearity in face a more diverse linear face model learnt from around 10,000
deformations. scans [9, 8].
Our model fully disentangles identity from expressions, To incorporate facial expressions, Vlasic et al. [45] pro-
and provides manipulation using a pair of low dimensional posed a multi-linear model to jointly estimate the varia-
feature vectors. To generate coupled geometry and albedo, tions in identity, viseme, and expression, and Cao et al. [12]
we designed a joint discriminator to ensure consistency, built a comprehensive bi-linear model (identity and expres-
along with two separate discriminators to maintain their sion) covering 20 different expressions from 150 subjects
individual quality. Inference and up-scaling of before- learned from RGBD data. Both of these models adopt a
mentioned skin intrinsics enable recovery of 4K resolution tensor-based method under the assumption that facial ex-
texture maps. pressions can be modeled using a small number of discrete
(a) (b) (c) (d) (e)
LS 4k × 4k 3.9M 4k × 4k 79 26
TG 8k × 8k 3.5M N/A 99 20

Table 1: Resolution and extent of the datasets. (a). Albedo resolu-


tion. (b). Geometry resolution. (c). Specular intensity resolution.
(d) # of subjects. (f). # of expressions per subject.

Figure 1: Capture system and camera setup. Left: Light Stage


capturing system. Right: camera layout.
as the coarse [44], medium [36] or even mesoscopic [21]
poses, corresponded between subjects. More recently, Li scale facial geometry inferred directly from images. Beside
et al. [27] released the FLAME model, which incorporates geometry, Yamaguchi et al. [47] presented a comprehen-
both pose-dependent corrective blendshapes, and additional sive method to infer facial reflectance maps (diffuse albedo,
global identity and expression blendshapes learnt from a specular intensity, and medium- and high-frequency dis-
large number of 4D scans. placement) based on single image inputs. More recently,
To enable adaptive, high level, semantic control over face Nagano et al. [31] proposed a framework for synthesiz-
deformations, various locality-based face models have been ing arbitrary expressions both in image space and UV tex-
proposed. Neumann et al. [32] extract sparse and spatially ture space, from a single portrait image. Although these
localized deformation modes, and Brunton et al. [10] use methods can synthesize facial geometry or/and texture maps
a large number of localized multilinear wavelet modes. As from a given image, they don’t provide explicit parametric
a framework for anatomically accurate local face deforma- controls of the generated result.
tions, the Facial Action Coding System (FACS) by Ekman
[13] is widely adopted. It decomposes facial movements
into basic action units attributed to the full range of motion 3. Database
of all facial muscles.
Morphable face models have been widely used for appli- 3.1. Data Capturing and Processing
cations like face fitting [7], expression manipulation [12],
real-time tracking [41], as well as in products like Apple’s Data Capturing Our Light Stage scan system employs
ARKit. However, their use cases are often limited by the photometric stereo [17] in combination with monochrome
resolution of the source data and restrictions of linear mod- color reconstruction using polarization promotion [25] to
els causing smoothing in middle and high frequency geom- allow for pore level accuracy in both the geometry re-
etry details (e.g. wrinkles, and pores). Moreover, to the best construction and the reflectance maps. The camera setup
of our knowledge, all existing morphable face models gen- (Fig.1) was designed for rapid, database scale, acquisition
erate texture and geometry separately, without considering by the use of Ximea machine vision cameras which en-
the correlation between them. Given the specific and var- able faster streaming and wider depth of field than tra-
ied ways in which age, gender, and ethnicity are manifested ditional DSLRs [25]. The total set of 25 cameras con-
within the spectrum of human life, ignoring such correla- sists of eight 12MP monochrome cameras, eight 12MP
tion will cause artifacts; e.g. pairing an African-influenced color cameras, and nine 4MP monochrome cameras. The
albedo to an Asian-influenced geometry. 12MP monochrome cameras allow for pore level geome-
try, albedo, and specular reflectance reconstruction, while
Image-based Detail Inference To augment the quality of the additional cameras aid in stereo base mesh-prior recon-
existing 3DMMs, many works have been proposed to infer struction.
the fine-level details from image data. Skin detail can be To capture consistent data across multiple subjects with
synthesized using data-driven texture synthesis [20] or sta- maximized expressiveness, we devised a FACS set [13]
tistical skin detail models [18]. Cao et al. [11] used a prob- which combines 40 action units to a condensed set of 26
ability map to locally regress the medium-scale geometry expressions. In total, 79 subjects, 34 female, and 45 male,
details, where a regressor was trained from captured patch ranging from age 18 to 67, were scanned performing the 26
pairs of high-resolution geometry and appearance. Saito et expressions. To increase diversity, we combined the data
al. [35] presented a texture inference technique using a deep set with a selection of 99 Triplegangers [2] full head scans;
neural network-based feature correlation analysis. each with 20 expressions. Resolution and extent of the two
GAN-based Image-to-Image frameworks [22] have data sets are shown in Table 1. Fig. 2 shows the age and
proven to be powerful for high-quality detail synthesis, such ethnicity (multiple choice) distributions of the source data.
1 mm
80 60
60
40
40
20 20

0 0
15 25 35 45 55 65 75 Asian Indian Black White Hispanic Middle
Eastern
LightStage Age interval (yrs) LightStage Ethnicity
Triplegangers Triplegangers

(a) Age distribution (b) Ethnicity distribution


4K 256x256 0 mm
Figure 2: Distribution of age (a) and ethnicity (b) in the data sets.
Figure 4: Comparison of base mesh geometry resolutions. Left:
Base geometry reconstructed in 4K resolution. Middle: Base ge-
(e)
(c) ometry reconstructed in 256 × 256 resolution. Right: Error map
(f) (j) showing the Hausdorff distance in the range (0mm, 1mm), with
(g) a mean error of 0.068mm.
(h)
(a) (d) (k) spread out in texture space, would
√ require a bitmap of res-
(i)
olution greater or equal to d 2 × me2 = 155 × 155, ac-
(b)
cording to Nyquist’s resampling theorem. As shown in
Linear Non-linear (DNN) Non-linear (Laplacian) (l) Fig. 4, the proposed resolution is enough to recover middle-
Figure 3: Our generic face model consists of multiple geometries frequency detail. This relatively low resolution base geom-
constrained by different types of deformation. In addition to face etry representation enables great simplification in training
(a), head, and neck (b), our model represents teeth (c), gums (d), data load.
eyeballs (e), eye blending (f), lacrimal fluid (g), eye occlusion
(h), and eyelashes (i). Texture maps provide high resolution (4K)
albedo (j), specularity (k), and geometry through displacement (l). Data Augmentation Since the number of subjects is lim-
ited to 178 individuals, we apply two strategies to augment
Processing Pipeline. Starting from the multi-view im- the data for identity training: 1) For each source albedo,
agery, a neutral scan base mesh is reconstructed using MVS. we randomly sample a target albedo within the same eth-
Then a linear PCA model in our topology (See Fig.3) based nicity and gender in the data set using [49] to transfer skin
on a combination and extrapolation of two existing mod- tones of target albedos to source albedos (these samples are
els (Basel [33] and Face Warehouse [12]) is used to fit the restricted to datapoints of the same ethnicity), followed by
mesh. Next, Laplacian deformation is applied to deform an image enhancement [19] to improve the overall quality
the face area to further minimize the surface-to-surface er- and remove artifacts. 2). For each neutral geometry, we
ror. Cases of inaccurate fitting were manually modeled and add a very small expression offset using FaceWarehouse ex-
fitted to retain the fitting accuracy of the eyeballs, mouth pression components with a small random weights(< ±0.5
sockets and skull shapes. The resulting set of neutral scans std) to loosen the constraints of “neutral”. To augment the
were immediately added to the PCA basis for registering expressions, we add random expression offsets to generate
new scans. We fit expressions using generic blendshapes fully controlled expressions.
and non-rigid ICP [26]. Additionally, to retain texture space
and surface correspondence, image space optical flow from 4. Generative Model
neutral to expression scan is added from 13 different vir-
An overview of our system is illustrated in Fig. 5. Given
tual camera views as additional dense constraint in the final
a sampled latent code Zid ∼ N (µid , σid ), our Identity net-
Laplacian deformation of the face surface.
work generates a consistent albedo and geometry pair of
3.2. Training Data Preparation neutral expression. We train an Expression network to gen-
erate the expression offset that can be added to the neu-
Data format. The full set of the generic model consists tral geometry. We use random blendshape weights Zexp ∼
of a hybrid geometry and texture maps (albedo, specular N (µexp , σexp ) as the expression network’s input to enable
intensity, and displacement) encoded in 4K resolution, as manipulation of target semantic expressions. We upscale
illustrated in Fig. 3. To enable joint learning of the cor- the albedo and geometry maps to 1K, and feed them into
relation between geometry and albedo, 3D vertex positions a transfer network [46] to synthesize the corresponding 1K
are rasterized to a three channel HDR bitmap of 256 × 256 specular and displacement maps. Finally, all the maps ex-
pixels resolution. The face area (pink in Fig. 3) used to cept for the middle frequency geometry map are upscaled
learn the geometry distribution in our non-linear generative to 4K using Super-resolution [24], as we observed that
model consists of m = 11892 vertices, which, if evenly 256 × 256 pixels are sufficient to represent the details of
256x256 1K 4K

Identity Super
Network Resolution

~ ( , ) 256x256 1K
Identity latent Texture
Inference Assemble
Super
and
Resolution
Render

256x256 1K

Expression
Network Expression
geometry
~ ( , ) Expression Inferred maps Up-sampled maps Rendered image of combined assets
Expression latent offset

Figure 5: Overview of generative pipeline. Latent vectors for identity and expression serve as input for generating the final face model.

ℒexp =∥ Zexp − Zexp




GT Dalbedo Real/
Fake?

Generated Gexp Rexp


Gid
albedo
Djoint Real/
Fake?
GT
~ ( , ) Generated offset ′
~ ( , ) Expression latent
Identity latent
Dgeometry Real/
Dexp Real/
Fake? Fake?
GT
Generated
geometry Ground truth offset

Figure 6: Identity generative network. The identity generator Figure 7: Expression generative network. The expression genera-
Gid produces albedo and geometry which get checked against tor Gexp generates offsets which get checked against ground truth
ground truth (GT) data by the discriminators, Dalbedo , Djoint , offsets by the discriminator Dexp . The regressor Rexp produces
0
and Dgeometry during training. an estimate of the latent code Zexp so that the L1 loss Lexp can be
modeled.
the base geometry (Section 3.2). The details of each com-
ponent are elaborated on in Section 4.1, 4.2, and 4.3. which also makes the learning of expressions independent
from identity. Similar to the Identity network, the expres-
4.1. Identity Network
sion network adopts Style-GAN as the base structure. To al-
The goal of our Identity network is to model the cross low for intuitive control over expressions, we use the blend-
correlation between geometry and albedo to generate con- shape weights, which correspond to the strength of 25 or-
sistent, diverse and biologically accurate identities. The net- thogonal facial activation units, as network input. We in-
work is built upon the Style-GAN architecture [23], that can troduce a pre-trained expression regression network Rexp
produce high-quality, style-controllable sample images. to predict the expression weights from the generated image,
To achieve consistency, we designed 3 discriminators and force this prediction to be similar to the input latent
as shown in Fig.6, including individual discriminators for code Zexp . We then force the generator to understand the in-
albedo (Dalbedo ) and geometry (Dgeometry ), to ensure the put latent code Zexp under the perspective of the pre-trained
quality and sharpness of the generated maps, and an addi- expression regression network. As a result, each dimension
tional joint discriminator (Djoint ) to learn their correlated of the latent code Zexp will control the corresponding ex-
distribution. Djoint is formulated as follows: pression defined in the original blendshape set. The loss we
  introduce here is:
Ladv = min max Ex∼pdata (x) log Djoint (A) +
Gid Djoint
  (1) Lexp =k Zexp − Zexp k
0
(2)
Ez∼pz (z) log (1 − Djoint (Gid (z))) .
This loss, Lexp , will be back propagated during training to
where pdata (x) and pz (z) represent the distributions of real enforce the orthogonality of each blending unit. We mini-
paired albedo and geometry x and noise variables z in the mize the following losses to train the network:
domain of A respectively.
4.2. Expression Network L = Lexp exp
l2 + β1 Ladv + β2 Lexp (3)

To simplify the learning of a wide range of diverse where Lexp


l2 is the L2 reconstruction loss of the offset map
expressions, we represent them using vector offset maps, and Lexp
adv is the discriminator loss.
4.3. Inference and Super-resolution
Similar to [47]; upon obtaining albedo and geometry
maps (256 × 256), we use them to infer specular and dis-
placement maps in 1K resolution. In contrast to [47], us-
ing only albedo as input, we introduce the geometry map to
form stronger constraints. For displacement, we adopted the
method of [47, 21] to separate displacement in to individ-
ual high-frequency and low-frequency components, which
makes the problem more tractable. Before feeding the
two inputs into the inference network [46], we up-sample
the albedo to 1K using a super-resolution network simi-
lar to [24]. The geometry map is super-sampled using bi- Figure 8: Non-linear identity interpolation between generated sub-
linear interpolation. The maps are further up-scaled from jects. Age (top) and gender (bottom) are interpolated from left to
1K to 4K using the same super-resolution network struc- right.
ture. Our method can be regarded as a two step cascading
up-sampling strategy (256 to 1K, and 1K to 4K). This Jaw open [std from neutral]
makes the training faster, and enables higher resolution in
the final results.

5. Implementation Details Right eye blink [std from neutral]


-5.0
Our framework is implemented using Pytorch and all our
networks are trained using two NVIDIA Quadro GV100s.
We follow the basic training schedule of Style-GAN [23]
with several modifications applied to the Expression net- 1.0
work, like by-passing the progressive training strategy as
expression offsets are only distinguishable on relatively
high resolution maps. We also remove the noise injec-
tion layer, due to the input latent code Zexp which enables 7.0
full control of the generated results. The regression mod- 0.0 2.5 4.0 5.5 7.0
ule (Rexp -block in Fig.7) has the same structure as the dis-
criminator Dexp , except for the number of channels in the Figure 9: Non linear expression interpolation using generative ex-
last layer, as it serves as a discriminator during training. pression network. Combinations of two example shapes are dis-
The regression module is initially trained using synthetic played in a grid where the number of standard deviations from the
generic neutral model define the extent of an expression shape.
unit expression data generated with neutral expression and
F aceW arehouse expression components, and then fine-
tuned on scanned expression data. During training, Rexp , is
6.2. Qualitative Evaluation
fixed without updating parameters. The Expression network
is trained with a constant batch size of 128 on 256x256- We show identity interpolation in Fig.8. The interpola-
pixel images for 40 hours. The Identity network is trained tion in latent space reflects both albedo and geometry. In
by progressively reducing the batch size from 1024 to 128 contrast to linear blending, our interpolation generates sub-
on growing image sizes ranging from 8x8 to 256x256 pix- jects belonging to a natural statistical distribution.
els, for 80 hours.
In Fig.9, we show the generation and interpolation of
6. Experiments And Evaluations our non-linear expression model. We pick two orthogonal
blendshapes for each axis and gradually change the input
6.1. Results weights. Smooth interpolation in vector space will lead to a
smooth interpolation in model space.
In Fig.10, we show the quality of our generated model
rendered using Arnold. The direct output of our genera- We show nearest neighbors for generated models in the
tive model provides all the assets necessary for physically- training set in Fig.11. These are found based on point-wise
based rendering in software such as Maya, Unreal Engine, Euclidean distance in geometry. Albedos are compared to
or Unity 3D. We also show the effect of each generated prove our ability to generate new models that are not merely
component. recreations of the training set.
(a) (b) (c) (d) (e) (f)
Figure 10: Rendered images of generated random samples. Column (a), (b), and (c) show images rendered under novel image-based HDRI
lighting [48]. Column (c), (d), and (e), show geometry with albedo, specular intensity, and displacement added one at the time.

6.3. Quantitative Evaluation


We evaluate the effectiveness of our identity network’s
joint generation in Table 2 by computing Frechet Inception
Distances (FID) and Inception-Scores (IS) on rendered im-
ages of three categories: randomly paired albedo and ge-
ometry, paired albedo and geometry generated using our
model, and ground truth pairs. Based on these results, we
conclude that our model generates more plausible faces,
Figure 11: Nearest neighbors for generated models in training set. similar to those using ground truth data pairs, than random
Top row: albedo from generated models. Bottom row: albedo of pairing.
geometrically nearest neighbor in training set. We also evaluate our identity network’s generalization to
unseen faces by fitting 48 faces from [1]. The average Haus-
Generation Method IS↑ FID↓ dorff distance is 2.8mm, which proves that our model’s ca-
independent 2.22 23.61 pacity is not limited by the training set.
joint 2.26 21.72 In addition, to evaluate the non-linearity of our expres-
groud truth 2.35 - sion network in comparison to the linear expression model
of FaceWarehouse [12], we first fit all the Light Stage scans
using FaceWarehouse, and get the 25 fitting weights, and
Table 2: Evaluation on our Identity generation. Both IS and FID
expression recoveries, for each scan. We then recover the
are calculated on images rendered with independently/jointly gen-
erated albedo and geometry.
same expressions by feeding the weights to our expres-
sion network. We evaluate the reconstruction loss with
mean-square error (MSE) for both FaceWarehouse’s and
40% 40%
30% 30%
20% 20%
10% 10%
0% 0%
15 25 35 45 55 65 75 15 25 35 45 55 65 75
Female Age interval (yrs) Female Age interval (yrs)
Male Male

(a) Training data (b) Generated data


Figure 12: The age distribution of the training data (a) VS. ran-
domly generated samples (b).
Figure 14: Low-quality data domain transfer. Top row: Models
5 mm
with low resolution geometry and albedo. Bottom row: Enhance-
ment result using our model.

7. Conclusion and Limitations


Conclusion We have introduced the first published use of
a high-fidelity face database, with physically-based mare-
rial attributes, in generative face modeling. Our model can
generate novel subjects and expressions in a controllable
Basel FaceWarehouse FLAME Ours Ground truth
0 mm manner. We have shown that our generative model performs
well on applications such as mesh registration and low res-
Figure 13: Comparison of 3D scan fitting with Basel [7], Face- olution data enhancement. We hope that this work will ben-
wareHouse [12], and FLAME [27]. Error maps are computed us-
efit many analysis-by-synthesis research efforts through the
ing Hausdorff distance between each fitted model and ground truth
scans.
provision of higher quality in face image rendering.

our model’s reconstructions. On average, our method’s Limitations and Future work In our model, expression
MSE is 1.2mm while FaceWarehouse’s is 2.4mm. This and identity are modeled separately without considering
shows that for expression fitting, our non-linear model nu- their correlation. Thus the reconstructed expression off-
merically outperforms a linear model of the same dimen- set will not include middle-frequency geometry of an in-
sionality. dividual’s expression, as different subjects will have unique
To demonstrate our generative identity model’s coverage representations of the same action unit. Our future work
of the training data, we show the gender, and age distribu- will include modeling of this correlation. Since our expres-
tions of the original training data and 5000 randomly gener- sion generation model requires neural network inference
ated samples in Fig.12. The generated distributions are well and re-sampling of 3D geometry it is not currently as user
aligned with the source. friendly as blendshape modeling. Its ability to re-target pre-
recorded animation sequences will have to be tested further
6.4. Applications
to be conclusive. One issue of our identity model arises in
To test the extent of our identity model’s parameter applications that require fitting to 2D imagery, which ne-
space, we apply it to scanned mesh registration by reversing cessitates an additional differentiable rendering component.
the GAN to fit the latent code of a target image [28]. As our A potential problem is fitting lighting in conjunction with
model requires a 2D parameterized geometry input, we first shape as complex material models make the problem less
use our linear model to align the scans using landmarks, and tractable. A possible solution could be an image-based re-
then parameterize it to UV space after Laplacian morphing lighting method [40, 30] applying a neural network to con-
of the surface. We compare our fitting results with widely vert the rendering process to an image manipulation prob-
used (linear) morphable face models in Fig.13. This evalua- lem. The model will be continuously updated with new fea-
tion does not prove the ability to register unconstrained data tures such as variable eye textures and hair as well as more
but shows that our model is able to reconstruct novel faces anatomically relevant components such as skull, jaw, and
by the virtue of it’s non-linearity, to a degree unobtainable neck joints by combining data sources through collabora-
by linear models. tive efforts. To encourage democratization and wide use
Another application of our model is transferring low- cases we will explore encryption techniques such as fed-
quality scans into the domain of our model by fitting using erated learning, homomorphic encryption, and zero knowl-
both MSE loss and discriminator loss. In Fig.14, we show edge proofs which have the effect of increasing subjects’
examples of data enhancement of low resolution scans. anonymity.
References try from monocular video. ACM Transactions on Graphics
(TOG), 32(6), Nov. 2013. 2
[1] 3d scan store: Male and female 3d head
[15] Baris Gecer, Alexander Lattas, Stylianos Ploumpis, Jiankang
model 48 x bundle. [Link]
Deng, Athanasios Papaioannou, Stylianos Moschoglou, and
[Link]/3d-head-models/
Stefanos Zafeiriou. Synthesizing coupled 3d face modali-
female-retopologised-3d-head-models/
ties by trunk-branch generative adversarial networks. arXiv
male-female-3d-head-model-48xbundle.
preprint arXiv:1909.02215, 2019. 2
Online; Accessed: 2019-11-22. 7
[16] Thomas Gerig, Andreas Morel-Forster, Clemens Blumer,
[2] Triplegangers. [Link] On-
Bernhard Egger, Marcel Luthi, Sandro Schönborn, and
line; Accessed: 2019-11-22. 3
Thomas Vetter. Morphable face models-an open framework.
[3] Oleg Alexander, Mike Rogers, William Lambeth, Matt Chi-
In IEEE International Conference on Automatic Face & Ges-
ang, and Paul Debevec. The digital emily project: Photo-
ture Recognition (FG 2018). IEEE, 2018. 2
real facial modeling and animation. In ACM SIGGRAPH
Courses, 2009. 2 [17] Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay
Busch, Xueming Yu, and Paul E. Debevec. Multiview face
[4] Thabo Beeler, Bernd Bickel, Paul A. Beardsley, Bob Sumner,
capture using polarized spherical gradient illumination. ACM
and Markus H. Gross. High-quality single-shot capture of
Transactions on Graphics (TOG), 30:129, 2011. 1, 2, 3
facial geometry. In ACM Transactions on Graphics (TOG),
2010. 2 [18] Aleksey Golovinskiy, Wojciech Matusik, Hanspeter Pfister,
[5] Thabo Beeler, Fabian Hahn, Derek Bradley, Bernd Bickel, Szymon Rusinkiewicz, and Thomas A. Funkhouser. A sta-
Paul A. Beardsley, Craig Gotsman, Robert W. Sumner, and tistical model for synthesis of detailed facial geometry. ACM
Markus Groß. High-quality passive facial performance cap- Transactions on Graphics (TOG), 25:1025–1034, 2006. 3
ture using anchor frames. In ACM Transactions on Graphics [19] Yoav HaCohen, Eli Shechtman, Dan B Goldman, and Dani
(TOG), 2011. 1 Lischinski. Non-rigid dense correspondence with applica-
[6] Ayush Tewari Stefanie Wuhrer Michael Zollhoefer tions for image enhancement. ACM transactions on graphics
Thabo Beeler Florian Bernard Timo Bolkart Adam (TOG), 30(4):70, 2011. 4
Kortylewski Sami Romdhani Christian Theobalt Volker [20] Antonio Haro, Irfan A. Essa, and Brian K. Guenter. Real-
Blanz Thomas Vetter Bernhard Egger, William A. P. Smith. time photo-realistic physically based rendering of fine scale
3d morphable face models–past, present and future. arXiv human skin structure. In Rendering Techniques, 2001. 3
preprint arXiv:1909.01815, 2019. 1, 2 [21] Loc Huynh, Weikai Chen, Shunsuke Saito, Jun Xing, Koki
[7] Volker Blanz and Thomas Vetter. A morphable model for Nagano, Andrew Jones, Paul E. Debevec, and Hao Li. Meso-
the synthesis of 3d faces. In ACM Transactions on Graphics scopic facial geometry inference using deep neural networks.
(TOG), SIGGRAPH ’99, 1999. 1, 2, 3, 8 Conference on Computer Vision and Pattern Recognition
[8] James Booth, Anastasios Roussos, Allan Ponniah, David (CVPR), 2018. 3, 6
Dunaway, and Stefanos Zafeiriou. Large scale 3d morphable [22] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A.
models. International Journal of Computer Vision, 126:233– Efros. Image-to-image translation with conditional adver-
254, 2017. 2 sarial networks. 2017 IEEE Conference on Computer Vision
[9] James Booth, Anastasios Roussos, Stefanos Zafeiriou, Allan and Pattern Recognition (CVPR), pages 5967–5976, 2016. 3
Ponniah, and David Dunaway. A 3d morphable model learnt [23] Tero Karras, Samuli Laine, and Timo Aila. A style-based
from 10,000 faces. In Proceedings of the IEEE Conference generator architecture for generative adversarial networks.
on Computer Vision and Pattern Recognition (CVPR), pages In Proceedings of the IEEE Conference on Computer Vision
5543–5552, 2016. 2 and Pattern Recognition, pages 4401–4410, 2019. 5, 6
[10] Alan Brunton, Timo Bolkart, and Stefanie Wuhrer. Multilin- [24] Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero,
ear wavelets: A statistical shape space for human faces. In Andrew Cunningham, Alejandro Acosta, Andrew Aitken,
Proceedings of the European Conference on Computer Vi- Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo-
sion (ECCV), 2014. 3 realistic single image super-resolution using a generative ad-
[11] Chen Cao, Derek Bradley, Kun Zhou, and Thabo Beeler. versarial network. In Proceedings of the IEEE conference on
Real-time high-fidelity facial performance capture. ACM computer vision and pattern recognition, pages 4681–4690,
Transactions on Graphics (TOG), 34(4), July 2015. 3 2017. 4, 6
[12] Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Zhou [25] Chloe LeGendre, Kalle Bladin, Bipin Kishore, Xinglei Ren,
Kun. Facewarehouse: A 3d facial expression database for Xueming Yu, and Paul Debevec. Efficient multispectral fa-
visual computing. IEEE Transactions on Visualization and cial capture with monochrome cameras. In Color and Imag-
Computer Graphics, 2014. 1, 2, 3, 4, 7, 8 ing Conference, volume 2018, pages 187–202, 2018. 2, 3
[13] Paul Ekman and Wallace V. Friesen. Facial action coding [26] Hao Li, Robert W. Sumner, and Mark Pauly. Global cor-
system: a technique for the measurement of facial move- respondence optimization for non-rigid registration of depth
ment. In Consulting Psychologists Press, 1978. 3 scans. In Proceedings of the Symposium on Geometry Pro-
[14] Pablo Garrido, Levi Valgaert, Chenglei Wu, and Christian cessing, SGP ’08, pages 1421–1430, Aire-la-Ville, Switzer-
Theobalt. Reconstructing detailed dynamic face geome- land, Switzerland, 2008. Eurographics Association. 4
[27] Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and using monocular videos. ACM Transactions on Graphics
Javier Romero. Learning a model of facial shape and expres- (TOG), 33:222:1–222:13, 2014. 2
sion from 4d scans. ACM Transactions on Graphics (TOG), [40] Tiancheng Sun, Jonathan T Barron, Yun-Ta Tsai, Zexiang
36(6):194, 2017. 2, 3, 8 Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay
[28] Zachary C Lipton and Subarna Tripathi. Precise recovery of Busch, Paul Debevec, and Ravi Ramamoorthi. Single image
latent vectors from generative adversarial networks. arXiv portrait relighting. ACM Transactions on Graphics (TOG),
preprint arXiv:1702.04782, 2017. 8 38(4):79, 2019. 8
[29] Wan-Chun Ma, Tim Hawkins, Pieter Peers, Charles-Félix [41] Justus Thies, Michael Zollhöfer, Marc Stamminger, Chris-
Chabert, Malte Weiss, and Paul E. Debevec. Rapid acqui- tian Theobalt, and Matthias Nießner. Face2face: real-time
sition of specular and diffuse normal maps from polarized face capture and reenactment of rgb videos. IEEE Confer-
spherical gradient illumination. In Rendering Techniques, ence on Computer Vision and Pattern Recognition (CVPR),
2007. 2 2016. 1, 3
[30] Abhimitra Meka, Christian Haene, Rohit Pandey, Michael [42] Luan Tran, Feng Liu, and Xiaoming Liu. Towards high-
Zollhoefer, Sean Fanello, Graham Fyffe, Adarsh Kowdle, fidelity nonlinear 3d face morphable model. In Proceed-
Xueming Yu, Jay Busch, Jason Dourgarian, Peter Denny, ings of the IEEE Conference on Computer Vision and Pattern
Sofien Bouaziz, Peter Lincoln, Matt Whalen, Geoff Harvey, Recognition (CVPR), pages 1126–1135, 2019. 2
Jonathan Taylor, Shahram Izadi, Andrea Tagliasacchi, Paul [43] Luan Tran and Xiaoming Liu. On learning 3d face mor-
Debevec, Christian Theobalt, Julien Valentin, and Christoph phable model from in-the-wild images. IEEE transactions
Rhemann. Deep reflectance fields - high-quality facial re- on pattern analysis and machine intelligence, 2019. 2
flectance field inference from color gradient illumination. [44] George Trigeorgis, Patrick Snape, Iasonas Kokkinos, and
volume 38, July 2019. 8 Stefanos Zafeiriou. Face normals ”in-the-wild” using fully
[31] Koki Nagano, Jaewoo Seo, Jun Xing, Lingyu Wei, Zimo convolutional networks. 2017 IEEE Conference on Com-
Li, Shunsuke Saito, Aviral Agarwal, Jens Fursund, and Hao puter Vision and Pattern Recognition (CVPR), pages 340–
Li. pagan: real-time avatars using dynamic textures. ACM 349, 2017. 3
Transactions on Graphics (TOG), 37:258:1–258:12, 2018. 3 [45] Daniel Vlasic, Matthew Brand, Hanspeter Pfister, and Jovan
[32] Thomas Neumann, Kiran Varanasi, Stephan Wenger, Markus Popovic. Face transfer with multilinear models. In ACM
Wacker, Marcus A. Magnor, and Christian Theobalt. Sparse Transactions on Graphics (TOG), 2005. 2
localized deformation components. ACM Transactions on [46] Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao,
Graphics (TOG), 32:179:1–179:10, 2013. 3 Jan Kautz, and Bryan Catanzaro. High-resolution image syn-
[33] Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami thesis and semantic manipulation with conditional gans. In
Romdhani, and Thomas Vetter. A 3d face model for pose and Proceedings of the IEEE Conference on Computer Vision
illumination invariant face recognition. IEEE International and Pattern Recognition, 2018. 4, 6
Conference on Advanced Video and Signal Based Surveil- [47] Shugo Yamaguchi, Shunsuke Saito, Koki Nagano, Yajie
lance, pages 296–301, 2009. 2, 4 Zhao, Weikai Chen, Kyle Olszewski, Shigeo Morishima, and
[34] Stylianos Ploumpis, Haoyang Wang, Nick Pears, Hao Li. High-fidelity facial reflectance and geometry infer-
William AP Smith, and Stefanos Zafeiriou. Combin- ence from an unconstrained image. ACM Transactions on
ing 3d morphable models: A large scale face-and-head Graphics (TOG), 37:162:1–162:14, 2018. 3, 6
model. In Proceedings of the IEEE Conference on Computer [48] G. Zaal. HDRI Haven. [Link]
Vision and Pattern Recognition (CVPR), 2019. 2 hdris/. Online; Accessed: 2019-11-22. 7
[35] Shunsuke Saito, Lingyu Wei, Liwen Hu, Koki Nagano, and [49] Yajie Zhao, Qingguo Xu, Weikai Chen, Chao Du, Jun Xing,
Hao Li. Photorealistic facial texture inference using deep Xinyu Huang, and Ruigang Yang. Mask-off: Synthesizing
neural networks. 2017 IEEE Conference on Computer Vision face images in the presence of head-mounted displays. In
and Pattern Recognition (CVPR), pages 2326–2335, 2016. 3 2019 IEEE Conference on Virtual Reality and 3D User In-
[36] Matan Sela, Elad Richardson, and Ron Kimmel. Unre- terfaces (VR), pages 267–276. IEEE, 2019. 4
stricted facial geometry reconstruction using image-to-image
translation. 2017 IEEE International Conference on Com-
puter Vision (ICCV), pages 1585–1594, 2017. 3
[37] Yeongho Seol, Wan-Chun Ma, and J. P. Lewis. Creating
an actor-specific facial rig from performance capture. In
Proceedings of the 2016 Symposium on Digital Production,
DigiPro ’16, 2016. 1
[38] Gil Shamai, Ron Slossberg, and Ron Kimmel. Syn-
thesizing facial photometries and corresponding geome-
tries using generative adversarial networks. arXiv preprint
arXiv:1901.06551, 2019. 2
[39] Fuhao Shi, Hsiang-Tao Wu, Xin Tong, and Jinxiang Chai.
Automatic acquisition of high-fidelity facial performances

You might also like