A 3D Face Model for Pose and Illumination Invariant Face
Recognition
Pascal Paysan Reinhard Knothe Brian Amberg
[Link]@[Link] [Link]@[Link] [Link]@[Link]
Sami Romdhani Thomas Vetter
[Link]@[Link] [Link]@[Link]
Abstract most important face recognition [8, 18, 11], but also
face image analysis [7] (estimating the 3D shape from
Generative 3D face models are a powerful tool in a single photograph), expression transfer between in-
computer vision. They provide pose and illumination dividuals [6, 17], animation of faces and whole bod-
invariance by modeling the space of 3D faces and the ies [6, 1], and stimuli generation for psychological ex-
imaging process. The power of these models comes at periments [14] to name a few.
the cost of an expensive and tedious construction pro- A 3DMM consists of a parameterized generative 3D
cess, which has led the community to focus on more eas- shape, and a parameterized albedo model together with
ily constructed but less powerful models. With this pa- an associated probability density on the model coef-
per we publish a generative 3D shape and texture model, ficients. A set of shape and albedo coefficients de-
the Basel Face Model (BFM), and demonstrate its ap- scribes a face. Together with projection and illumina-
plication to several face recognition task. We improve tion parameters a rendering of the face can be gener-
on previous models by offering higher shape and texture ated. Given a face image one can also solve the inverse
accuracy due to a better scanning device and less cor- problem of finding the coefficients which most likely
respondence artifacts due to an improved registration generated the image. Identification and manipulation
algorithm. tasks in coefficient space are trivial, because the gen-
The same 3D face model can be fit to 2D or 3D im- erating factors (light, pose, camera, and identity) have
ages acquired under different situations and with dif- been separated. Solving this inverse problem is termed
ferent sensors using an analysis by synthesis method. “model fitting”, and was introduced for faces in [7] and
The resulting model parameters separate pose, lighting, subsequently refined in [18]. A similar method has also
imaging and identity parameters, which facilitates in- been applied to stereo data [3] and 3D scans [2].
variant face recognition across sensors and data sets by However, the widespread use of 3DMMs has been
comparing only the identity parameters. We hope that held back by their difficult construction process, which
the availability of this registered face model will spur requires a precise and fast 3D scanner, the scanning of
research in generative models. Together with the model several hundreds of individuals and the computation
we publish a set of detailed recognition and reconstruc- of dense correspondence between the scans. Numer-
tion results on standard databases to allow complete ous face recognition articles acknowledge the fact that
algorithm comparisons. 3DMM based face image analysis constitutes the state
of the art, but note that the main obstacle resides in the
complications of their construction (e.g. [22, 13, 12, 5]).
1. Introduction For example, quoting Zhou and Chellappa [23]: “Its
Automatic face recognition from a single image is only weakness is the requirement of the 3D models”.
still difficult for non-frontal views and complex illumi- Hence, there is a demand from the face image analysis
nation conditions. To achieve pose and light invari- community for a publicly available 3D Morphable Face
ance, 3D information of the object is useful. For this Model. The aim of this paper is to fill this gap.
reason, 3D Morphable Models (3DMM) have been in- We describe a 3D Morphable Face Model - the
troduced a decade ago [7]. They have become a well Basel Face Model (BFM) - that is publicly available
established technology able to perform various tasks, ([Link] The usage of the
1
BFM is free for non-commercial purposes. This model
not only allows development of 3DMM based image
analysis algorithms but will also permit new practices
that were impossible before:
First, the 3DMM allows generalization over a va-
riety of different test data sets. Currently, there ex- BFM MPI Model
ist several publicly available face image databases (e.g.
Figure 1. Correspondence artifacts caused by the surface
CMU-PIE [20], FERET [15], etc.) and databases with parametrization occur especially for larger values of model
unregistered 3D face scans (e.g. UND [9]). Each of the coefficients. Rendering the MPI model (right) and the BFM
image databases provides gigabytes of face photographs (left) with the same coefficient vector (ci ∼ N (0, (2.5)2 ))
taken at different poses and illumination conditions. shows that the new BFM displays less artifacts.
These images are either used to train or test new algo-
rithms. Unfortunately, in most cases (e.g. [23, 10]) the Age Weight
same face database is used for both training and test- 90 35
80 30
70
ing. Such recognition systems usually have difficulties 60 25
20
50
to generalize from one database to another, because 40
30
15
20 10
the imaging conditions are too different. 10 5
0 0
The 3DMM, however, can generate face images at
5
10
15
20
25
30
35
40
45
50
45
50
55
60
65
70
75
80
85
90
years kilogramm
any pose and under any illumination. As mentioned
before the face rendering can be used directly in an Figure 2. The BFM was trained on 200 individuals (100f/
100m). Age (avg 25y) and weight (avg 66kg) are distributed
analysis by synthesis approach [7], to fit the model to
over a large range, but peak at students age.
images. Or it can be used indirectly to generate train-
ing or test images at any imaging condition. Hence, in
addition to being a valuable model for use in face anal- reproduce the input images and how accurate is the re-
ysis it can also be viewed as a meta-database which construction? Do the coefficients of different images of
allows the creation of an infinity of accurately labeled the same individual cluster nicely in model space? Can
synthetic training and testing images. the recognition be improved with a different metric in
Registered scans of ten individuals, which are not model space? These and similar questions can only
part of the BFM training set, are provided together be answered if the model coefficients are made public.
with a fixed test set of 270 renderings with pose and Moreover, numerous face image analysis articles de-
light variations. Additionally, synthetic faces can be scribe algorithms but do not release training data. This
generated from random model coefficients. This flexi- hinders reproducible research and fair comparison with
bility can also be used to test specific aspects of face im- other algorithms. To address these two restrictions we
age analysis algorithms: For instance, how the depar- provide both the training data set (the BFM) and the
ture from the Lambertian assumptions affects the per- model fitting results for several standard image data
formance of an algorithm (with or without cast shad- sets (CMU-PIE, FERET and UND) obtained with the
ows and specular lobe, with a sparse set of lights or an state of the art fitting algorithms [18, 2]. We hope
environment map, etc.). The pose can be varied contin- that researchers developing future algorithms based on
uously such that the extent of the pose generalization of the BFM will likewise provide the coefficients to enable
an algorithm can be easily analyzed. For example, test deeper algorithm comparison and accuracy analysis.
images for stereo algorithms with variable baseline, or Currently, to the best of our knowledge, there ex-
for photogrammetric stereo with programmable light ist only two comparable 3DMMs of faces: the Max-
directions can also be easily generated. The bottom Planck-Institut Tübingen (MPI) MM [7] and the Uni-
line is that it is now possible and easy to test the lim- versity of South Florida (USF) MM [19]. Compared
itations of face image analysis algorithms in terms of with these, the BFM is superior in two aspects: Our 3D
pose and illumination generalization. scanner (ABW-3D) offers a higher resolution and pre-
Secondly, the vast majority of face recognition ar- cision in shorter scan time than the Cyberware (TM)
ticles provide results in terms of percentage of rank-1 scanner (used for the MPI and USF models). This
correct identification or False Acceptance Rate (FAR) results in a more accurate model. Secondly, a differ-
/ False Rejection Rate (FRR) curves on a standard ent registration method is used yielding less correspon-
database. However, these numbers do not fully de- dence artifacts (Fig. 1). Additionally, renderings of
scribe the behavior of an algorithm and leave the reader fitting results are more realistic with the new model
with open questions such as: Is the algorithm able to (Sec. 3). We now describe the construction of the
x1 r1
y1 g1
z1 b1
x r
2 2
y g
2 2
Si = Ti =
z2 b2
. .
. .
. .
x r
m m
ym gm
zm bm
Figure 3. Registration establishes a common parametriza-
tion between the original scans (left) and fills in missing Figure 4. Each entry in the data vectors correspond to the
data (right). same point on the faces. In this example the first entry
corresponds to the tip of the nose
model followed by baseline experiments against which
to compare future computer vision algorithms. We de-
scribe results obtained with the BFM in terms of iden- 2.2. Registration
tification experiments and visual quality of the fittings.
To make the raw data usable it needs to be brought
in correspondence. This means, that the scans are
2. Model Construction
re-parameterized such that semantically correspond-
The construction of a 3DMM requires a training set ing points (i.e.} the nose tips or eye corners) share
with a large variety in face shapes and appearance. The the same position in the parametrization domain (Fig.
training data should be a representative sample of the 4). Registration establishes this correspondence for all
target population. The training data set for the BFM points of the face, including unstructured regions like
consists of face scans of 100 female and 100 male per- the cheek. After bringing the scans into correspondence
sons, most of them Europeans. The age of the persons linear combinations of scans, are again faces. The ef-
is between 8 and 62 years with an average of 24.97 years fect of a bad registration on the model quality can be
and the weight is between 40 and 123 kilogram with an seen in figure 1. To establish correspondence we use a
average of 66.48 kilogram (Fig. 2). Each person was modified version of the Optimal Step Nonrigid ICP Al-
scanned three times with neutral expression, and the gorithm [4]. The registration method is applied in the
most natural looking scan was selected. 3D domain on triangulated meshes. It progressively de-
forms a template towards the measured surface while
2.1. 3D face scanning ensuring a smooth deformation. In addition to estab-
Scanning of human faces is a challenging task. To lishing correspondence this method fills in missing re-
capture natural looking faces, the acquisition time is gions by using a robust distance measure. To improve
critical. We use a coded light system with an acquisi- the model quality we manually added landmarks at the
tion time of ∼ 1s. This leads to more accurate results lips, eyebrows and ears.
compared to laser scanners with a acquisition time of
around ∼ 15s. The structured light system was built
2.3. Texture Extraction and Inpainting
by ABW-3D. It uses a sequence of light patterns which
uniquely encode each pixel of the projectors, such that The face albedo is represented by one color per ver-
triangulation can be performed even on unstructured tex, which is calculated from the photographs. The in-
regions like the cheeks. To capture the full face the formation from the three photographs is blended based
system uses two projectors and three cameras resulting on the distance from the visibility boundaries and the
in four depths images. The system captures the facial orientation of the normal relative to the viewing di-
surface from ear to ear with outstanding precision (Fig. rection. To improve the albedo model we manually
3, left). The 3D shape of the eyes and hair cannot be removed hair and completed the missing data using
captured with our system, due to their reflection prop- diffusion.
erties. The resolution of the geometry measurement is
higher than all comparable systems: ABW-3D ∼ 200
k, Cyberware ∼ 75 k and 3Dmd ∼ 20 k. 2.4. Model
Simultaneously with each scan, three photos are
taken with SLR cameras (sRGB color profile). Three After registration the faces are parameterized as tri-
studio flashes with diffuser umbrellas are used to angular meshes with m = 53490 vertices and shared
achieve a homogeneous illumination. This ensures a topology. The vertices (xj , yj , zj )T ∈ R3 have an asso-
higher color fidelity compared to other systems. ciated color (rj , gj , bj )T ∈ [0, 1]3 . A face is then repre-
sented by two 3m dimensional vectors test database by splitting the database into test and
training set, instead it is possible to apply the same
s = (x1 , y1 , z1 , . . . xm , ym , zm )T (1) model to all data sets. With our face identification ex-
T periments, we show that the BFM is general enough to
t = (r1 , g1 , b1 , . . . rm , gm , bm ) .
be used with different 2D and 3D sensors.
The BFM assumes independence between shape and
texture, constructing two independent Linear Models 3.1. Face Identification on 2D images
as described in [7]. A Gaussian distributed is fit to To demonstrate the quality of the presented model
the data using Principle Component Analysis (PCA), we compare it with the MPI model by 2D identifica-
resulting in a parametric face model consisting of tion experiments (CMU-PIE/FERET) [18]. To allow a
Ms = (µs , σs , Us ) and Mt = (µt , σt , Ut ), (2) detailed and transparent analysis of the results and to
enable other researchers to compare their results with
where µ{s,t} ∈ R3m are the mean, σ{s,t} ∈ Rn−1 ours, we provide the reconstructed 3D shapes and tex-
the standard deviations and U{s,t} = [u1 , . . . un ] ∈ tures, coefficients and rendered faces for each test im-
R3m×n−1 are an orthonormal basis of principle com- age. None of the individuals in the test sets is part of
ponents of shape and texture. New faces are generated the training data for the Morphable Model. The test
from the model as linear combinations of the principal sets cover a large ethnic variety.
components
Test Set 1: FERET Subset The subset of the
s(α) = µs + Us diag(σs )α (3) FERET data set [16], consists of 194 individuals across
t(β) = µt + Ut diag(σt )β 9 poses at constant lighting condition except the frontal
view taken under a different illumination condition. In
The coefficients are independent and normally dis- the FERET nomenclature these images correspond to
tributed with unit variance under the assumption of the series ba through bk. We omitted the images bj as
normally distributed training examples and a correct the subjects present a smile and our model can only
mean estimation. represent neutral expressions.
The data necessary to synthesize faces (i.e. the
model data Ms , Mt and the triangulation) together
Test Set 2: CMU-PIE Subset The subset of the
with test data and test results are available at our web
site. We provide: CMU-PIE data set [20], consists of 68 individuals (28
wearing glasses) at 3 poses (frontal, side and profile)
Shape and albedo PCA model Ms , Mt (U, σ, µ) com-
under illumination from 21 different directions and am-
puted from the 200 face scans.
bient light only. To perform the identification we fit the
Ten additional registered 3D face scans together with
2D renderings of the scans with light and pose varia- BFM to the images of the test sets. In the fitting three
tion. error terms based on landmarks, the contour and the
Vertex indices of MPEG and Farkas points together shading are optimized. To extend the flexibility, four
with 2D projections within the above renderings. facial regions (eyes, nose, mouth and the rest Fig. 5)
Model coefficients obtained by the fitting of FERET are fitted separately and combined later by blending
and CMU-PIE together with 2D renderings of the re- them together. The obtained shape and albedo model
constructed shapes. parameters for the global fitting α0 and β0 and for
Model coefficients obtained by the fitting of the unreg- the facial segments α1 , . . . α4 and β1 , . . . β4 represent
istered 3D shapes of the UND database. the identity in the model space. These parameters are
Mask with four segments (Fig. 5) that is used in the
stacked together into one identity vector
identification experiments (Sec. 3.1).
Matlab code for own experiments, e.g. generation of
c = α0 , β0 , . . . α4 , β4 . (4)
random faces.
Similarity of two scans is measured by the angle be-
3. Experiments tween their identity vectors.
Table 1 and 2 list the percentages of correct rank
With the BFM a standard training set for face recog- 1 identification obtained on the CMU-PIE and the
nition algorithms is provided to the public. Together FERET subset, respectively. The overall identification
with test sets such as FERET, CMU-PIE and UND, rate with the BFM model is better than the MPI re-
this allows for a fair, data independent comparison of sults. For CMU-PIE 91.3% (vs. 89.4%) and for FERET
face identification algorithms. We demonstrate that it 95.8% (vs. 92.4%) were obtained. As in previous exper-
is not necessary to train a model specifically for each iments the best results are obtained for frontal views.
shape shape components texture texture components
mean 1st. (+5σ) 2nd. (+5σ) 3rd. (+5σ) mean 1st. (+5σ) 2nd. (+5σ) 3rd. (+5σ) Mask
1st. (−5σ) 2nd. (−5σ) 3rd. (−5σ) 1st. (−5σ) 2nd. (−5σ) 3rd. (−5σ)
Figure 5. The mean together with the first three principle components of the shape (left) and texture (right) PCA model.
Shown is the mean shape resp. texture plus/minus five standard deviations σ. Mask with the four manually chosen segments
(eyes, nose, mouth and rest) used in the fitting to extend the flexibility.
Gallery / Probe front side profile mean
front 98.9 % 96.1 % 75.7 % 90.2 %
side 96.9 % 99.9 % 87.8 % 94.9 %
profile 79.0 % 89.0 % 98.3 % 88.8 %
mean 91.6 % 95.0 % 87.3 % 91.3 %
Table 1. Rank 1 identification results obtained on a CMU-
PIE subset. The mean identification rate is 91.3%. With
the former MPI model a identification rate of 89.4% was
obtained.
Figure 6. Exemplary fitting result for CMU-PIE with BFM
Face Model. Left the original image, middle row the fitting
result rendered into the image and right the resulting 3D
Gallery / Probe Pose Φ Identification rate model.
bb 38.9 ° 97.4 %
bc 27.4 ° 99.5 %
bd 18.9 ° 100.0 % 3.2. Face Identification on 3D scans
be 11.2 ° Gallery For the 3D identification experiments, we fit the
ba 1.1 ° 99.0 % BFM to shape data without using the texture. The
bf -7.1 ° 99.5 % fitting algorithm [2] is a variant of the nonrigid ICP
bg -16.3 ° 97.9 % work in [4]. We initialize the fitting by locating the tip
bh -26.5 ° 94.8 % of the nose with the method of [21]. As test set we
bi -37.9 ° 83.0 % use the UND database [9] that consists of 953 unregis-
bk 0.1 ° 90.7 % tered 3D scans, with one to eight scans per subject. As
mean 95.8 % for the 2D experiments, we measure the similarity be-
Table 2. Rank 1 identification results obtained on a FERET tween two faces as the angle between their coefficients
subset. The mean identification rate is 95.8%. With the for-
in Mahalanobis space. The recognition performance
mer MPI model a identification rate of 92.4% was obtained.
for different distance thresholds is shown in Fig. 7.
4. Conclusion
Compared with the MPI, the visual quality of the BFM We presented a publicly available 3D Morphable
fitting results (Fig. 6) is much better since the overfit- Model of faces, together with basic experiments. The
ting in the texture reconstruction has been reduced. model addresses the lack of universal training data for
Shape-based Recognition Performance
[4] B. Amberg, S. Romdhani, and T. Vetter. Optimal
step nonrigid ICP algorithms for surface registration.
8
7 Basel Face Model In CVPR ’07.
6 [5] O. Arandjelovic, G. Shakhnarovich, J. Fisher,
5 R. Cipolla, and T. Darrell. Face recognition with im-
FRR %
4 age sets using manifold density divergence. CVPR ’05,
3
1, 2005.
2
1 [6] V. Blanz, C. Basso, T. Poggio, and T. Vetter. Rean-
0 imating faces in images and video. In EuroGraphics,
0 1 2 3 4 5 2003.
FAR % [7] V. Blanz and T. Vetter. A morphable model for the
Figure 7. Identification results obtained on the UND synthesis of 3D faces. In SIGGRAPH ’99.
database of unregistered 3D shapes. Varying the distance [8] V. Blanz and T. Vetter. Face recognition based on
threshold leads to varying false acceptance rates (FAR) and fitting a 3D morphable model. PAMI, 25(9), 2003.
false rejection rates (FRR). [9] K. I. Chang, K. W. Bowyer, and P. J. Flynn. An eval-
uation of multimodal 2D+3D face biometrics. PAMI,
27(4), 2005.
face recognition. Although many test data sets exist, [10] R. Gross, I. Matthews, and S. Baker. Appearance-
there are no standard training data sets. The reason based face recognition and light-fields. PAMI,
is that such a training set must be general enough to 26(4):449–465, 2004.
represent all appearance of faces under any pose and [11] B. Heisele, T. Serre, and T. Poggio. A component-
illumination condition. Since we believe that a stan- based framework for face detection and identification.
dard training set is necessary for a fair comparison, IJCV, 74(2), 2007.
we make the model publicly available. Due to its 3D [12] Y. Hu, D. Jiang, S. Yan, L. Zhang, and H. Zhang.
structure it can be used indirectly to generate images Automatic 3D reconstruction for face recognition. fg,
with any kind of pose and light variation or directly 0, 2004.
for 2D and 3D face recognition. It is planned to ex- [13] K.-C. Lee, J. Ho, M.-H. Yang, and D. Kriegman.
Video-based face recognition using probabilistic ap-
tend the data collection further and provide it on the
pearance manifolds. CVPR, 01, 2003.
web site. We also plan to provide results of experi-
[14] D. A. Leopold, A. J. O’Toole, T. Vetter, and V. Blanz.
ments and renderings with more complex illumination Prototype-referenced shape encoding revealed by high-
models. Using these standardized training and test level aftereffects. Nature Neuroscience, 4(1):89–94,
sets makes it possible for researchers to focus on the 2001.
comparison of algorithms independent of the data. We [15] P. J. Phillips, H. Moon, S. A. Rizvi, and P. J. Rauss.
trained our previously published face recognition algo- The feret evaluation methodology for face-recognition
rithm and provide detailed results (parameters for the algorithms. PAMI, 22(10), 2000.
model). Other researchers are invited to use the same [16] P. J. Phillips, H. Moon, S. A. Rizvi, and P. J.
standardized test set and present the results on our Rauss. The FERET evaluation methodology for face-
web site ([Link] recognition algorithms. PAMI, 22, 2000.
[17] S. Romdhani. Face Image Analysis Using a Multiple
4.1. Acknowledgment Features Fitting Strategy. PhD thesis, 2005.
[18] S. Romdhani and T. Vetter. Estimating 3D shape
This work was funded in part by the Swiss National and texture using pixel intensity, edges, specular high-
Science Foundation (200021-103814, NCCR CO-ME lights, texture constraints and a prior. In CVPR ’05.
5005-66380) and Microsoft Research. [19] S. Sarkar. USF HumanID 3D face dataset, 2005.
[20] T. Sim, S. Baker, and M. Bsat. The CMU pose, il-
References lumination, and expression database. PAMI, 25(12),
2003.
[1] B. Allen, B. Curless, and Z. Popović. The space of [21] F. B. ter Haar and R. C. Veltkamp. A 3D Face Match-
human body shapes: reconstruction and parameteri- ing Framework. In Shape Modeling Int. ’08.
zation from range scans. In SIGGRAPH ’03.
[22] W. Zhao, R. Chellappa, P. J. Phillips, and A. Rosen-
[2] B. Amberg, R. Knothe, and T. Vetter. Expression feld. Face recognition: A literature survey. ACM Com-
invariant 3D face recognition with a morphable model. put. Surv., 35(4), 2003.
In FG’08, 2008.
[23] S. Zhou and R. Chellappa. Illuminating light field:
[3] B. Amberg, S. Romdhani, A. Fitzgibbon, A. Blake, image-based face recognition across illuminations and
and T. Vetter. Accurate surface extraction using model poses. FG, 2004.
based stereo. In ICCV ’07, 2007.