How far are we from solving the 2D & 3D Face Alignment problem?
(and a dataset of 230,000 3D facial landmarks)
Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017
Abstract
This paper investigates how far a very deep neural network is from attaining close to saturating performance on existing 2D and 3D face alignment datasets. To this end, we make the following five contributions: (a) we construct, for the first time, a very strong baseline by combining a state-of-the-art architecture for landmark localization with a state-of-the-art residual block, train it on a very large yet synthetically expanded 2D facial landmark dataset and finally evaluate it on all other 2D facial landmark datasets.
(b) We create a guided by 2D landmarks network which converts 2D landmark annotations to 3D and unifies all existing datasets, leading to the creation of LS3D-W, the largest and most challenging 3D facial landmark dataset to date (~230,000 images). (c) Following that, we train a neural network for 3D face alignment and evaluate it on the newly introduced LS3D-W.
(d) We further look into the effect of all "traditional" factors affecting face alignment performance like large pose, initialization and resolution, and introduce a "new" one, namely the size of the network. (e) We show that both 2D and 3D face alignment networks achieve performance of remarkable accuracy which is probably close to saturating the datasets used.
Key contributions
-
A strong 2D baseline
Combine a powerful landmark localization architecture with improved residual blocks, then evaluate across facial landmark datasets.
-
Guided annotation conversion
Use images and existing 2D landmarks to produce annotations following a consistent 3D landmark convention.
-
Large-scale 3D evaluation
Introduce LS3D-W and evaluate a dedicated 3D face alignment network on the unified annotations.
-
Robustness and model capacity
Study the effects of pose, resolution, initialization and network size.
-
Benchmark accuracy
Report performance approaching saturation on the datasets evaluated in the 2017 study.
Models
The PyTorch library downloads the appropriate weights for your installed version automatically. Use TWO_D for 2D landmarks or THREE_D for 3D landmarks. The manual downloads below are the checkpoints used by version 1.5.0.
2D-FAN
Predict 68 facial landmarks as image coordinates (x, y) using the PyTorch library’s TWO_D mode.
PyTorch v1.5 checkpoints · 91 MiB
3D-FAN
Predict 68 landmarks with image coordinates x, y and estimated depth z. The PyTorch library’s THREE_D mode combines 3D-FAN predictions with the depth network.
PyTorch v1.5 checkpoints · 91 MiB + 224 MiB depth
Install
pip install face-alignment
Python example
import face_alignment
# Change TWO_D to THREE_D for 3D landmarks.
model = face_alignment.FaceAlignment(
face_alignment.LandmarksType.TWO_D, device="cpu"
)
landmarks = model.get_landmarks_from_image("face.jpg")
Original Torch7 release
For reproducing the paper’s numerical evaluations, consult the original Torch7 implementation and the checkpoints below.
2D-FAN
Predict facial landmarks in 2D.
Torch7 weights · 183 MiB
3D-FAN
Predict image-plane projections of landmarks following the 3D annotation convention.
Torch7 weights · 183 MiB
2D-to-3D-FAN
Convert compatible 2D annotations to the projected 3D landmark convention, guided by the input image.
Torch7 archive · 338 MiB
3D-FAN-depth
Estimate landmark depth. Combine with landmark predictions to obtain full x, y and z coordinates.
Torch7 weights · 455 MiB
Dataset: LS3D-W
LS3D-W brings approximately 230,000 face images into a consistent 68-point landmark annotation scheme. The paper’s 2D-to-3D network automatically converts compatible 2D annotations, creating unified data for training and evaluating 3D face alignment.
The full release and a balanced subset are available through the request form. The original Torch7 release also provides the 2D-to-3D-FAN model for converting compatible 2D landmarks.
Request the LS3D-W dataset
Enter your email address to display the full dataset and balanced-subset download links.
Questions about data access? Contact adrian@adrianbulat.com.
Citation
If you use this work, please cite:
@inproceedings{bulat2017far,
title={How far are we from solving the 2D \& 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)},
author={Bulat, Adrian and Tzimiropoulos, Georgios},
booktitle={International Conference on Computer Vision},
year={2017}
}
Watch the video demonstration
References
- M. Köstinger, P. Wohlhart, P. M. Roth and H. Bischof. Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization ICCV Workshops, 2011.
- J. Shen, S. Zafeiriou, G. G. Chrysos, J. Kossaifi, G. Tzimiropoulos and M. Pantic. The first facial landmark tracking in-the-wild challenge: Benchmark and results ICCV Workshops, 2015.
- C. Sagonas, G. Tzimiropoulos, S. Zafeiriou and M. Pantic. 300 faces in-the-wild challenge: The first facial landmark localization challenge ICCV Workshops, 2013.
- V. Jain and E. Learned-Miller. FDDB: A Benchmark for Face Detection in Unconstrained Settings UMass Amherst Technical Report, 2010.