Unconstrained Still/Video-Based Face Verification with Deep Convolutional Neural Networks

Springer Science and Business Media LLC - Tập 126 - Trang 272-291 - 2017
Jun-Cheng Chen1, Rajeev Ranjan1, Swami Sankaranarayanan1, Amit Kumar1, Ching-Hui Chen1, Vishal M. Patel2, Carlos D. Castillo1, Rama Chellappa1
1University of Maryland, College Park, College Park, USA
2Department of Electrical and Computer Engineering, Rutgers University, Piscataway, USA

Tóm tắt

Over the last 5 years, methods based on Deep Convolutional Neural Networks (DCNNs) have shown impressive performance improvements for object detection and recognition problems. This has been made possible due to the availability of large annotated datasets, a better understanding of the non-linear mapping between input images and class labels as well as the affordability of GPUs. In this paper, we present the design details of a deep learning system for unconstrained face recognition, including modules for face detection, association, alignment and face verification. The quantitative performance evaluation is conducted using the IARPA Janus Benchmark A (IJB-A), the JANUS Challenge Set 2 (JANUS CS2), and the Labeled Faces in the Wild (LFW) dataset. The IJB-A dataset includes real-world unconstrained faces of 500 subjects with significant pose and illumination variations which are much harder than the LFW and Youtube Face datasets. JANUS CS2 is the extended version of IJB-A which contains not only all the images/frames of IJB-A but also includes the original videos. Some open issues regarding DCNNs for face verification problems are then discussed.

Tài liệu tham khảo

AbdAlmageed, W., Wu, Y., Rawls, S., Harel, S., Hassne, T., Masi, I., et al. (2016). Face recognition using deep multi-pose representations. In IEEE winter conference on applications of computer vision (WACV). Ahonen, T., Hadid, A., & Pietikainen, M. (2006). Face description with local binary patterns: Application to face recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(12), 2037–2041. Ahuja, R., Magnanti, T., & Orlin, J. (1993). Network flows: Theory, algorithms, and applications. Englewood Cliffs: Prentice Hall. Asthana, A., Zafeiriou, S., Cheng, S., & Pantic, M. (2013). Robust discriminative response map fitting with constrained local models. In IEEE conference on computer vision and pattern recognition (pp. 3444–3451). Asthana, A., Zafeiriou, S., Cheng, S. Y., & Pantic, M. (2013). Robust discriminative response map fitting with constrained local models. In IEEE conference on computer vision and pattern recognition (pp. 3444–3451). Babenko, B., Yang, M. H., & Belongie, S. (2009). Visual tracking with online multiple instance learning. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 983–990). IEEE. Bae, S. H., & Yoon, K. J. (2014). Robust online multi-object tracking based on tracklet confidence and online discriminative appearance learning. In IEEE conference on computer vision and pattern recognition (CVPR). Belhumeur, P. N., Jacobs, D. W., Kriegman, D. J., & Kumar, N. (2013). Localizing parts of faces using a consensus of exemplars. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(12), 2930–2940. Bodla, N., Zheng, J., Xu, H., Chen, J. C., Castillo, C. D., & Chellappa, R. (2017). Deep heterogeneous feature fusion for template-based face recognition. In IEEE winter conference on applications of computer vision (WACV). Breitenstein, M. D., Reichlin, F., Leibe, B., Koller-Meier, E., & Gool, L. V. (2009). Robust tracking-by-detection using a detector confidence particle filter. In IEEE international conference on computer vision (ICCV). Bruna, J., & Mallat, S. (2013). Invariant scattering convolution networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8), 1872–1886. Burgos-Artizzu, X. P., Perona, P., & Dollár, P. (2013). Robust face landmark estimation under occlusion. In IEEE international conference on computer vision, ICCV ’13 (pp. 1513–1520). Washington, DC: IEEE Computer Society. doi:10.1109/ICCV.2013.191. Cao, X., Wei, Y., Wen, F., & Sun, J. (2014). Face alignment by explicit shape regression. http://www.google.com/patents/US20140185924. US Patent App. 13/728,584. Chen, D., Cao, X. D., Wang, L. W., Wen, F., & Sun, J. (2012a). Bayesian face revisited: A joint formulation. In European conference on computer vision (pp. 566–579). Chen, D., Cao, X. D., Wen, F., & Sun, J. (2013). Blessing of dimensionality: High-dimensional feature and its efficient compression for face verification. In IEEE conference on computer vision and pattern recognition. Chen, J. C., Patel, V. M., & Chellappa, R. (2015a). Unconstrained face verification using deep cnn features. arXiv:1508.01722. Chen, Y. C., Patel, V. M., Phillips, P. J., & Chellappa, R. (2012b). Dictionary-based face recognition from video. In European conference on computer vision (ECCV). Chen, J. C., Ranjan, R., Kumar, A., Chen, C. H., Patel, V. M., & Chellappa, R. (2015b). An end-to-end system for unconstrained face verification with deep convolutional neural networks. In IEEE international conference on computer vision workshop on ChaLearn looking at people (pp. 118–126). Chen, D., Ren, S., Wei, Y., Cao, X., & Sun, J. (2014) Joint cascade face detection and alignment. In D. Fleet, T. Pajdla, B. Schiele & T. Tuytelaars (eds.) European conference on computer vision (Vol. 8694, pp. 109–122). Chen, J. C., Sankaranarayanan, S., Patel, V. M., & Chellappa, R. (2015c). Unconstrained face verification using Fisher vectors computed from frontalized faces. In IEEE international conference on biometrics: Theory, applications and systems. Cheney, J., Klein, B., Jain, A. K., & Klare, B. F. (2015). Unconstrained face detection: State of the art baseline and challenges. In International conference on biometrics. Comaschi, F., Stuijk, S., Basten, T., & Corporaal, H. (2015). Online multi-face detection and tracking using detector confidence and structured SVMs. In IEEE international conference on advanced video and signal based surveillance (AVSS). Cootes, T. F., Edwards, G. J., & Taylor, C. J. (2001). Active appearance models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 6, 681–685. Cootes, T. F., Taylor, C. J., Cooper, D. H., & Graham, J. (1995). Active shape models-their training and application. Computer Vision and Image Understanding, 61(1), 38–59. Cristinacce, D., & Cootes, T. F. (2006). Feature detection and tracking with constrained local models. In British machine vision conference (Vol. 1, p. 3). Crosswhite, N., Byrne, J., Parkhi, O. M., Stauffer, C., Cao, Q., & Zisserman, A. (2016). Template adaptation for face verification and identification. arXiv:1603.03958. Davis, J. V., Kulis, B., Jain, P., Sra, S., & Dhillon, I. S. (2007). Information-theoretic metric learning. In International conference on machine learning (pp. 209–216). Ding, C., & Tao, D. (2015). Robust face recognition via multimodal deep face representation. arXiv:1509.00244. Dollár, P., Welinder, P., & Perona, P. (2010). Cascaded pose regression. In IEEE conference on computer vision and pattern recognition (pp. 1078–1085). IEEE. Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., et al. (2013). Decaf: A deep convolutional activation feature for generic visual recognition. arXiv:1310.1531. Du, M., & Chellappa, R. (2012). Face association across unconstrained video frames using conditional random fields. In European conference on computer vision (ECCV). Duffner, S., & Odobez, J. (2013). Track creation and deletion framework for long-term online multiface tracking. IEEE Transactions on Image Processing, 22(1), 272–285. Everingham, M., Gool, L. V., Williams, C. K. I., Winn, J., & Zisserman, A. (2009). The pascal visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303–338. Farfade, S. S., Saberian, M. J., & Li, L. J. (2015). Multi-view face detection using deep convolutional neural networks. In International conference on multimedia retrieval. Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014a). Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE conference on computer vision and pattern recognition (pp. 580–587). Girshick, R., Iandola, F., Darrell, T., & Malik, J. (2014b). Deformable part models are convolutional neural networks. In IEEE conference on computer vision and pattern recognition. Giryes, R., Sapiro, G., & Bronstein, A. M. (2014). On the stability of deep networks. arXiv:1412.5896. Gross, R., Matthews, I., Cohn, J., Kanade, T., & Baker, S. (2010). Multi-pie. Image and Vision Computing, 28(5), 807–813. Guillaumin, M., Verbeek, J., & Schmid, C. (2009) Is that you? Metric learning approaches for face identification. In IEEE international conference on computer vision (pp. 498–505). Haeffele, B. D., & Vidal, R. (2015). Global optimality in tensor factorization, deep learning, and beyond. arXiv:1506.07540. Hassner, T., Harel, S., Paz, E., & Enbar, R. (2015). Effective face frontalization in unconstrained images. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 4295–4304). He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. arXiv:1502.01852. Henriques, J. F., Caseiro, R., Martins, P., & Batista, J. (2015). High-speed tracking with kernelized correlation filters. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37(3), 583–596. Hu, J., Lu, J., & Tan, Y. P. (2014). Discriminative deep metric learning for face verification in the wild. In IEEE conference on computer vision and pattern recognition (pp. 1875–1882). Huang, G. B., Mattar, M., Berg, T., & Learned-Miller, E. (2008a). Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In Workshop on faces in real-life images: detection, alignment, and recognition. Huang, C., Wu, B., & Nevatia, R. (2008b). Robust object tracking by hierarchical association of detection responses. In European conference on computer vision (ECCV). Jain, V., & Learned-Miller, E. (2010). Fddb: A benchmark for face detection in unconstrained settings. UM-CS-2010-009. Kalal, Z., Mikolajczyk, K., & Matas, J. (2012). Tracking-learning-detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(7), 1409–1422. Kazemi, V., & Sullivan, J. (2014). One millisecond face alignment with an ensemble of regression trees. In IEEE conference on computer vision and pattern recognition (pp. 1867–1874). Klare, B. F., Klein, B., Taborsky, E., Blanton, A., Cheney, J., Allen, K., et al. (2015). Pushing the frontiers of unconstrained face detection and recognition: IARPA Janus Benchmark A. In IEEE conference on computer vision and pattern recognition. Koestinger, M., Wohlhart, P., Roth, P. M., & Bischof, H. (2011). Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization. In First IEEE international workshop on benchmarking facial image analysis technologies. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems (pp. 1097–1105). Kumar, A., Ranjan, R., Patel, V., & Chellappa, R. (2016). Face alignment by local deep descriptor regression. arXiv:1601.07950. Le, V., Brandt, J., Lin, Z., Bourdev, L., & Huang, T. S. (2012). Interactive facial feature localization. In European conference on computer vision (pp. 679–692). Springer. Li, J., & Zhang, Y. (2013). Learning surf cascade for fast and accurate object detection. In 2013 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 3468–3475). doi:10.1109/CVPR.2013.445. Li, H., Hua, G., Lin, Z., Brandt, J., & Yang, J. (2013). Probabilistic elastic part model for unsupervised face detector adaptation. In IEEE international conference on computer vision (pp. 793–800). Li, H., Lin, Z., Shen, X., Brandt, J., & Hua, G. (2015). A convolutional neural network cascade for face detection. In IEEE conference on computer vision and pattern recognition (pp. 5325–5334). Liao, S., Jain, A. K., & Li, S. Z. (2016). A fast and accurate unconstrained face detector. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(2), 211–223. Long, M., & Wang, J. (2015). Learning transferable features with deep adaptation networks. arXiv:1502.02791. Lui, Y. M., Beveridge, R. J., & Whitley, L. D. (2010). Adaptive appearance model and condensation algorithm for robust face tracking. IEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans, 40(3), 437–448. Mallat, S. (2016). Understanding deep convolutional networks. arXiv:1601.04920. Masi, I., Tran, A. T., Leksut, J. T., Hassner, T., & Medioni, G. (2016). Do we really need to collect millions of faces for effective face recognition? arXiv:1603.07057. Mathias, M., Benenson, R., Pedersoli, M., & Gool, L. V. (2014). Face detection without bells and whistles. In European conference on computer vision (Vol. 8692, pp. 720–735). Mignon, A., & Jurie, F. (2012). Pcca: A new approach for distance learning from sparse pairwise constraints. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 2666–2672). National institute of standards and technology (NIST): IARPA Janus benchmark-a performance report. http://biometrics.nist.gov/cs_links/face/face_challenges/IJBA_reports.zip. Nguyen, H. V., & Bai, L. (2010). Cosine similarity metric learning for face verification. In Asian conference on computer vision (pp. 709–720). Springer. Parkhi, O. M., Vedaldi, A., & Zisserman, A. (2015). Deep face recognition. In British machine vision conference. Ranjan, R., Castillo, C. D., & Chellappa, R. (2017). L2-constrained softmax loss for discriminative face verification. arXiv:1703.09507. Ranjan, R., Patel, V. M., & Chellappa, R. (2015). A deep pyramid deformable part model for face detection. In IEEE international conference on biometrics: Theory, applications and systems. Ranjan, R., Patel, V. M., & Chellappa, R. (2016a). HyperFace: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition. arXiv:1603.01249. Ranjan, R., Sankaranarayanan, S., Castillo, C. D., & Chellappa, R. (2016b). An all-in-one convolutional neural network for face analysis. arXiv:1611.00851. Ren, S., Cao, X., Wei, Y., & Sun, J. (2014). Face alignment at 3000 fps via regressing local binary features. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 1685–1692). doi:10.1109/CVPR.2014.218 Ross, G. (2015). Fast r-cnn. In IEEE international conference on computer vision (pp. 1440–1448). Roth, M., Bauml, M., Nevatia, R., & Stiefelhagen, R. (2012). Robust multi-pose face tracking by multi-stage tracklet association. In International conference on pattern recognition (ICPR). RoyChowdhury, A., Lin, T. Y., Maji, S., & Learned-Miller, E. (2016). One-to-many face recognition with bilinear cnns. In IEEE winter conference on applications of computer vision (WACV). Sankaranarayanan, S., Alavi, A., Castillo, C. D., & Chellappa, R. (2016a). Triplet probabilistic embedding for face verification and clustering. In IEEE international conference on biometrics theory, applications and systems (BTAS) (pp. 1–8). Sankaranarayanan, S., Alavi, A., & Chellappa, R. (2016b). Triplet similarity embedding for face verification. arXiv:1602.03418. Schroff, F., Kalenichenko, D., & Philbin, J. (2015). Facenet: A unified embedding for face recognition and clustering. arXiv:1503.03832. Shi, J., & Tomasi, C. (1994). Good features to track. In IEEE conference on computer vision and pattern recognition (CVPR). Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Simonyan, K., Parkhi, O. M., Vedaldi, A., & Zisserman, A. (2013). Fisher vector faces in the wild. In British machine vision conference (Vol. 1, p. 7). Sun, Y., Liang, D., Wang, X., & Tang, X. (2015). Deepid3: Face recognition with very deep neural networks. arXiv:1502.00873. Sun, Y., Wang, X., & Tang, X. (2013). Deep convolutional network cascade for facial point detection. In IEEE conference on computer vision and pattern recognition (pp. 3476–3483). Sun, Y., Wang, X., & Tang, X. (2014). Deeply learned face representations are sparse, selective, and robust. arXiv:1412.1265. Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., et al. (2014). Going deeper with convolutions. arXiv:1409.4842. Taigman, Y., Wolf, L., & Hassner, T. (2009). Multiple one-shots for utilizing class label information. In British machine vision conference (pp. 1–12). Taigman, Y., Yang, M., Ranzato, M. A., & Wolf, L. (2014). Deepface: Closing the gap to human-level performance in face verification. In IEEE conference on computer vision and pattern recognition (pp. 1701–1708). Tzimiropoulos, G., Pantic, M. (2014). Gauss–Newton deformable part models for face alignment in-the-wild. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 1851–1858). doi:10.1109/CVPR.2014.239. Viola, P., & Jones, M. J. (2004). Robust real-time face detection. International Journal of Computer Vision, 57(2), 137–154. Wang, D., Otto, C., & Jain, A. K. (2015). Face search at scale: 80 million gallery. arXiv:1507.07242. Wang, P., & Ji, Q. (2008). Robust face tracking via collaboration of generic and specific models. IEEE Transactions on Image Processing, 17(7), 1189–1199. Weinberger, K. Q., & Saul, L. K. (2009). Distance metric learning for large margin nearest neighbor classification. Journal of Machine Learning Research, 10, 207–244. Wolf, L., Hassner, T., & Taigman, Y. (2009). The one-shot similarity kernel. In International conference on computer vision (pp. 897–902). IEEE. Xiong, X., & la Torre, F. D. (2013). Supervised descent method and its applications to face alignment. In 2013 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 532–539). doi:10.1109/CVPR.2013.75. Xiong, L., Karlekar, J., Zhao, J., Feng, J., Pranata, S., & Shen, S. (2017). A good practice towards top performance of face recognition: Transferred deep feature fusion. arXiv:1704.00438. Yan, J., Zhang, X., Lei, Z., & Li, S. Z. (2014). Face detection by structural models. Image and Vision Computing, 32(10), 790–799. doi:10.1016/j.imavis.2013.12.004. http://www.sciencedirect.com/science/article/pii/S0262885613001765. Best of Automatic Face and Gesture Recognition 2013. Yang, S., Luo, P., Loy, C. C., & Tang, X. (2015a). From facial parts responses to face detection: A deep learning approach. In IEEE international conference on computer vision. Yang, J., Ren, P., Chen, D., Wen, F., Li, H., & Hua, G. (2016). Neural aggregation network for video face recognition. arXiv:1603.05474. Yang, J., Ren, P., Zhang, D., Chen, D., Wen, F., Li, H., et al. (2017). Neural aggregation network for video face recognition. In IEEE conference on computer vision and pattern recognition (CVPR). Yang, B., Yan, J., Lei, Z., & Li, S. Z. (2015b). Convolutional channel features. In IEEE international conference on computer vision. Yi, D., Lei, Z., Liao, S., & Li, S. Z. (2014). Learning face representation from scratch. arXiv:1411.7923. Yoon, J. H., Yang, M. H., Lim, J., & Yoon, K. J. (2015). Bayesian multi-object tracking using motion context from multiple objects. In IEEE winter conference on applications of computer vision (WACV). Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? In Advances in neural information processing systems (NIPS) (pp. 3320–3328). Zhang, J., Shan, S., Kan, M., & Chen, X. (2014). Coarse-to-fine auto-encoder networks (CFAN) for real-time face alignment. In European conference on computer vision ECCV (pp. 1–16). doi:10.1007/978-3-319-10605-2_1. Zhao, W. Y., Chellappa, R., Phillips, P. J., & Rosenfeld, A. (2003). Face recognition: A literature survey. ACM Computing Surveys, 35(4), 399–458. Zhu, S., Li, C., Chen, C. L., & Tang, X. (2015). Face alignment by coarse-to-fine shape searching. In IEEE conference on computer vision and pattern recognition (CVPR). Zhu, X., & Ramanan, D. (2012). Face detection, pose estimation, and landmark localization in the wild. In IEEE conference on computer vision and pattern recognition (CVPR) (pp. 2879–2886). IEEE.