Real-time 3D reconstruction techniques applied in dynamic scenes: A systematic literature review

Computer Science Review - Tập 39 - Trang 100338 - 2021
Anupama K. Ingale1, Divya Udayan J.1
1School of Information Technology and Engineering, Vellore Institute of Technology, Vellore, India

Tài liệu tham khảo

J. Divya Udayan, H. Kim, Constrained procedural modeling of real buildings from single facade layout, Int. J. Comput. Vis. Signal Process. 6 (1). Udayan, 2016, An analysis of reconstruction algorithms applied to 3d building modeling, Indian J. Sci. Technol., 9, 33 Curless, 1996, A volumetric method for building complex models from range images, 303 Newcombe, 2011, Kinectfusion: Real-time dense surface mapping and tracking, 127 Zach, 2008, Fast and high quality fusion of depth maps De Aguiar, 2008 Vlasic, 2009, Dynamic shape capture using multi-view photometric stereo, 174 Sarbolandi, 2015, Kinect range sensing: Structured-light versus time-of-flight kinect, Comput. Vis. Image Underst., 139, 1, 10.1016/j.cviu.2015.05.006 Marani, 2015, A compact 3d omnidirectional range sensor of high resolution for robust reconstruction of environments, Sensors, 15, 2283, 10.3390/s150202283 Li, 2009, Robust single-view geometry and motion reconstruction, ACM Trans. Graph., 28, 175, 10.1145/1618452.1618521 Kitchenham, 2004, Procedures for performing systematic literature reviews, 33 Dou, 2016, Fusion4d: Real-time performance capture of challenging scenes, ACM Trans. Graph., 35, 114, 10.1145/2897824.2925969 S. Izadi, D. Kim, O. Hilliges, D. Molyneaux, R. Newcombe, P. Kohli, J. Shotton, S. Hodges, D. Freeman, A. Davison, et al. Kinectfusion: real-time 3d reconstruction and interaction using a moving depth camera, in: Proceedings of the 24th Annual ACM Symposium on User Interface Software and Technology, 2011, pp. 559–568. Newcombe, 2011, Kinectfusion: Real-time dense surface mapping and tracking, 127 Keller, 2013, Real-time 3d reconstruction in dynamic scenes using point-based fusion, 1 R.A. Newcombe, D. Fox, S.M. Seitz, Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 343–352. Innmann, 2016, Volumedeform: Real-time volumetric non-rigid reconstruction, 362 Dou, 2017, Motion2fusion: Real-time volumetric performance capture, ACM Trans. Graph., 36, 246, 10.1145/3130800.3130801 M. Slavcheva, M. Baust, D. Cremers, S. Ilic, Killingfusion: Non-rigid 3d reconstruction without correspondences, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1386–1395. Wang, 2018, Dynamic non-rigid objects reconstruction with a single rgb-d sensor, Sensors, 18, 886, 10.3390/s18030886 Kozlov, 2018, Patch-based non-rigid 3d reconstruction from a single depth stream, 42 Guo, 2017, Real-time geometry albedo and motion reconstruction using a single rgb-d camera, ACM Trans. Graph., 36, 32, 10.1145/3083722 M. Slavcheva, M. Baust, S. Ilic, Sobolevfusion: 3d reconstruction of scenes undergoing free non-rigid motion, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2646–2655. Lu, 2018, Reconstructing non-rigid object with large movement using a single depth camera, Comput. Aided Geom. Design, 64, 15, 10.1016/j.cagd.2018.06.002 Zollhöfer, 2014, Real-time non-rigid reconstruction using an rgb-d camera, ACM Trans. Graph., 33, 156, 10.1145/2601097.2601165 Gu, 2018, Monocular 3d reconstruction of multiple non-rigid objects by union of non-linear spatial–temporal subspaces, 14 T. Yu, K. Guo, F. Xu, Y. Dong, Z. Su, J. Zhao, J. Li, Q. Dai, Y. Liu, Bodyfusion: Real-time capture of human motion and surface geometry using a single depth camera, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 910–919. Xu, 2018, Monoperfcap: Human performance capture from monocular video, ACM Trans. Graph., 37, 27, 10.1145/3181973 Lowe, 1999, Object recognition from local scale-invariant features, 1150 Lowe, 2004, Distinctive image features from scale-invariant keypoints, Int. J. Comput. Vis., 60, 91, 10.1023/B:VISI.0000029664.99615.94 Bay, 2006, Surf: Speeded up robust features, 404 Rublee, 2011, Orb: An efficient alternative to sift or surf, 2564 Henry, 2012, Rgb-d mapping: Using kinect-style depth cameras for dense 3d modeling of indoor environments, Int. J. Robot. Res., 31, 647, 10.1177/0278364911434148 Dai, 2017, Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration, ACM Trans. Graph., 36, 1, 10.1145/3072959.3054739 Whelan, 2013, Robust real-time visual odometry for dense rgb-d mapping, 5724 Li, 2015, Real-time head pose tracking with online face template reconstruction, IEEE Trans. Pattern Anal. Mach. Intell., 38, 1922, 10.1109/TPAMI.2015.2500221 T. Yu, Z. Zheng, K. Guo, J. Zhao, Q. Dai, H. Li, G. Pons-Moll, Y. Liu, Doublefusion: Real-time capture of human performances with inner body shapes from a single depth sensor, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7287–7296. Besl, 1992, Method for registration of 3-d shapes, 586 Chen, 1992, Object modeling by registration of multiple range images., Image Vis. Comput., 10, 145, 10.1016/0262-8856(92)90066-C S. Wang, S. Ryan Fanello, C. Rhemann, S. Izadi, P. Kohli, The global patch collider in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 127–135. Sorkine, 2007, As-rigid-as-possible surface modeling, 109 A. Mustafa, H. Kim, J.-Y. Guillemaut, A. Hilton, General dynamic scene reconstruction from multiple view video, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 900–908. A. Mustafa, H. Kim, J.-Y. Guillemaut, A. Hilton, Temporally coherent 4d reconstruction of complex dynamic scenes, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4660–4669. T. Yu, Z. Zheng, Y. Zhong, J. Zhao, Q. Dai, G. Pons-Moll, Y. Liu, Simulcap: Single-view human performance capture with cloth simulation, arXiv preprint arXiv:1903.06323. Wang, 2017, Templateless non-rigid reconstruction and motion tracking with a single rgb-d camera, IEEE Trans. Image Process., 26, 5966, 10.1109/TIP.2017.2740624 Zhang, 2019, Interactionfusion: real-time reconstruction of hand poses and deformable objects in hand-object interactions, ACM Trans. Graph., 38, 48, 10.1145/3306346.3322998 Du, 2010, Affine iterative closest point algorithm for point set registration, Pattern Recognit. Lett., 31, 791, 10.1016/j.patrec.2010.01.020 Dempster, 1977, Maximum likelihood from incomplete data via the em algorithm, J. R. Stat. Soc. Ser. B Stat. Methodol., 39, 1 Kavan, 2007, Kinning with dual quaternions, 39 Sumner, 2007, Embedded deformation for shape manipulation, ACM Trans. Graph., 26, 80, 10.1145/1276377.1276478 Garrido, 2016, Reconstruction of personalized 3d face rigs from monocular video, ACM Trans. Graph., 35, 28, 10.1145/2890493 Tkach, 2016, Sphere-meshes for real-time hand modeling and tracking, ACM Trans. Graph., 35, 222, 10.1145/2980179.2980226 Taylor, 2016, Efficient and precise interactive hand tracking through joint, continuous optimization of pose and correspondences, ACM Trans. Graph., 35, 143, 10.1145/2897824.2925965 Habermann, 2019, Livecap: Real-time human performance capture from monocular video, ACM Trans. Graph., 38, 14, 10.1145/3311970 Mehta, 2017, Vnect: Real-time 3d human pose estimation with a single rgb camera, ACM Trans. Graph., 36, 44, 10.1145/3072959.3073596 Saragih, 2009, Face alignment through subspace constrained mean-shifts, 1034 Guo, 2017, Robust non-rigid motion tracking and surface reconstruction using l_0 regularization, IEEE Trans. Vis. Comput. Graphics, 24, 1770, 10.1109/TVCG.2017.2688331 Weise, 2011, Realtime performance-based facial animation, 77 Li, 2013, Realtime facial animation with on-the-fly correctives, ACM Trans. Graph., 32, 10.1145/2461912.2462019 Helten, 2013, Personalization and evaluation of a real-time depth-based full body tracker, 279 F. Bogo, M.J. Black, M. Loper, J. Romero, Detailed full-body reconstructions of moving people from monocular rgb-d sequences, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2300–2308. T. Alldieck, M. Magnor, W. Xu, C. Theobalt, G. Pons-Moll, Video based reconstruction of 3d people models, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8387–8397. H. Joo, T. Simon, Y. Sheikh, Total capture: A 3d deformation model for tracking faces, hands, and bodies, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8320–8329. Laurentini, 1994, The visual hull concept for silhouette-based image understanding, IEEE Trans. Pattern Anal. Mach. Intell., 16, 150, 10.1109/34.273735 P. Fechteler, A. Hilsmann, P. Eisert, Markerless multiview motion capture with 3d shape model adaptation, in: Computer Graphics Forum, Wiley Online Library. Zhang, 2003, Spacetime stereo: Shape recovery for dynamic scenes, II Bleyer, 2011, Patchmatch stereo-stereo matching with slanted support windows, 1 Zienkiewicz, 2016, Monocular, real-time surface reconstruction using dynamic level of detail, 37 C.R. Qi, H. Su, K. Mo, L.J. Guibas, Pointnet: Deep learning on point sets for 3d classification and segmentation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 652–660. Riegler, 2017, Octnetfusion: Learning depth fusion from data, 57 J.J. Park, P. Florence, J. Straub, R. Newcombe, S. Lovegrove, Deepsdf: Learning continuous signed distance functions for shape representation, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 165–174. Wang, 2019, Dynamic graph cnn for learning on point clouds, ACM Trans. Graph., 38, 1, 10.1145/3326362 Q. Tan, L. Gao, Y.-K. Lai, S. Xia, Variational autoencoders for deforming 3d mesh models, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5841–5850. Gao, 2018, Automatic unpaired shape deformation transfer, ACM Trans. Graph., 37, 1, 10.1145/3197517.3201309 A. Dai, A.X. Chang, M. Savva, M. Halber, T. Funkhouser, M. Nießner, Scannet: Richly-annotated 3d reconstructions of indoor scenes, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5828–5839. Wang, 2018, Adversarial semantic scene completion from a single depth image, 426 Sun, 2020, Pointgrow: Autoregressively learned point cloud generation with self-attention, 61