Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation

Acta Numerica - Tập 30 - Trang 203-248 - 2021
Mikhail A. Belkin1
1Halıcıoğlu Data Science Institute, University of California San Diego, 10100 Hopkins Drive, La Jolla, CA 92093, USA E-mail:

Tóm tắt

In the past decade the mathematical theory of machine learning has lagged far behind the triumphs of deep neural networks on practical challenges. However, the gap between theory and practice is gradually starting to close. In this paper I will attempt to assemble some pieces of the remarkable and still incomplete mathematical mosaic emerging from the efforts to understand the foundations of deep learning. The two key themes will be interpolation and its sibling over-parametrization. Interpolation corresponds to fitting data, even noisy data, exactly. Over-parametrization enables interpolation and provides flexibility to select a suitable interpolating model.As we will see, just as a physical prism separates colours mixed within a ray of light, the figurative prism of interpolation helps to disentangle generalization and optimization properties within the complex picture of modern machine learning. This article is written in the belief and hope that clearer understanding of these issues will bring us a step closer towards a general theory of deep learning and machine learning.

Từ khóa


Tài liệu tham khảo

Oymak, 2019, Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 97, 4951

10.1006/jcss.1997.1504

10.1006/jmva.1997.1725

Nagarajan, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)

Mei, S. and Montanari, A. (2019), The generalization error of random features regression: Precise asymptotics and double descent curve. Available at arXiv:1908.05355.

Hui, L. and Belkin, M. (2021), Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks, in 9th International Conference on Learning Representations (ICLR 2021). Available at https://openreview.net/forum?id=hsFN92eQEla.

Du, S. S. , Zhai, X. , Poczos, B. and Singh, A. (2019b), Gradient descent provably optimizes over-parameterized neural networks, in 7th International Conference on Learning Representations (ICLR 2019). Available at https://openreview.net/forum?id=S1eK3i09YQ.

Ghorbani, 2020, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 14820

Jacot, 2018, Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 8571

10.1109/TEVC.2019.2890858

Mitra, P. P. (2019), Understanding overfitting peaks in generalization error: Analytical risk curves for l 2 and l 1 penalized interpolation. Available at arXiv:1906.03667.

10.1109/JSAIT.2020.2984716

Moulines, 2011, Advances in Neural Information Processing Systems 24 (NIPS 2011), 451

10.1214/aos/1176345451

10.1162/neco.1992.4.1.1

Nadaraya, 1964, On estimating regression, Theory Probab, Appl, 9, 141

Neyshabur, B. , Tomioka, R. and Srebro, N. (2015), In search of the real inductive bias: On the role of implicit regularization in deep learning, in ICLR (Workshop) 2015. Available at https://openreview.net/forum?id=6AzZb_7Qo0e.

Ji, 2019, Proceedings of the 32nd Conference on Learning Theory (COLT 2019), 99, 1772

Bartlett, 2021, Acta Numerica, 30, 1

Kaczmarz, 1937, Angenäherte Auflösung von Systemen linearer Gleichungen, Bull, Int. Acad. Sci. Pologne A, 35, 355

LeCun, Y. (2019), The epistemology of deep learning. Available at https://www.youtube.com/watch?v=gG5NCkMerHU&t=3210s.

10.1214/aoms/1177697089

Goodfellow, I. , Bengio, Y. and Courville, A. (2016), Deep Learning, MIT Press. Available at http://www.deeplearningbook.org.

Ma, S. and Belkin, M. (2019), Kernel machines that adapt to GPUs for effective large batch training, in Proceedings of Machine Learning and Systems ( Talwalkar, A. et al., eds), pp. 360–373. Available at https://proceedings.mlsys.org/paper/2019/file/a4a042cf4fd6bfb47701cbc8a1653ada-Paper.pdf.

Needell, 2014, Advances in Neural Information Processing Systems 27 (NIPS 2014), 1017

10.1038/s41586-019-1923-7

Karimi, H. , Nutini, J. and Schmidt, M. (2016), Linear convergence of gradient and proximal-gradient methods under the Polyak–Łojasiewicz condition, in Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2016) ( Frasconi, P. et al., eds), Vol. 9851 of Lecture Notes in Computer Science, Springer, pp. 795–811.

Schapire, 1998, Boosting the margin: A new explanation for the effectiveness of voting methods, Ann, Statist, 26, 1651

10.1137/20M1336072

Nichani, E. , Radhakrishnan, A. and Uhler, C. (2020), Do deeper convolutional networks perform better? Available at arXiv:2010.09610.

Muthukumar, V. , Narang, A. , Subramanian, V. , Belkin, M. , Hsu, D. and Sahai, A. (2020a), Classification vs regression in overparameterized regimes: Does the loss function matter? Available at arXiv:2005.08054.

10.1007/978-0-387-21606-5

Arora, S. , Du, S. S. , Li, Z. , Salakhutdinov, R. , Wang, R. and Yu, D. (2020), Harnessing the power of infinitely wide deep nets on small-data tasks, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=rkl8sJBYvH.

Xu, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)

Canziani, A. , Paszke, A. and Culurciello, E. (2016), An analysis of deep neural network models for practical applications. Available at arXiv:1605.07678.

Defazio, 2014, Advances in Neural Information Processing Systems 27 (NIPS 2014), 1646

Woodworth, 2020, Proceedings of the 33rd Conference on Learning Theory (COLT 2020), 125, 3635

Sindhwani, V. , Niyogi, P. and Belkin, M. (2005), Beyond the point cloud: From transductive to semi-supervised learning, in Proceedings of the 22nd International Conference on Machine Learning (ICML 2005), ACM, pp. 824–831.

Rahimi, 2007, Advances in Neural Information Processing Systems 20 (NIPS 2007), 1177

Lee, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 8570

Allen-Zhu, Z. and Li, Y. (2020), Backward feature correction: How deep learning performs deep learning. Available at arXiv:2001.04413.

10.1016/0020-0190(87)90114-1

Rahimi, A. and Recht, B. (2017), Reflections on random kitchen sinks. Available at http://www.argmin.net/2017/12/05/kitchen-sinks/.

10.1073/pnas.1903070116

Belkin, 2018a, Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 2306

10.1007/s00365-006-0663-2

10.1007/978-1-4757-2440-0

10.1214/07-STS242B

10.1109/TIT.1967.1053964

Belkin, 2019b, Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS 2019), 89, 1611

Shankar, 2020, Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 119, 8614

10.1007/b97848

10.1561/2200000050

10.1007/s00041-008-9030-4

Zhou, 2020, Advances in Neural Information Processing Systems 33 (Neur-IPS), 6867

Bassily, R. , Belkin, M. and Ma, S. (2018), On exponential convergence of SGD in non-convex over-parametrized learning. Available at arXiv:1811.02564.

10.1073/pnas.2001875117

Liu, 2020a, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 15954

Halton, J. H. (1991), Simplicial multivariable linear interpolation. Report TR91-002, Department of Computer Science, University of North Carolina at Chapel Hill.

Breiman, 1995, The Mathematics of Generalization, 11

Watson, 1964, Smooth regression analysis, Sankhyā A, 26, 359

Hastie, T. , Montanari, A. , Rosset, S. and Tibshirani, R. J. (2019), Surprises in high-dimensional ridgeless least squares interpolation. Available at arXiv:1903.08560.

10.1007/978-3-540-28650-9_8

10.1017/CBO9780511617539

Polyak, 1963, Gradient methods for minimizing functionals, Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki, 3, 643

Fawzi, 2016, Advances in Neural Information Processing Systems 29 (NIPS 2016), 1632

Roux, 2012, Advances in Neural Information Processing Systems 25 (NIPS 2012), 2663

Kingma, D. P. and Ba, J. (2015), Adam: A method for stochastic optimization, in 3rd International Conference on Learning Representations (ICLR 2015) ( Bengio, Y. and LeCun, Y. , eds).

Du, 2019a, Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 97, 1675

Li, 2020, Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS 2020), 108, 4313

Lojasiewicz, 1963, A topological property of real analytic subsets, Coll. du CNRS, Les Équations aux Dérivées Partielles, 117, 87

Salakhutdinov, R. (2017), Tutorial on deep learning. Available at https://simons.berkeley.edu/talks/ruslan-salakhutdinov-01-26-2017-1.

Liu, C. , Zhu, L. and Belkin, M. (2020b), Toward a theory of optimization for overparameterized systems of non-linear equations: The lessons of deep learning. Available at arXiv:2003.00307.

Mai, X. and Liao, Z. (2019), High dimensional classification via regularized and unreg-ularized empirical risk minimization: Precise error and optimal loss. Available at arXiv:1905.13742.

Liu, C. and Belkin, M. (2020), Accelerating SGD with momentum for over-parameterized learning, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=r1gixp4FPH.

Bai, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 690

10.1214/19-AOS1849

Fedus, W. , Zoph, B. and Shazeer, N. (2021), Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Available at arXiv:2101.03961.

10.1073/pnas.2005013117

Warmuth, 2005, International Conference on Computational Learning Theory (COLT 2005), 3559, 366

Bruna, J. , Szegedy, C. , Sutskever, I. , Goodfellow, I. , Zaremba, W. , Fergus, R. and Er-han, D. (2014), Intriguing properties of neural networks, in 2nd International Conference on Learning Representations (ICLR 2014). Available at https://openreview.net/forum?id=kklr_MTHMRQjG.

Bartlett, P. L. and Long, P. M. (2020), Failures of model-dependent generalization bounds for least-norm interpolation. Available at arXiv:2010.08479.

Madry, A. , Makelov, A. , Schmidt, L. , Tsipras, D. and Vladu, A. (2018), Towards deep learning models resistant to adversarial attacks, in 6th International Conference on Learning Representations (ICLR 2018). Available at https://openreview.net/forum?id=rJzIBfZAb.

Rifkin, R. M. (2002), Everything old is new again: A fresh look at historical approaches in machine learning. PhD thesis, Massachusetts Institute of Technology.

Ma, 2018, Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 80, 3325

10.1088/1751-8121/ab4c8b

Johnson, 2013, Advances in Neural Information Processing Systems 26 (NIPS 2013), 315

Nakkiran, P. , Kaplun, G. , Bansal, Y. , Yang, T. , Barak, B. and Sutskever, I. (2020), Deep double descent: Where bigger models and more data hurt, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=B1g5sA4twr.

Cutler, 2001, PERT: perfect random tree ensembles, Comput, Sci. Statist, 33, 490

Belkin, 2018b, Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 80, 541

Thrampoulidis, 2020, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 8907

Bartlett, 2002, Rademacher and Gaussian complexities: Risk bounds and structural results, J. Mach. Learn. Res, 3, 463

Lee, J. , Schoenholz, S. S. , Pennington, J. , Adlam, B. , Xiao, L. , Novak, R. and Sohl-Dickstein, J. (2020), Finite versus infinite neural networks: An empirical study. Available at arXiv:2007.15801.

Shepard, D. (1968), A two-dimensional interpolation function for irregularly-spaced data, in Proceedings of the 1968 23rd ACM National Conference, ACM, pp. 517–524.

Nocedal, 2006, Numerical Optimization

Meanti, G. , Carratino, L. , Rosasco, L. and Rudi, A. (2020), Kernel methods through the roof: Handling billions of points efficiently. Available at arXiv:2006.10350.

10.1073/pnas.1907378117

Zhang, C. , Bengio, S. , Hardt, M. , Recht, B. and Vinyals, O. (2017), Understanding deep learning requires rethinking generalization, in 5th International Conference on Learning Representations (ICLR 2017). Available at https://openreview.net/forum?id=Sy8gdB9xx.

Negrea, 2020, Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 119, 7263

Wyner, 2017, Explaining the success of AdaBoost and random forests as interpolating classifiers, J. Mach. Learn. Res, 18, 1

Pravesh, 2020, Algorithmic Learning Theory (ALT 2020), 117, 422

Ilyas, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)