Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation
Tóm tắt
Từ khóa
Tài liệu tham khảo
Oymak, 2019, Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 97, 4951
Nagarajan, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)
Mei, S. and Montanari, A. (2019), The generalization error of random features regression: Precise asymptotics and double descent curve. Available at arXiv:1908.05355.
Hui, L. and Belkin, M. (2021), Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks, in 9th International Conference on Learning Representations (ICLR 2021). Available at https://openreview.net/forum?id=hsFN92eQEla.
Du, S. S. , Zhai, X. , Poczos, B. and Singh, A. (2019b), Gradient descent provably optimizes over-parameterized neural networks, in 7th International Conference on Learning Representations (ICLR 2019). Available at https://openreview.net/forum?id=S1eK3i09YQ.
Ghorbani, 2020, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 14820
Jacot, 2018, Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 8571
Mitra, P. P. (2019), Understanding overfitting peaks in generalization error: Analytical risk curves for l 2 and l 1 penalized interpolation. Available at arXiv:1906.03667.
Moulines, 2011, Advances in Neural Information Processing Systems 24 (NIPS 2011), 451
Nadaraya, 1964, On estimating regression, Theory Probab, Appl, 9, 141
Neyshabur, B. , Tomioka, R. and Srebro, N. (2015), In search of the real inductive bias: On the role of implicit regularization in deep learning, in ICLR (Workshop) 2015. Available at https://openreview.net/forum?id=6AzZb_7Qo0e.
Ji, 2019, Proceedings of the 32nd Conference on Learning Theory (COLT 2019), 99, 1772
Bartlett, 2021, Acta Numerica, 30, 1
Kaczmarz, 1937, Angenäherte Auflösung von Systemen linearer Gleichungen, Bull, Int. Acad. Sci. Pologne A, 35, 355
LeCun, Y. (2019), The epistemology of deep learning. Available at https://www.youtube.com/watch?v=gG5NCkMerHU&t=3210s.
Goodfellow, I. , Bengio, Y. and Courville, A. (2016), Deep Learning, MIT Press. Available at http://www.deeplearningbook.org.
Ma, S. and Belkin, M. (2019), Kernel machines that adapt to GPUs for effective large batch training, in Proceedings of Machine Learning and Systems ( Talwalkar, A. et al., eds), pp. 360–373. Available at https://proceedings.mlsys.org/paper/2019/file/a4a042cf4fd6bfb47701cbc8a1653ada-Paper.pdf.
Needell, 2014, Advances in Neural Information Processing Systems 27 (NIPS 2014), 1017
Karimi, H. , Nutini, J. and Schmidt, M. (2016), Linear convergence of gradient and proximal-gradient methods under the Polyak–Łojasiewicz condition, in Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2016) ( Frasconi, P. et al., eds), Vol. 9851 of Lecture Notes in Computer Science, Springer, pp. 795–811.
Schapire, 1998, Boosting the margin: A new explanation for the effectiveness of voting methods, Ann, Statist, 26, 1651
Nichani, E. , Radhakrishnan, A. and Uhler, C. (2020), Do deeper convolutional networks perform better? Available at arXiv:2010.09610.
Muthukumar, V. , Narang, A. , Subramanian, V. , Belkin, M. , Hsu, D. and Sahai, A. (2020a), Classification vs regression in overparameterized regimes: Does the loss function matter? Available at arXiv:2005.08054.
Arora, S. , Du, S. S. , Li, Z. , Salakhutdinov, R. , Wang, R. and Yu, D. (2020), Harnessing the power of infinitely wide deep nets on small-data tasks, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=rkl8sJBYvH.
Xu, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)
Canziani, A. , Paszke, A. and Culurciello, E. (2016), An analysis of deep neural network models for practical applications. Available at arXiv:1605.07678.
Defazio, 2014, Advances in Neural Information Processing Systems 27 (NIPS 2014), 1646
Woodworth, 2020, Proceedings of the 33rd Conference on Learning Theory (COLT 2020), 125, 3635
Sindhwani, V. , Niyogi, P. and Belkin, M. (2005), Beyond the point cloud: From transductive to semi-supervised learning, in Proceedings of the 22nd International Conference on Machine Learning (ICML 2005), ACM, pp. 824–831.
Rahimi, 2007, Advances in Neural Information Processing Systems 20 (NIPS 2007), 1177
Lee, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 8570
Allen-Zhu, Z. and Li, Y. (2020), Backward feature correction: How deep learning performs deep learning. Available at arXiv:2001.04413.
Rahimi, A. and Recht, B. (2017), Reflections on random kitchen sinks. Available at http://www.argmin.net/2017/12/05/kitchen-sinks/.
Belkin, 2018a, Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 2306
Belkin, 2019b, Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS 2019), 89, 1611
Shankar, 2020, Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 119, 8614
Zhou, 2020, Advances in Neural Information Processing Systems 33 (Neur-IPS), 6867
Bassily, R. , Belkin, M. and Ma, S. (2018), On exponential convergence of SGD in non-convex over-parametrized learning. Available at arXiv:1811.02564.
Liu, 2020a, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 15954
Halton, J. H. (1991), Simplicial multivariable linear interpolation. Report TR91-002, Department of Computer Science, University of North Carolina at Chapel Hill.
Breiman, 1995, The Mathematics of Generalization, 11
Watson, 1964, Smooth regression analysis, Sankhyā A, 26, 359
Hastie, T. , Montanari, A. , Rosset, S. and Tibshirani, R. J. (2019), Surprises in high-dimensional ridgeless least squares interpolation. Available at arXiv:1903.08560.
Polyak, 1963, Gradient methods for minimizing functionals, Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki, 3, 643
Fawzi, 2016, Advances in Neural Information Processing Systems 29 (NIPS 2016), 1632
Roux, 2012, Advances in Neural Information Processing Systems 25 (NIPS 2012), 2663
Kingma, D. P. and Ba, J. (2015), Adam: A method for stochastic optimization, in 3rd International Conference on Learning Representations (ICLR 2015) ( Bengio, Y. and LeCun, Y. , eds).
Du, 2019a, Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 97, 1675
Li, 2020, Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS 2020), 108, 4313
Lojasiewicz, 1963, A topological property of real analytic subsets, Coll. du CNRS, Les Équations aux Dérivées Partielles, 117, 87
Salakhutdinov, R. (2017), Tutorial on deep learning. Available at https://simons.berkeley.edu/talks/ruslan-salakhutdinov-01-26-2017-1.
Liu, C. , Zhu, L. and Belkin, M. (2020b), Toward a theory of optimization for overparameterized systems of non-linear equations: The lessons of deep learning. Available at arXiv:2003.00307.
Mai, X. and Liao, Z. (2019), High dimensional classification via regularized and unreg-ularized empirical risk minimization: Precise error and optimal loss. Available at arXiv:1905.13742.
Liu, C. and Belkin, M. (2020), Accelerating SGD with momentum for over-parameterized learning, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=r1gixp4FPH.
Bai, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019), 690
Fedus, W. , Zoph, B. and Shazeer, N. (2021), Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Available at arXiv:2101.03961.
Warmuth, 2005, International Conference on Computational Learning Theory (COLT 2005), 3559, 366
Bruna, J. , Szegedy, C. , Sutskever, I. , Goodfellow, I. , Zaremba, W. , Fergus, R. and Er-han, D. (2014), Intriguing properties of neural networks, in 2nd International Conference on Learning Representations (ICLR 2014). Available at https://openreview.net/forum?id=kklr_MTHMRQjG.
Bartlett, P. L. and Long, P. M. (2020), Failures of model-dependent generalization bounds for least-norm interpolation. Available at arXiv:2010.08479.
Madry, A. , Makelov, A. , Schmidt, L. , Tsipras, D. and Vladu, A. (2018), Towards deep learning models resistant to adversarial attacks, in 6th International Conference on Learning Representations (ICLR 2018). Available at https://openreview.net/forum?id=rJzIBfZAb.
Rifkin, R. M. (2002), Everything old is new again: A fresh look at historical approaches in machine learning. PhD thesis, Massachusetts Institute of Technology.
Ma, 2018, Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 80, 3325
Johnson, 2013, Advances in Neural Information Processing Systems 26 (NIPS 2013), 315
Nakkiran, P. , Kaplun, G. , Bansal, Y. , Yang, T. , Barak, B. and Sutskever, I. (2020), Deep double descent: Where bigger models and more data hurt, in 8th International Conference on Learning Representations (ICLR 2020). Available at https://openreview.net/forum?id=B1g5sA4twr.
Cutler, 2001, PERT: perfect random tree ensembles, Comput, Sci. Statist, 33, 490
Belkin, 2018b, Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 80, 541
Thrampoulidis, 2020, Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 8907
Bartlett, 2002, Rademacher and Gaussian complexities: Risk bounds and structural results, J. Mach. Learn. Res, 3, 463
Lee, J. , Schoenholz, S. S. , Pennington, J. , Adlam, B. , Xiao, L. , Novak, R. and Sohl-Dickstein, J. (2020), Finite versus infinite neural networks: An empirical study. Available at arXiv:2007.15801.
Shepard, D. (1968), A two-dimensional interpolation function for irregularly-spaced data, in Proceedings of the 1968 23rd ACM National Conference, ACM, pp. 517–524.
Nocedal, 2006, Numerical Optimization
Meanti, G. , Carratino, L. , Rosasco, L. and Rudi, A. (2020), Kernel methods through the roof: Handling billions of points efficiently. Available at arXiv:2006.10350.
Zhang, C. , Bengio, S. , Hardt, M. , Recht, B. and Vinyals, O. (2017), Understanding deep learning requires rethinking generalization, in 5th International Conference on Learning Representations (ICLR 2017). Available at https://openreview.net/forum?id=Sy8gdB9xx.
Negrea, 2020, Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 119, 7263
Wyner, 2017, Explaining the success of AdaBoost and random forests as interpolating classifiers, J. Mach. Learn. Res, 18, 1
Pravesh, 2020, Algorithmic Learning Theory (ALT 2020), 117, 422
Ilyas, 2019, Advances in Neural Information Processing Systems 32 (NeurIPS 2019)
