Một khảo sát về các thuật toán tối ưu siêu tham số đa mục tiêu cho học máy

Artificial Intelligence Review - Tập 56 - Trang 8043-8093 - 2022
Alejandro Morales-Hernández1,2, Inneke Van Nieuwenhuyse1,2, Sebastian Rojas Gonzalez1,2,3
1Faculty of Sciences, Hasselt University, Hasselt, Belgium
2VCCM Core Lab and Data Science Institute, Hasselt University, Hasselt, Belgium
3Surrogate Modeling Lab, Ghent University, Ghent, Belgium

Tóm tắt

Tối ưu siêu tham số (HPO) là bước cần thiết để đảm bảo hiệu suất tốt nhất có thể cho các thuật toán Học máy (ML). Nhiều phương pháp đã được phát triển để thực hiện HPO; hầu hết tập trung vào việc tối ưu hóa một thước đo hiệu suất (thường là thước đo dựa trên lỗi), và tài liệu về những vấn đề HPO đơn mục tiêu như vậy là rất phong phú. Gần đây, tuy nhiên, đã xuất hiện các thuật toán tập trung vào việc tối ưu hóa nhiều mục tiêu mâu thuẫn cùng một lúc. Bài báo này trình bày một khảo sát hệ thống về tài liệu đã được xuất bản từ năm 2014 đến 2020 về các thuật toán HPO đa mục tiêu, phân biệt các thuật toán dựa trên metaheuristic, các thuật toán dựa trên mô hình meta và các phương pháp sử dụng sự kết hợp của cả hai. Chúng tôi cũng thảo luận về các chỉ số chất lượng được sử dụng để so sánh các quy trình HPO đa mục tiêu và trình bày các hướng nghiên cứu trong tương lai.

Từ khóa

#tối ưu siêu tham số #học máy #thuật toán đa mục tiêu #metaheuristic #mô hình meta

Tài liệu tham khảo

Abdolsh M, Shilton A, Rana S, Gupta S, Venkatesh S (2019) Multi-objective Bayesian optimisation with preferences over objectives. Advances in neural information processing systems pp 12235–12245 Ab Wahab MN, Nefti-Meziani S, Atyabi A (2015) A comprehensive review of swarm optimization algorithms. PLoS ONE 10(5):e0122827 Akiba T, Sano S, Yanase T, Ohta T, Koyama M (2019) Optuna: a next-generation hyperparameter optimization framework. In: Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. pp 2623–2631 Alaya I, Solnon C, Ghedira K (2007) Ant colony optimization for multi-objective optimization problems. In: 19th IEEE international conference on tools with artificial intelligence (ICTAI 2007) vol 1, pp 450–457. https://doi.org/10.1109/ICTAI.2007.108 Albelwi S, Mah A (2016) Automated optimal architecture of deep convolutional neural networks for image recognition. In: 2016 15th IEEE international conference on machine learning and applications (icmla) pp 53–60. https://doi.org/10.1109/ICMLA.2016.0018 Andreopoulos A, Tsotsos JK (2013) 50 years of object recognition: directions forward. Comput Vis Image Understand 117(8):827–891. https://doi.org/10.1016/j.cviu.2013.04.005 Babu B, Gujarathi AM (2007) Multi-objective differential evolution (mode) algorithm for multi-objective optimization: parametric study on benchmark test problems. J Future Eng Technol 3(1):47–59. https://doi.org/10.26634/jfet.3.1.697 Baldeon M, Lai-Yuen SK (2020) Adaresu-net: multiobjective adaptive convolutional neural network for medical image segmentation. Neu-rocomputing 392:325–340. https://doi.org/10.1016/j.neucom.2019.01.110 Barán B, Schaerer M (2003) A multiobjective ant colony system for vehicle routing problem with time windows. Applied informatics pp 97–102 Belgiu M, Drăguţ L (2016) Random forest in remote sensing: a review of applications and future directions. ISPRS J Photogramm Remote Sens 114:24–31 Bergstra J, Bardenet R, Bengio Y, Kégl B (2011) Algorithms for hyper-parameter optimization. In: 25th annual conference on neural information processing systems (NIPS 2011) vol 24 Bergstra J, Bengio Y (2012) Random search for hyper-parameter optimization. J Mach Learn Res 13(1):281–305 Bergstra J, Yamins D, Cox D (2013) Making a science of model search: hyperparameter optimization in hundreds of dimensions for vision archi-tectures. In: International conference on machine learning, pp 115–123. http://proceedings.mlr.press/v28/bergstra13.html Bianchi L, Dorigo M, Gambardella LM, Gutjahr WJ (2009) A survey on metaheuristics for stochastic combinatorial optimization. Natural Comput 8(2):239–287. https://doi.org/10.1007/s11047-008-9098-4 Binder M, Moosbauer J, Thomas J, Bischl B (2020) Multi-objective hyperparameter tuning and feature selection using filter ensembles. vol 1050, p 13. https://doi.org/10.1145/3377930.3389815 Binois M, Huang J, Gramacy RB, Ludkovski M (2019) Replication or exploration? Sequential design for stochastic simulation experiments. Technometrics 61(1):7–23. https://doi.org/10.1080/00401706.2018.1469433 Bischl B, Mersmann O, Trautmann H, Weihs C (2012) Resampling methods for meta-model validation with recommendations for evolutionary computation. Evol Comput 20(2):249–275. https://doi.org/10.1162/EVCO_a_00069 Blumer A, Ehrenfeucht A, Haussler D, Warmuth MK (1987) Occam’s razor. Inf Process Lett 24(6):377–380. https://doi.org/10.1016/0020-0190(87)90114-1 Bouraoui A, Jamoussi S, BenAyed Y (2018) A multi-objective genetic algorithm for simultaneous model and feature selection for support vector machines. Artif Intell Rev 50(2):261–281. https://doi.org/10.1007/s10462-017-9543-9 Bui K-HN, Yi H (2020) Optimal hyperparameter tuning using meta-learning for big traffic datasets. In: Lee W et al. (ed) 2020 IEEE international conference on big data and smart computing (bigcomp 2020) pp 48–54. IEEE. https://doi.org/10.1109/BigComp48618.2020.0-100 Cai X, Hu Z, Zhao P, Zhang W, Chen J (2020) A hybrid recommendation system with many-objective evolutionary algorithm. Expert Syst Appl 159:113648. https://doi.org/10.1016/j.eswa.2020.113648 Calisto MB, Lai-Yuen SK (2020) Adaen-net: an ensemble of adaptive 2d–3d fully convolutional networks for medical image segmentation. Neural Netw. https://doi.org/10.1016/j.neunet.2020.03.007 Calisto MB, Lai-Yuen SK (2021) Emonas-net: efficient multiobjective neural architecture search using surrogate-assisted evolutionary algorithm for 3d medical image segmentation. Artif Intell Med 119:102154. https://doi.org/10.1016/j.artmed.2021.102154 Chandra A, Lane I (2016) Automated optimization of decoder hyper-parameters for online lvcsr. In: 2016 IEEE spoken language technol-ogy workshop (slt) pp 454–460. https://doi.org/10.1109/SLT.2016.7846303 Chatelain C, Adam S, Lecourtier Y, Heutte L, Paquet T (2007) Multi-objective optimization for svm model selection. In: Ninth international conference on document analysis and recognition (ICDAR 2007) vol 1, pp 427–431. https://doi.org/10.1109/ICDAR.2007.4378745 Chen J, Li K, Tang Z, Bilal K, Yu S, Weng C, Li K (2016) A parallel random forest algorithm for big data in a spark cloud computing environment. IEEE Trans Parallel Distrib Syst 28(4):919–933 Chen W-C, Jiang X-Y, Chang H-P, Chen H-P (2014) An effective system for parameter optimization in photolithography process of a lgp stamper. Neural Comput Appl 24(6):1391–1401. https://doi.org/10.1007/s00521-013-1353-7 Chin T-W, Morcos AS, Marculescu D (2020) Pareco: Pareto-aware channel optimization for slimmable neural networks. In: 2nd Workshop on Adversarial Learning Methods for Machine Learning and Data Mining, KDD’2020. https://openreview.net/forum?id=SPyxaz%5Fh9Nd Cho H, Kim Y, Lee E, Choi D, Lee Y, Rhee W (2020) Basic enhancement strategies when using Bayesian optimization for hyperparameter tuning of deep neural networks. IEEE Access 8:52588–52608. https://doi.org/10.1109/ACCESS.2020.2981072 Cooney C, Korik A, Folli R, Coyle D (2020) Evaluation of hyperparameter optimization in machine and deep learning methods for decoding imagined speech EEG. Sensors 20(16):4629. https://doi.org/10.3390/s20164629 Cowen-Rivers AI, Lyu W, Wang Z, Tutunov R, Jianye H, Wang J, Ammar HB (2020) Hebo: heteroscedastic evolutionary bayesian optimisation. Workshop at NeurIPS 2020 Competition Track on Black-Box Optimization Challenge Dai Z, Damianou A, Hensman J, Lawrence N (2014) Gaussian process models with parallelization and GPU acceleration. arXiv:1410.4984 Dai Z, Yu H, Low BKH, Jaillet P (2019) Bayesian optimization meets bayesian optimal stopping. International conference on machine learning. pp 1496–1506. http://proceedings.mlr.press/v97/dai19a.html Deb K, Pratap A, Agarwal S, Meyarivan T (2002) A fast and elitist multiobjective genetic algorithm: Nsga-ii. IEEE Trans Evol Comput 6(2):182–197. https://doi.org/10.1109/4235.996017 Deighan DS, Field SE, Capano CD, Khanna G (2021) Genetic-algorithm-optimized neural networks for gravitational wave classification. Neural Comput Appl. https://doi.org/10.1007/s00521-021-06024-4 de Toro F, Ros E, Mota S, Ortega J (2002) Multi-objective optimization evolutionary algorithms applied to paroxysmal atrial fibrillation diagnosis based on the k-nearest neighbours classifier. Ibero-american conference on artificial intelligence pp 313–318. https://doi.org/10.1007/3-540-36131-6_32 Doerner K, Gutjahr WJ, Hartl RF, Strauss C, Stummer C (2004) Pareto ant colony optimization: a metaheuristic approach to multiobjective portfolio selection. Ann Oper Res 131(1):79–99. https://doi.org/10.1023/B:ANOR.0000039513.99038.c6 Doerr C (2020) Complexity theory for discrete black-box optimization heuristics. Theory of evolutionary computation. Springer pp 133–212 Dorigo M, Blum C (2005) Ant colony optimization theory: a survey. Theor Comput Sci 344(2–3):243–278. https://doi.org/10.1016/j.tcs.2005.05.020 Dorigo M, Maniezzo V, Colorni A (1996) Ant system: optimization by a colony of cooperating agents. IEEE Trans Syst Man Cybern B 26(1):29–41. https://doi.org/10.1109/3477.484436 Durillo JJ, Nebro AJ, Luna F, Alba E (2008) A study of master-slave approaches to parallelize nsga-ii. In: 2008 IEEE international symposium on parallel and distributed processing pp 1–8 Dutta S, Gandomi AH (2020) Surrogate model-driven evolutionary algorithms: theory and applications. Evolution in action: past, present and future. Springer. pp 435–451. https://doi.org/10.1007/978-3-030-39831-6_29 Eberhart R, Kennedy J (1995) A new optimizer using particle swarm theory. Mhs’95. In: Proceedings of the sixth international symposium on micro machine and human science. pp 39–43. https://doi.org/10.1109/MHS.1995.494215 Ekbal A, Saha S (2015) Joint model for feature selection and parameter optimization coupled with classifier ensemble in chemical mention recog-nition. Knowl-Based Syst 85:37–51. https://doi.org/10.1016/j.knosys.2015.04.015 Ekbal A, Saha S (2016) Simultaneous feature and parameter selection using multiobjective optimization: application to named entity recogni-tion. Int J Mach Learn Cybern 7(4):597–611. https://doi.org/10.1007/s13042-014-0268-7 Emmerich MT, Deutz AH (2018) A tutorial on multiobjective optimization: fundamentals and evolutionary methods. Natural Comput 17(3):585–609 Emmerich MT, Deutz AH, Klinkenberg JW (2011) Hypervolume-based expected improvement: Monotonicity properties and exact computation. In: 2011 IEEE congress of evolutionary computation (CEC) pp 2147–2154. https://doi.org/10.1109/CEC.2011.5949880 Ertel W (2018) Introduction to artificial intelligence. Springer, Cham Falkner S, Klein A, Hutter F (2018) Bohb: robust and efficient hyper-parameter optimization at scale. In: International conference on machine learning pp 1437–1446. http://proceedings.mlr.press/v80/falkner18a.html Faris H, Habib M, Faris M, Alomari M, Alomari A (2020) Medical speciality classification system based on binary particle swarms and ensemble of one vs rest support vector machines. J Biomed Inform 109:103525. https://doi.org/10.1016/j.jbi.2020.103525 Feliot P, Bect J, Vazquez E (2017) A bayesian approach to constrained single-and multi-objective optimization. J Glob Optim 67(1–2):97–133. https://doi.org/10.1007/s10898-016-0427-3 Feurer M, Hutter F (2019) Hyperparameter optimization. Automated machine learning: methods, systems, challenges. Springer, Cham, pp 3–33 Fieldsend JE, Everson RM (2014) The rolling tide evolutionary algorithm: a multiobjective optimizer for noisy optimization problems. IEEE Trans Evol Comput 19(1):103–117. https://doi.org/10.1109/TEVC.2014.2304415 Forrester AI, Keane AJ, Bressloff NW (2006) Design and analysis of “noisy’’ computer experiments. AIAA J 44(10):2331–2339. https://doi.org/10.2514/1.20068 Gad AG (2022) Particle swarm optimization algorithm and its applications: a systematic review. Arch Comput Methods Eng 8:1–31 Garrido EC, Hernández D (2019) Predictive entropy search for multi-objective Bayesian optimization with constraints. Neurocomputing 361:50–68. https://doi.org/10.1016/j.neucom.2019.06.025 Glover F (1986) Future paths for integer programming and links to artificial intelligence. Comput Oper Res 13(5):533–549. https://doi.org/10.1016/0305-0548(86)90048-1 Gonzalez SR, Jalali H, Van Nieuwenhuyse I (2020) A multiobjective stochastic simulation optimization algorithm. Eur J Oper Res 284(1):212–226. https://doi.org/10.1016/j.ejor.2019.12.014 Gravel M, Price WL, Gagné C (2002) Scheduling continuous casting of aluminum using a multiple objective ant colony optimization meta-heuristic. Eur J Oper Res 143(1):218–229. https://doi.org/10.1016/S0377-2217(01)00329-0 Gülcü A, Kuş Z (2021) Multi-objective simulated annealing for hyper-parameter optimization in convolutional neural networks. PeerJ Comput Sci 7:e338. https://doi.org/10.7717/peerj-cs.338 Guo C, Li L, Hu Y, Yan J (2020) A deep learning based fault diagnosis method with hyperparameter optimization by using parallel computing. IEEE Access 8:131248–131256. https://doi.org/10.1109/ACCESS.2020.3009644 Guo J, Yang L, Bie R, Yu J, Gao Y, Shen Y, Kos A (2019) An xgboost-based physical fitness evaluation model using advanced feature selection and bayesian hyper-parameter optimization for wearable running monitoring. Comput Netw 151:166–180. https://doi.org/10.1016/j.comnet.2019.01.026 Gupta S, Shilton A, Rana S, Venkatesh S (2018) Exploiting strategy-space diversity for batch bayesian optimization. In: International conference on artificial intelligence and statistics pp 538–547. http://proceedings.mlr.press/v84/gupta18a.html Han S, Pool J, Tran J, Dally W (2015) Learning both weights and connections for efficient neural network. In: Cortes C, Lawrence N, Lee D, Sugiyama M, Garnett R (eds) Advances in neural information processing systems vol 28, pp 1135-1143. Curran Associates, Inc. https://dl.acm.org/doi/10.5555/2969239.2969366 Hansen N, Müller SD, Koumoutsakos P (2003) Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (CMA-ES). Evol Comput 11(1):1–18. https://doi.org/10.1162/106365603321828970 Hegde S, Mundada MR (2020) Early prediction of chronic disease using an efficient machine learning algorithm through adaptive probabilistic divergence based feature selection approach. Int J Pervas Comput Commun. https://doi.org/10.1108/IJPCC-04-2020-0018 Hernández D, Hernandez-Lobato J, Shah A, Adams R (2016) Predictive entropy search for multi-objective bayesian optimization. In: International conference on machine learning pp 1492–1501. http://proceedings.mlr.press/v48/hernandez-lobatoa16.html Hernández-Lobato JM, Gelbart MA, Reagen B, Adolf R, Hernández-Lobato D, Whatmough PN, Adams RP (2016) Designing neural network hardware accelerators with decoupled objective evaluations. Nips workshop on bayesian optimization. p 10 Ho TK (1995) Random decision forests. In: Proceedings of 3rd international conference on document analysis and recognition vol 1, pp 278–282. https://doi.org/10.1109/ICDAR.1995.598994 Horn D, Bischl B (2016) Multi-objective parameter configuration of machine learning algorithms using model-based optimization. In: 2016 IEEE symposium series on computational intelligence (SSCI) pp 1–8. https://doi.org/10.1109/SSCI.2016.7850221 Horn D, Dagge M, Sun X, Bischl B (2017) First investigations on noisy model-based multi-objective optimization. International conference on evolutionary multi-criterion optimization pp 298–313. https://doi.org/10.1007/978-3-319-54157-0_21 Hsu C-H, Juang C-F (2013) Multi-objective continuous-ant-colony-optimized fc for robot wall-following control. IEEE Comput Intell Mag 8(3):28–40. https://doi.org/10.1109/MCI.2013.2264233 Hu W, Jin J, Liu T-Y, Zhang C (2019) Automatically design convolutional neural networks by optimization with submodularity and supermodularity. IEEE Trans Neural Netw Learn Syst. https://doi.org/10.1109/TNNLS.2019.2939157 Hutter F, Hoos HH, Leyton-Brown K, Stützle T (2009) Paramils: an automatic algorithm configuration framework. J Artif Intell Res 36:267–306. https://doi.org/10.1613/jair.2861 Hutter F, Lücke J, Schmidt-Thieme L (2015) Beyond manual tuning of hyperparameters. KI-Künstliche Intelligenz 29(4):329–337. https://doi.org/10.1007/s13218-015-0381-0 Iredi S, Merkle D, Middendorf M (2001) Bi-criterion optimization with multi colony ant algorithms. International conference on evolutionary multi-criterion optimization pp 359–372. https://doi.org/10.1007/3-540-44719-9_25 Jaderberg M, Dalibard V, Osindero S, Czarnecki WM, Donahue J, Razavi A, et al (2017) Population based training of neural networks. arXiv:1711.09846 Jalali H, Van Nieuwenhuyse I, Picheny V (2017) Comparison of kriging-based algorithms for simulation optimization with heterogeneous noise. Eur J Oper Res 261(1):279–301. https://doi.org/10.1016/j.ejor.2017.01.035 Jiang F, Jiang Y, Zhi H, Dong Y, Li H, Ma S, Wang Y (2017) Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol 2(4):230–243. https://doi.org/10.1136/svn-2017-000101 Jin Y (2011) Surrogate-assisted evolutionary computation: recent advances and future challenges. Swarm Evol Comput 1(2):61–70. https://doi.org/10.1016/j.swevo.2011.05.001 Jing W, Lin J, Wang H (2020) Building nas: Automatic designation of efficient neural architectures for building extraction in high-resolution aerial images. IMAGE AND VISION COMPUTING 103. https://doi.org/10.1016/j.imavis.2020.104025 Jomaa HS, Grabocka J, Schmidt-Thieme L (2019) Hyp-rl: Hyper-parameter optimization by reinforcement learning. arXiv:1906.11527 Jones DR, Schonlau M, Welch WJ (1998) Efficient global optimization of expensive black-box functions. J Global Optim 13(4):455–492. https://doi.org/10.1023/A:1008306431147 Joorabian M, Afzalan E (2014) Optimal power flow under both normal and contingent operation conditions using the hybrid fuzzy particle swarm optimisation and nelder-mead algorithm (hfpso-nm). Appl Soft Comput 14:623–633 Juang C-F (2002) A tsk-type recurrent fuzzy network for dynamic systems processing by neural network and genetic algorithms. IEEE Trans Fuzzy Syst 10(2):155–170. https://doi.org/10.1109/91.995118 Juang C-F, Hsu C-H (2014) Structure and parameter optimization of fnns using multi-objective aco for control and prediction. In: 2014 ieee international conference on fuzzy systems (fuzz-ieee) pp 928–933. https://doi.org/10.1109/FUZZ-IEEE.2014.6891545 Karnin Z, Koren T, Somekh O (2013) Almost optimal exploration in multi-armed bandits. International conference on machine learning pp 1238–1246 Kim Y, Reddy B, Yun S, Seo C (2017) Nemo: Neuro-evolution with mul-tiobjective optimization of deep neural network for speed and accuracy. Icml 2017 automl workshop Kirkpatrick S, Gelatt CD, Vecchi MP (1983) Optimization by simulated annealing. Science 220(4598):671–680. https://doi.org/10.1126/science.220.4598.671 Knowles J (2006) Parego: a hybrid algorithm with on-line landscape approximation for expensive multiobjective optimization problems. IEEE Trans Evol Comput 10(1):50–66. https://doi.org/10.1109/TEVC.2005.851274 Koch P, Wagner T, Emmerich MT, Bäck T, Konen W (2015) Efficient multi-criteria optimization on noisy machine learning problems. Appl Soft Comput 29:357–370. https://doi.org/10.1016/j.asoc.2015.01.005 Kohavi R, John GH (1995) Automatic parameter selection by minimizing estimated error. Machine learning proceedings 1995 pp 304–312. Elsevier. https://doi.org/10.1016/B978-1-55860-377-6.50045-1 Kong W, Dong ZY, Luo F, Meng K, Zhang W, Wang F, Zhao X (2017) Effect of automatic hyperparameter tuning for residential load forecasting via deep learning. 2017 australasian universities power engineering conference (aupec) (pp 1–6). https://doi.org/10.1109/AUPEC.2017.8282478 Kuncheva LI (2014) Combining pattern classifiers: methods and algorithms. Wiley, Hoboken Laskaridis S, Venieris SI, Kim H, Lane ND (2020) Hapi: hardware-aware progressive inference. In: 2020 ieee/acm international conference on computer aided design (iccad) (pp 1–9) Laumanns M, Thiele L, Deb K, Zitzler E (2002) Combining convergence and diversity in evolutionary multiobjective optimization. Evol Comput 10(3):263–282. https://doi.org/10.1162/106365602760234108 León J, Ortega J, Ortiz A (2019) Convolutional neural networks and feature selection for bci with multiresolution analysis. International work-conference on artificial neural networks (pp 883-894). https://doi.org/10.1007/978-3-030-20521-8_72 Li M, Yao X (2019) Quality evaluation of solution sets in multiobjective optimisation: a survey. ACM Comput Surv (CSUR) 52(2):1–38. https://doi.org/10.1145/3300148 Li H, Zhang Q, Tsang E, Ford JA (2004) Hybrid estimation of distribution algorithm for multiobjective knapsack problem. J. Gottlieb & G.R. Raidl (Eds.), Evolutionary computation in combinatorial optimization (pp 145–154). https://doi.org/10.1007/978-3-540-24652-7_15 Li L, Jamieson K, DeSalvo G, Rostamizadeh A, Talwalkar A (2017) Hyperband: a novel bandit-based approach to hyperparameter optimization. J Mach Learn Res 18(1):6765–6816 Li S, Gong W, Yan X, Hu C, Bai D, Wang L (2019) Parameter estimation of photovoltaic models with memetic adaptive differential evolution. Sol Energy 190:465–474. https://doi.org/10.1016/j.solener.2019.08.022 Liang J, Meyerson E, Hodjat B, Fink D, Mutch K, Miikkulainen R (2019) Evolutionary neural automl for deep learning. Proceedings of the genetic and evolutionary computation conference (pp 401–409). https://doi.org/10.1145/3321707.3321721 Liu H, Cai J, Ong Y-S (2018) Remarks on multi-output gaussian process regression. Knowl-Based Syst 144:102–121. https://doi.org/10.1016/j.knosys.2017.12.034 Liu J, Tunguz B, Titericz G (2020) GPU accelerated exhaustive search for optimal ensemble of black-box optimization algorithms. Workshop at NeurIPS 2020 Competition Track on Black-Box Optimization Challenge Loni M, Zoljodi A, Sinaei S, Daneshtalab M, Sjödin M (2019) Neuropower: Designing energy efficient convolutional neural network architecture for embedded systems. International conference on artificial neural networks (pp 208-222). https://doi.org/10.1007/978-3-030-30487-4_17 Loni M, Sinaei S, Zoljodi A, Daneshtalab M, Sjödin M (2020) Deepmaker: a multi-objective optimization framework for deep neural networks in embedded systems. Microprocess Microsyst 73:102989. https://doi.org/10.1016/j.micpro.2020.102989 López-Ibáñez M, Dubois-Lacoste J, Cáceres LP, Birattari M, Stützle T (2016) The irace package: iterated racing for automatic algorithm configuration. Oper Res Perspect 3:43–58. https://doi.org/10.1016/j.orp.2016.09.002 Lu Z, Whalen I, Boddeti V, Dhebar Y, Deb K, Goodman E, Banzhaf W (2019) Nsga-net: neural architecture search using multi-objective ge-netic algorithm. Proceedings of the genetic and evolutionary computation conference (pp 419–427). https://doi.org/10.1145/3321707.3321729 Lu Z, Deb K, Goodman E, Banzhaf W, Boddeti VN (2020) Nsganetv2: Evolutionary multi-objective surrogate-assisted neural architecture search. European conference on computer vision (pp 35–51). https://doi.org/10.1007/978-3-030-58452-8_3 Luo G (2016) A review of automatic selection methods for machine learning algorithms and hyper-parameter values. Netw Model Anal Health Inform Bioinform 5(1):18. https://doi.org/10.1007/s13721-016-0125-6 Magda M, Martinez-Alvarez A, Cuenca-Asensi S (2017) Mooga parameter optimization for onset detection in emg signals. International conference on image analysis and processing (pp 171–180). https://doi.org/10.1007/978-3-319-70742-6_16 Makarova A, Shen H, Perrone V, Klein A, Faddoul JB, Krause A, Archambeau C (2021) Overfitting in bayesian optimization: an empirical study and early-stopping solution. https://www.amazon.science/publications/overfitting-in-bayesian-optimization-an-empirical-study-and-early-stopping-solution Martinez-de Pison FJ, Gonzalez-Sendino R, Aldama A, Ferreiro J, Fraile E (2017) Hybrid methodology based on bayesian optimization and ga-parsimony for searching parsimony models by combining hyperparameter optimization and feature selection. International conference on hybrid artificial intelligence systems (pp 52-62). https://doi.org/10.1016/j.neucom.2018.05.136 McKinnon KI (1998) Convergence of the nelder-mead simplex method to a nonstationary point. SIAM J Optim 9(1):148–158. https://doi.org/10.1137/S1052623496303482 Mei J, Li Y, Lian X, Jin X, Yang L, Yuille A, Yang J (2020) Atomnas: Fine-grained end-to-end neural architecture search. International conference on learning representations. https://openreview.net/forum?id=BylQSxHFwr Meinshausen N, Ridgeway G (2006) Quantile regression forests. J Mach Learn Res 7(6). http://jmlr.org/papers/v7/meinshausen06a.html Mentch L, Hooker G (2016) Quantifying uncertainty in random forests via confidence intervals and hypothesis tests. J Mach Learn Res 17(1):841–881 Miettinen K (2012) Nonlinear multiobjective optimization. Springer, Cham Miettinen K, Mäkelä MM (2002) On scalarizing functions in multiobjective optimization. OR Spectrum 24(2):193–213. https://doi.org/10.1007/s00291-001-0092-9 Mitchell M (1998) An introduction to genetic algorithms. MIT Press, Cambridge Mitchell TM et al (1997) Machine learning. Burr Ridge 45(37):870–877 Montgomery DC (2017) Design and analysis of experiments. Wiley, Hoboken Mostafa SS, Mendonça F, Ravelo-Garcia A, Julia-Serda G, Morgado-Dias F (2020) Multi-objective hyperparameter optimization of convolutional neural network for obstructive sleep apnea detection. IEEE Access 8:129586–129599. https://doi.org/10.1109/ACCESS.2020.3009149 Nabil M, Mahmoud M, Ismail M, Serpedin E (2019) Deep recurrent electricity theft detection in ami networks with evolutionary hyper-parameter tuning. 2019 international conference on internet of things (ithings) and ieee green computing and communications (greencom) and ieee cyber, physical and social computing (cpscom) and ieee smart data (smartdata) (pp 1002–1008) Negrinho R, Gormley M, Gordon GJ, Patil D, Le N, Ferreira D (2019) Towards modular and programmable architecture search. Advances in neural information processing systems (pp 13715-13725). https://dl.acm.org/doi/abs/10.5555/3454287.3455517 Olsson DM, Nelson LS (1975) The Nelder-mead simplex procedure for function minimization. Technometrics 17(1):45–51. https://doi.org/10.1080/00401706.1975.10489269 Ounpraseuth ST (2008) Gaussian processes for machine learning. Taylor & Francis, Milton Park Ozaki Y, Tanigaki Y, Watanabe S, Onishi M (2020) Multiobjective tree-structured parzen estimator for computationally expensive optimization problems. Proceedings of the 2020 genetic and evolutionary computation conference (pp 533–541). https://doi.org/10.1145/3377930.3389817 Parker-Holder J, Nguyen V, Roberts SJ (2020) Provably efficient online hyperparameter optimization with population-based bandits. Adv Neural Inf Process Syst 33:17200–17211 Parsa M, Ankit A, Ziabari A, Roy K (2019) Pabo: Pseudo agent-based multi-objective bayesian hyperparameter optimization for efficient neural accelerator design. 2019 ieee/acm international conference on computer-aided design (iccad) (pp 1-8). https://doi.org/10.1109/ICCAD45719.2019.8942046 Pathak Y, Shukla PK, Arya K (2020) Deep bidirectional classification model for covid-19 disease infected patients. IEEE/ACM Trans Comput Biol Bioinf. https://doi.org/10.1109/TCBB.2020.3009859 Phillips PJ, Flynn PJ, Scruggs T, Bowyer KW, Chang J, Hoffman K, Worek W (2005) Overview of the face recognition grand challenge. 2005 ieee computer society conference on computer vision and pattern recognition (cvpr’05) (Vol. 1, pp 947-954). https://doi.org/10.1109/CVPR.2005.268 Picheny V (2014) A stepwise uncertainty reduction approach to constrained global optimization. Artificial intelligence and statistics (pp 787–795). http://proceedings.mlr.press/v33/picheny14.html Ponweiser W, Wagner T, Biermann D, Vincze M (2008) Multiobjective optimization on a limited budget of evaluations using model-assisted \(s\)-metric selection. International conference on parallel problem solving from nature (pp 784-794). https://doi.org/10.1007/978-3-540-87700-4_78 Provost F, Jensen D, Oates T (1999) Efficient progressive sampling. Proceedings of the fifth acm sigkdd international conference on knowledge discovery and data mining (pp 23-32). https://doi.org/10.1145/312129.312188 Qin AK, Huang VL, Suganthan PN (2008) Differential evolution algorithm with strategy adaptation for global numerical optimization. IEEE Trans Evol Comput 13(2):398–417. https://doi.org/10.1109/TEVC.2008.927706 Qin H, Shinozaki T, Duh K (2017) Evolution strategy based automatic tuning of neural machine translation systems. Proceeding of international workshop on spoken language translation (iwslt) (pp 120-128) Rajagopal A, Joshi GP, Ramachandran A, Subhalakshmi R, Khari M, Jha S, You J (2020) A deep learning model based on multi-objective particle swarm optimization for scene classification in unmanned aerial vehicles. IEEE Access 8:135383–135393. https://doi.org/10.1109/ACCESS.2020.3011502 Richter J, Kotthaus H, Bischl B, Marwedel P, Rahnenführer J, Lang M (2016) Faster model-based optimization through resource-aware scheduling strategies. International conference on learning and intelligent optimization (pp 267-273). https://doi.org/10.1007/978-3-319-50349-3_22 Rojas-Gonzalez S, Van Nieuwenhuyse I (2020) A survey on kriging-based infill algorithms for multiobjective simulation optimization. Comput Oper Res 116:104869. https://doi.org/10.1016/j.cor.2019.104869 Russell S, Norvig P (2010) Artificial intelligence: a modern approach, 3rd edn. Prentice Hall, Hoboken Salt L, Howard D, Indiveri G, Sandamirskaya Y (2019) Parameter optimization and learning in a spiking neural network for uav obstacle avoidance targeting neuromorphic processors. IEEE Trans Neural Netw Learn Syst. https://doi.org/10.1109/TNNLS.2019.2941506 Sanz-García A, Fernández-Ceniceros J, Antonanzas-Torres F, Pernia-Espinoza A, Martinez-De-Pison F (2015) Ga-parsimony: a ga-svr approach with feature selection and parameter optimization to obtain parsimonious solutions for predicting temperature settings in a continuous annealing furnace. Appl Soft Comput 35:13–28. https://doi.org/10.1016/j.asoc.2015.06.012 Schaffer JD (1985) Multiple objective optimization with vector evaluated genetic algorithms. Proceedings of the 1st international conference on genetic algorithms (p. 93-100). USA: L. Erlbaum Associates Inc Shah A, Ghahramani Z (2016) Pareto frontier learning with expensive correlated objectives. International conference on machine learning (pp 1919-1927). http://proceedings.mlr.press/v48/shahc16.html Shimizu H, Toyoda M (2021) Cma-es with coordinate selection for high-dimensional and ill-conditioned functions. Proceedings of the genetic and evolutionary computation conference companion (pp 209–210) Shinozaki T, Watanabe S, Duh K (2020) Automated development of dnn based spoken language systems using evolutionary algorithms. Deep neural evolution (pp 97-129). Springer. https://doi.org/10.1007/978-981-15-3685-4_4 Sierra MR, Coello CAC (2005) Improving pso-based multi-objective optimization using crowding, mutation and \(\epsilon\)-dominance. International conference on evolutionary multi-criterion optimization (pp 505-519). https://doi.org/10.1007/978-3-540-31880-4_35 Silva LF, Santos AAS, Bravo RS, Silva AC, Muchaluat-Saade DC, Conci A (2016) Hybrid analysis for indicating patients with breast cancer using temperature time series. Comput Methods Programs Biomed 130:142–153. https://doi.org/10.1016/j.cmpb.2016.03.002 Singh D, Kumar V, Kaur M (2020) Classification of covid-19 patients from chest ct images using multi-objective differential evolution-based convolutional neural networks. Eur J Clin Microbiol Infect Dis. https://doi.org/10.1007/s10096-020-03901-z Sjöberg A, Önnheim M, Gustavsson E, Jirstrand M (2019) Architecture-aware bayesian optimization for neural network tuning. International conference on artificial neural networks (pp 220-231). https://doi.org/10.1007/978-3-030-30484-3_19 Smithson SC, Yang G, Gross WJ, Meyer BH (2016) Neural networks designing neural networks: multi-objective hyper-parameter optimization. Proceedings of the 35th international conference on computer-aided design (pp 1-8). https://doi.org/10.1145/2966986.2967058 Snoek J, Larochelle H, Adams RP (2012) Practical bayesian optimization of machine learning algorithms. Advances in neural information processing systems, 25 Socha K, Dorigo M (2008) Ant colony optimization for continuous domains. Eur J Oper Res 185(3):1155–1173. https://doi.org/10.1016/j.ejor.2006.06.046 Sopov E, Ivanov I (2015) Self-configuring ensemble of neural network classifiers for emotion recognition in the intelligent human-machine in-teraction. 2015 ieee symposium series on computational intelligence (pp 1808-1815). https://doi.org/10.1109/SSCI.2015.252 Srinivas N, Deb K (1994) Multiobjective optimization using nondom-inated sorting in genetic algorithms. Evol Comput 2(3):221–248. https://doi.org/10.1162/evco.1994.2.3.221 Stamoulis D, Cai E, Juan D-C, Marculescu D (2018) Hyperpower: Power-and memory-constrained hyper-parameter optimization for neural networks. 2018 design, automation & test in europe conference & exhibition (date) (pp 19-24). https://doi.org/10.23919/DATE.2018.8341973 Stamoulis D, Chin T-W, Prakash AK, Fang H, Sajja S, Bognar M, Marculescu D (2018) Designing adaptive neural networks for energy-constrained image classification. Proceedings of the international conference on computer-aided design (pp 1-8). https://doi.org/10.1145/3240765.3240796 Storn R, Price K (1997) Differential evolution -a simple and efficient heuristic for global optimization over continuous spaces. J Global Optim 11(4):341–359. https://doi.org/10.1023/A:1008202821328 Swersky K, Snoek J, Adams RP (2013) Multi-task bayesian optimization. Advances in neural information processing systems, 26 Swersky K, Snoek J, Adams RP (2014) Freeze-thaw bayesian optimization. arXiv:1406.3896 Talbi E-G (2021) Automated design of deep neural networks: a survey and unified taxonomy. ACM Comput Surv (CSUR) 54(2):1–37. https://doi.org/10.1145/3439730 Tanabe R, Fukunaga A (2013) Success-history based parameter adap-tation for differential evolution. 2013 ieee congress on evolutionary computation (pp 71-78). https://doi.org/10.1109/CEC.2013.6557555 Tanaka T, Moriya T, Shinozaki T, Watanabe S, Hori T, Duh K (2016) Automated structure discovery and parameter tuning of neural network language model based on evolution strategy. 2016 ieee spoken language technology workshop (slt) (pp 665-671). https://doi.org/10.1109/SLT.2016.7846334 Tharwat A (2020) Classification assessment methods. Appl Comput Inf. https://doi.org/10.1016/j.aci.2018.08.003 Thornton C, Hutter F, Hoos HH, Leyton-Brown K (2013) Auto-weka: Combined selection and hyperparameter optimization of classification algorithms. Proceedings of the 19th acm sigkdd international conference on knowledge discovery and data mining (pp 847-855). https://doi.org/10.1145/2487575.2487629 Tripathy R, Bilionis I, Gonzalez M (2016) Gaussian processes with built-in dimensionality reduction: applications to high-dimensional uncertainty propagation. J Comput Phys 321:191–223 van Rijn JN, Abdulrahman SM, Brazdil P, Vanschoren J (2015) Fast algorithm selection using learning curves. International symposium on intelligent data analysis (pp 298-309). https://doi.org/10.1007/978-3-319-24465-5_26 Vanschoren J (2019) Meta-learning. Automated machine learning: methods, systems, challenges (pp 35-61). Springer, Cham Victoria AH, Maragatham G (2021) Automatic tuning of hyper-parameters using bayesian optimization. Evolving Systems 217–223. https://doi.org/10.1007/s12530-020-09345-2 Wang D, Tan D, Liu L (2018) Particle swarm optimization algorithm: an overview. Soft Comput 22(2):387–408 Wang B, Sun Y, Xue B, Zhang M (2019) Evolving deep neural networks by multi-objective particle swarm optimization for image classification. Proceedings of the genetic and evolutionary computation conference (pp 490-498). https://doi.org/10.1145/3321707.3321735 Wang B, Xue B, Zhang M (2020) Particle swarm optimization for evolving deep convolutional neural networks for image classification: Single-and multi-objective approaches. Deep neural evolution (pp 155-184). Springer. https://doi.org/10.1007/978-981-15-3685-4_6 Wang F, Zhang H, Zhou A (2021) A particle swarm optimization algorithm for mixed-variable optimization problems. Swarm Evol Comput 60:100808 Wawrzyński P (2017) Asd+ m: automatic parameter tuning in stochastic optimization and on-line learning. Neural Netw 96:1–10. https://doi.org/10.1016/j.neunet.2017.07.007 Wistuba M, Schilling N, Schmidt-Thieme L (2018) Scalable gaussian process-based transfer surrogates for hyperparameter optimization. Mach Learn 107(1):43–78. https://doi.org/10.1007/s10994-017-5684-y Wu J, Chen S, Liu X (2020) Efficient hyperparameter optimization through model-based reinforcement learning. Neurocomputing 409:381–393 Yang L, Shami A (2020) On hyperparameter optimization of machine learning algorithms: theory and practice. Neurocomputing 415:295–316. https://doi.org/10.1016/j.neucom.2020.07.061 Zames G, Ajlouni N, Ajlouni N, Ajlouni N, Holland J, Hills W, Gold-berg D (1981) Genetic algorithms in search, optimization and machine learning. Inf Technol J 3(1):301–302 Zhang C, Lim P, Qin AK, Tan KC (2016) Multiobjective deep belief networks ensemble for remaining useful life estimation in prognostics. IEEE Trans Neural Netw Learn Syst 28(10):2306–2318. https://doi.org/10.1109/TNNLS.2016.2582798 Zhang M, Ni Q, Zhao S, Wang Y, Shen C (2020) A combined prediction method for short-term wind speed using variational mode decomposition based on parameter optimization. 2020 ieee international conference on systems, man, and cybernetics (smc) (pp 2607-2614) Zhang Q, Li H (2007) Moea/d: a multiobjective evolutionary algorithm based on decomposition. IEEE Trans Evol Comput 11(6):712–731. https://doi.org/10.1109/TEVC.2007.892759 Zhao S-Z, Suganthan PN, Zhang Q (2012) Decomposition-based multiobjective evolutionary algorithm with an ensemble of neighborhood sizes. IEEE Trans Evol Comput 16(3):442–446. https://doi.org/10.1109/TEVC.2011.2166159 Zitzler E, Deb K, Thiele L (2000) Comparison of multiobjective evolutionary algorithms: empirical results. Evol Comput 8(2):173–195. https://doi.org/10.1162/106365600568202 Zitzler E, Thiele L (1999) Multiobjective evolutionary algorithms: a comparative case study and the strength pareto approach. IEEE Trans Evol Comput 3(4):257–271. https://doi.org/10.1109/4235.797969 Zitzler E, Laumanns M, Thiele L (2001) Spea 2: Improving the strength pareto evolutionary algorithm. TIK-report 103. https://doi.org/10.3929/ethz-a-004284029