Kiểm thử vi phân cho học máy: phân tích cho các thuật toán phân loại ngoài học sâu

Empirical Software Engineering - Tập 28 - Trang 1-38 - 2023
Steffen Herbold1, Steffen Tunkel2
1Faculty of Computer Science and Mathematics, University of Passau, Passau, Germany
2Institute of Computer Science, University of Goettingen, Goettingen, Germany

Tóm tắt

Kiểm thử vi phân là một phương pháp hữu ích sử dụng các triển khai khác nhau của cùng một thuật toán và so sánh các kết quả để kiểm thử phần mềm. Trong những năm gần đây, phương pháp này đã được sử dụng thành công cho các chiến dịch kiểm thử của các khung học sâu. Tuy nhiên, có rất ít kiến thức về việc áp dụng kiểm thử vi phân ngoài lĩnh vực học sâu. Trong bài báo này, chúng tôi muốn lấp đầy khoảng trống này cho các thuật toán phân loại. Chúng tôi thực hiện một nghiên cứu trường hợp sử dụng Scikit-learn, Weka, Spark MLlib và Caret, trong đó chúng tôi xác định tiềm năng của kiểm thử vi phân bằng cách xem xét các thuật toán nào có sẵn trong nhiều framework, khả thi bằng cách xác định các cặp thuật toán mà nên thể hiện hành vi tương tự, và hiệu quả bằng cách thực hiện các bài kiểm tra cho các cặp đã xác định và phân tích các sai lệch. Trong khi chúng tôi phát hiện tiềm năng lớn cho các thuật toán phổ biến, khả thi dường như bị hạn chế vì thường không thể xác định các cấu hình giống nhau trong các framework khác. Việc thực hiện các bài kiểm tra khả thi cho thấy có rất nhiều sai lệch về điểm số và lớp. Chỉ có phương pháp nhân nhượng dựa trên ý nghĩa thống kê của các lớp không dẫn đến một lượng lớn các lỗi kiểm thử. Tiềm năng của kiểm thử vi phân ngoài học sâu dường như bị hạn chế cho nghiên cứu về chất lượng của các thư viện học máy. Các nhà thực hành vẫn có thể sử dụng phương pháp này nếu họ có kiến thức sâu về các triển khai, đặc biệt nếu một oracle thô chỉ cân nhắc những khác biệt có ý nghĩa của các lớp là đủ.

Từ khóa

#kiểm thử vi phân #học máy #thuật toán phân loại #học sâu #Scikit-learn #Weka #Spark MLlib #Caret

Tài liệu tham khảo

Abadi M, Agarwal A , Barham P, Brevdo E , Chen Z , Citro C , Corrado GS, Davis A , Dean J , Devin M , Ghemawat S, Goodfellow I , Harp A, Irving G , Isard M , Jia Y , Jozefowicz R , Kaiser L , Kudlur M , Levenberg J , Mané D , Monga R , Moore S , Murray D , Olah C , Schuster M, Shlens J, Steiner B , Sutskever I , Talwar K , Tucker P, Vanhoucke V, Vasudevan V , Viégas F , Vinyals O , Warden P , Wattenberg M , Wicke M, Yu Y, Zheng X (2015) TensorFlow : large-scale machine learning on heterogeneous systems https://www.tensorflow.org/softwareavailablefromtensorflow.org Abadi M, Barham P, Chen J, Chen Z, Davis A, Dean J, Devin M, Ghemawat S, Irving G, Isard M, Kudlur M, Levenberg J, Monga R, Moore S, Murray DG, Steiner B, Tucker P, Vasudevan V, Warden P, Wicke M, Yu Y, Zheng X (2016) Tensorflow: a system for large-scale machine learning. In: Proceedings of the 12th USENIX conference on operating systems design and implementation. OSDI’16, USENIX Association, Berkeley, pp 265–283. http://dl.acm.org/citation.cfm?id=3026877.3026899 Asyrofi MH, Thung F, Lo D, Jiang L (2020) Crossasr: efficient differential testing of automatic speech recognition via text-to-speech Barash G, Farchi E, Jayaraman I, Raz O, Tzoref-Brill R, Zalmanovici M (2019) Bridging the gap between ml solutions and their business requirements using feature interactions. In: Proceedings of the 2019 27th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. Association for Computing Machinery, New York, ESEC/FSE 2019, pp 1048–1058 . https://doi.org/10.1145/3338906.3340442 Benavoli A, Corani G, Demsar J, Zaffalon M (2016) Time for a change: a tutorial for comparing multiple classifiers through Bayesian analysis. ArXiv:1606.04316 Braiek HB, Khomh F (2020) On testing machine learning programs. J Syst Softw 164:110542. https://doi.org/10.1016/j.jss.2020.110542 Breiman L (1996) Bagging predictors. Mach Learn 24:123–140. Breiman L (2001) Random forests. Mach Learn 45(1):5–32. https://doi.org/10.1023/a:1010933404324 Brieman L, Friedman J, Olshen R, Stone C (1984) Classification and regression trees. Wadsworth, Belmont Cliff N (1993) Dominance statistics: ordinal analyses to answer ordinal questions. Psychol Bull 114(3):494–509. https://doi.org/10.1037/0033-2909.114.3.494 Cook TD, Campbell DT, Day A (1979) Quasi-experimentation: design & analysis issues for field settings, vol 351. Houghton Mifflin, Boston Davis MD, Weyuker EJ (1981) Pseudo-oracles for non-testable programs. In: Proceedings of the ACM ’81 conference, ACM, New York, NY, USA, ACM ’81, pp 254–257, https://doi.org/10.1145/800175.809889 Ding J, Kang X, Hu XH (2017) Validating a deep learning framework by metamorphic testing. In: 2017 IEEE/ACM 2nd international workshop on metamorphic testing (MET), pp 28–34, 10.1109/MET.2017.2 Frank E, Hall MA, Witten IH (2016) The WEKA workbench online appendix for Data Mining: Practical Machine Learning Tools and Techniques. Morgan Kaufmann Freund Y, Schapire RE (1997) A decision-theoretic generalization of on-line learning and an application to boosting. J Comput Syst Sci 55(1):119 – 139. https://doi.org/10.1006/jcss.1997.1504 Giray G (2021) A software engineering perspective on engineering machine learning systems: state of the art and challenges. J Syst Softw 180:111031. https://doi.org/10.1016/j.jss.2021.111031 Groce A, Kulesza T, Zhang C, Shamasunder S, Burnett M, Wong WK, Stumpf S, Das S, Shinsel A, Bice F, McIntosh K (2014) You are the only possible oracle: effective test selection for end users of interactive machine learning systems. IEEE Trans Softw Eng 40(3):307–323. https://doi.org/10.1109/TSE.2013.59 Gross P, Boulanger A, Arias M, Waltz D, Long PM, Lawson C, Anderson R, Koenig M, Mastrocinque M, Fairechio W, Johnson JA, Lee S, Doherty F, Kressner A (2006) Predicting electricity distribution feeder failures using machine learning susceptibility analysis. In: Proceedings of the 18th conference on innovative applications of artificial intelligence, vol 2. AAAI Press, IAAI’06, pp 1705–1711 http://dl.acm.org/citation.cfm?id=1597122.1597127 Guo Q, Xie X, Li Y, Zhang X, Liu Y, Li X, Shen C (2020) Audee: automated testing for deep learning frameworks. In: 2020 35th IEEE/ACM international conference on automated software engineering (ASE), pp 486–498 Guo J, Zhao Y, Song H, Jiang Y (2021) Coverage guided differential adversarial testing of deep learning systems. IEEE Trans Netw Sci Eng 8(2):933–942. https://doi.org/10.1109/TNSE.2020.2997359 Herbold S, Haar T (2022) Smoke testing for machine learning: simple tests to discover severe bugs. Empir Softw Eng 27(2). https://doi.org/10.1007/s10664-021-10073-7 ISO/IEC/IEEE (2017) Iso/iec/ieee international standard - systems and software engineering–vocabulary. ISO/IEC/IEEE, 24765:2017(E) pp 1–541. https://doi.org/10.1109/IEEESTD.2017.8016712 Karpathy A (2018) Cs231n: convolutional neural networks for visual recognition. https://cs231n.github.io/neural-networks-3/ Kuhn M (2018) Caret: classification and regression training. https://CRAN.R-project.org/package=caret,rpackageversion6.0-80 Mann HB, Whitney DR (1947) On a test of whether one of two random variables is stochastically larger than the other. Ann Math Stat 18(1):50 – 60. https://doi.org/10.1214/aoms/1177730491 Marijan D, Gotlieb A (2020) Software testing for machine learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol 34. pp 13576-13582 Martínez-Fernández S, Bogner J, Franch X, Oriol M, Siebert J, Trendowicz A, Vollmer A M, Wagner S (2021) Software engineering for ai-based systems: a survey. ACM Trans Softw Eng Methodol McCullough BD, Mokfi T, Almaeenejad M (2019) On the accuracy of linear regression routines in some data mining packages. WIREs Data Min Knowl Disc 9(3):e1279. https://doi.org/10.1002/widm.1279 Meng X, Bradley J, Yavuz B, Sparks E, Venkataraman S, Liu D, Freeman J, Tsai D, Amde M, Owen S et al (2016) Mllib: machine learning in apache spark. J Mach Learn Res 17(1):1235–1241 Murphy C, Kaiser GE, Arias M (2007) An approach to software testing of machine learning applications. In: SEKE, vol 167 Murphy C, Kaiser GE, Hu L, Wu L (2008) Properties of machine learning applications for use in metamorphic testing. In: SEKE, vol 8. pp 867-872 Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, Kopf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J, Chintala S (2019) Pytorch: an imperative style, high-performance deep learning library. In: Wallach H, Larochelle H, Beygelzimer A, d'Alché-Buc F, Fox E, Garnett R (eds) advances in neural information processing systems 32, Curran Associates, Inc, pp 8024–8035. http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf Patton MQ (2014) Qualitative research & evaluation methods: integrating theory and practice. Sage Publications Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E (2011) Scikit-learn: machine learning in Python. J Mach Learn Res 12:2825–2830 Pei K, Cao Y, Yang J, Jana S (2019) Deepxplore: automated whitebox testing of deep learning systems. Commun ACM 62(11):137–145. https://doi.org/10.1145/3361566 Pham HV, Lutellier T, Qi W, Tan L (2019) Cradle: cross-backend validation to detect and localize bugs in deep learning libraries. In: 2019 IEEE/ACM 41st international conference on software engineering (ICSE), pp 1027–1038. https://doi.org/10.1109/ICSE.2019.00107 Quinlan JR (1986) Induction of decision trees. Mach Learn 1 (1):81–106 Quinlan JR (1993) C4.5: programs for machine learning. Morgan Kaufmann Publishers Inc, San Francisco Runeson P, Höst M (2009) Guidelines for conducting and reporting case study research in software engineering. Empir Softw Eng 14(2):131–164 Theano Development Team (2016) Theano: a Python framework for fast computation of mathematical expressions. arXiv:1605.02688 Tunkel S, Herbold S (2022) Replication Kit for: differential testing for machine learning: an analysis for classification algorithms beyond deep learning. https://doi.org/10.5281/zenodo.7341092 Wang Z, Yan M, Chen J, Liu S, Zhang D (2020) Deep learning library testing via effective model generation. In: Proceedings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering, Association for Computing Machinery, New York, NY, USA, ESEC/FSE 2020, pp 788–799. https://doi.org/10.1145/3368089.3409761 Wilcoxon F (1945) Individual comparisons by ranking methods. Biometrics Bulletin 1(6):80–83. http://www.jstor.org/stable/3001968 Wohlin C, Runeson P, Höst M, Ohlsson MC, Regnell B, Wesslen A (2012) Experimentation in Software Engineering. Springer Publishing Company, Incorporated Xie X, Ho JW, Murphy C, Kaiser G, Xu B, Chen TY (2011) Testing and validating machine learning classifiers by metamorphic testing. J Syst Softw 84(4):544–558,. https://doi.org/10.1016/j.jss.2010.11.920 Zaharia M, Chowdhury M, Franklin MJ, Shenker S, Stoica I (2010) Spark: cluster computing with working sets. In: Proceedings of the 2nd USENIX conference on hot topics in cloud computing (HotCloud) Zhang JM, Harman M, Ma L, Liu Y (2020) Machine learning testing: survey , landscapes and horizons. IEEE Trans Softw Eng, pp 1–1