Tăng tốc độ xử lý dựa trên FPGA cho mạng nơ-ron tích chập đầy đủ rời rạc dưới dạng cân bằng trọng số theo bộ lọc với giải thuật lát chồng chéo

Journal of Signal Processing Systems - Tập 93 - Trang 499-512 - 2021
Masayuki Shimoda1, Youki Sada1, Hiroki Nakahara1
1Tokyo Institute of Technology, Tokyo, Japan

Tóm tắt

Mạng nơ-ron tích chập (CNN) thể hiện hiệu suất hàng đầu trong các tác vụ thị giác máy tính. CNN cần phần cứng có tốc độ cao, tiêu thụ điện năng thấp và độ chính xác cao cho nhiều tình huống khác nhau, chẳng hạn như môi trường cạnh. Tuy nhiên, số lượng trọng số rất lớn khiến các hệ thống nhúng không thể lưu trữ do bộ nhớ trong chip hạn chế. Một phương pháp khác được sử dụng để giảm kích thước hình ảnh đầu vào cho việc xử lý thời gian thực, nhưng điều đó gây ra sự sụt giảm đáng kể trong độ chính xác. Mặc dù đã có các đề xuất cho CNN rời rạc được tỉa, và các bộ tăng tốc đặc biệt, yêu cầu truy cập ngẫu nhiên dẫn đến số lượng lớn bộ đa hợp mở rộng cho mức độ song song cao, điều này trở nên phức tạp hơn và không phù hợp cho việc triển khai FPGA. Để giải quyết vấn đề này, chúng tôi đề xuất phương pháp tỉa theo bộ lọc với quá trình tinh chế và bộ tăng tốc bỏ qua trọng số bằng bộ nhớ RAM khối (BRAM). Nó loại bỏ trọng số sao cho mỗi bộ lọc có cùng số lượng trọng số không bằng không, thực hiện đào tạo lại với quá trình tinh chế, trong khi vẫn giữ được độ chính xác tương đương. Hơn nữa, việc tỉa theo bộ lọc cho phép bộ tăng tốc của chúng tôi khai thác tính song song giữa các bộ lọc, trong đó một khối xử lý cho một lớp thực hiện các bộ lọc đồng thời với kiến trúc đơn giản. Chúng tôi cũng đề xuất một thuật toán lát chồng chéo, trong đó các viên gạch được khai thác với sự chồng chéo để ngăn ngừa cả việc giảm độ chính xác và việc sử dụng cao các BRAM lưu trữ hình ảnh độ phân giải cao. Đánh giá của chúng tôi sử dụng các tác vụ phân đoạn ngữ nghĩa cho thấy tốc độ tăng gấp 1,8 lần và hiệu quả năng lượng tăng 18,0 lần trong thiết kế FPGA của chúng tôi so với GPU để bàn. Thêm vào đó, so với việc triển khai FPGA thông thường, tốc độ tăng và sự cải thiện độ chính xác lần lượt là 1,09 lần và 6,6 điểm. Do đó, phương pháp của chúng tôi hữu ích cho việc triển khai FPGA và thể hiện độ chính xác đáng kể cho các ứng dụng trong các hệ thống nhúng.

Từ khóa

#mạng nơ-ron tích chập #tăng tốc FPGA #trọng số rời rạc #tỉa theo bộ lọc #bộ nhớ RAM khối #phân đoạn ngữ nghĩa

Tài liệu tham khảo

Albericio, J., Judd, P., Hetherington, T., Aamodt, T., Jerger, N. E., & Moshovos, A. (2016). Cnvlutin: ineffectual-neuron-free deep neural network computing. In 2016 ACM/IEEE 43rd annual international symposium on computer architecture (ISCA). https://doi.org/10.1109/ISCA.2016.11 (pp. 1–13). Alvarez, J. M., & Salzmann, M. (2017). Compression-aware training of deep networks. In Advances in neural information processing systems (pp. 856–867). Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). Segnet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481–2495. https://doi.org/10.1109/TPAMI.2016.2644615. Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L. D., Monfort, M., Muller, U., Zhang, J., & et al. (2016). End to end learning for self-driving cars. arXiv:1604.07316. Cao, S., Zhang, C., Yao, Z., Xiao, W., Nie, L., Zhan, D., Liu, Y., Wu, M., & Zhang, L. (2019). Efficient and effective sparse lstm on fpga with bank-balanced sparsity. In Proceedings of the 2019 ACM/SIGDA international symposium on field-programmable gate arrays, FPGA ’19. https://doi.org/10.1145/3289602.3293898 (pp. 63–72). New York: Association for Computing Machinery. Chollet, F. (2017). Xception: deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1251– 1258). Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213–3223). Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., & Bengio, Y. (2016). Binarized neural networks: training deep neural networks with weights and activations constrained to + 1 or − 1. arXiv:1602.02830. Deng, C., Liao, S., Xie, Y., Parhi, K. K., Qian, X., & Yuan, B. (2018). Permdnn: efficient compressed dnn architecture with permuted diagonal matrices. In 2018 51st Annual IEEE/ACM international symposium on microarchitecture (MICRO). https://doi.org/10.1109/MICRO.2018.00024 (pp. 189–202). Fan, H., Liu, S., Ferianc, M., Ng, H., Que, Z., Liu, S., Niu, X., & Luk, W. (2018). A real-time object detection accelerator with compressed ssdlite on fpga. In 2018 International conference on field-programmable technology (FPT). https://doi.org/10.1109/FPT.2018.00014 (pp. 14–21). Fang, S., Tian, L., Wang, J., Liang, S., Xie, D., Chen, Z., Sui, L., Yu, Q., Sun, X., Yao, S., Shan, Y., & Wang, Y. (2018). Real-time object detection and semantic segmentation hardware system with deep learning networks. In International conference on field programmable technology, ICFPT 2018, Okinawa, Japan, December 10–14, 2018, p. (to be appear). Gray, S., Radford, A., & Kingma, D. P. (2017). Gpu kernels for block-sparse weights. arXiv:1711.09224, 3. Han, S., Pool, J., Tran, J., & Dally, W. (2015). Learning both weights and connections for efficient neural network. In Advances in neural information processing systems (pp. 1135–1143). Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M. A., & Dally, W. J. (2016). Eie: efficient inference engine on compressed deep neural network. In 2016 ACM/IEEE 43rd annual international symposium on computer architecture (ISCA) (pp. 243–254). IEEE. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In 2016 IEEE Conference on computer vision and pattern recognition (CVPR). https://doi.org/10.1109/CVPR.2016.90 (pp. 770–778). He, Y., Zhang, X., & Sun, J. (2017). Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE international conference on computer vision (pp. 1389–1397). He, Y., Liu, P., Wang, Z., Hu, Z., & Yang, Y. (2019). Filter pruning via geometric median for deep convolutional neural networks acceleration. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4340–4349). Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv:1503.02531. Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., & Adam, H. (2017). Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Ioffe, S., & Szegedy, C. (2015). Batch normalization: accelerating deep network training by reducing internal covariate shift. arXiv:1502.03167. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2704–2713). Kang, H. (2019). Real-time object detection on 640x480 image with vgg16+ssd. In 2019 International conference on field-programmable technology (ICFPT). https://doi.org/10.1109/ICFPT47387.2019.00082 (pp. 419–422). Kang, H. J. (2019). Accelerator-aware pruning for convolutional neural networks. IEEE Transactions on Circuits and Systems for Video Technology, 30(7), 2093–2103. IEEE. Krishnamoorthi, R. (2018). Quantizing deep convolutional networks for efficient inference: a whitepaper. arXiv:1806.08342. Krizhevsky, A., Sutskever, I., & Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. In Proceedings of the 25th international conference on neural information processing systems, NIPS’12. http://dl.acm.org/citation.cfm?id=2999134.2999257, (Vol. 1 pp. 1097–1105). Curran Associates Inc. Lai, B., Pan, J., & Lin, C. (2019). Enhancing utilization of simd-like accelerator for sparse convolutional neural networks. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 27(5), 1218–1222. https://doi.org/10.1109/TVLSI.2019.2897052. Lecun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444. https://doi.org/10.1038/nature14539. Li, J., Yan, G., Lu, W., Jiang, S., Gong, S., Wu, J., & Li, X. (2018). Ccr: a concise convolution rule for sparse neural network accelerators. In 2018 Design, automation test in europe conference exhibition (DATE). https://doi.org/10.23919/DATE.2018.8342001 (pp. 189–194). Lin, C.Y., & Lai, B.C. (2018). Supporting compressed-sparse activations and weights on simd-like accelerator for sparse convolutional neural networks. In Proceedings of the 23rd Asia and South Pacific design automation conference, ASPDAC ’18. http://dl.acm.org/citation.cfm?id=3201607.3201630 (pp. 105–110). Piscataway: IEEE Press. Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S. E., Fu, C. Y., & Berg, A. C. (2016). Ssd: single shot multibox detector. European conference on computer vision. Luo, J. H., Wu, J., & Lin, W. (2017). Thinet: a filter level pruning method for deep neural network compression. In Proceedings of the IEEE international conference on computer vision (pp. 5058–5066). Lyu, Y., Bai, L., & Huang, X. (2018). Real-time road segmentation using lidar data processing on an fpga. In 2018 IEEE international symposium on circuits and systems (ISCAS). https://doi.org/10.1109/ISCAS.2018.8351244(pp. 1–5). Mao, H., Han, S., Pool, J., Li, W., Liu, X., Wang, Y., & Dally, W. J. (2017). Exploring the regularity of sparse structure in convolutional neural networks. arXiv:1705.08922. Molchanov, D., Ashukha, A., & Vetrov, D. (2017). Variational dropout sparsifies deep neural networks. arXiv:1701.05369. Narang, S., Undersander, E., & Diamos, G. (2017). Block-sparse recurrent neural networks. arXiv:1711.02782. NVIDIA. (2020). TensorRT. https://developer.nvidia.com/tensorrt. Parashar, A., Rhu, M., Mukkara, A., Puglielli, A., Venkatesan, R., Khailany, B., Emer, J., Keckler, S. W., & Dally, W. J. (2017). Scnn: an accelerator for compressed-sparse convolutional neural networks. In 2017 ACM/IEEE 44th annual international symposium on computer architecture (ISCA). https://doi.org/10.1145/3079856.3080254 (pp. 27–40). Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., & Chintala, S. (2019). Pytorch: an imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d’ Alché-Buc, E. Fox, & R. Garnett (Eds.) Advances in neural information processing systems (Vol. 32, pp. 8024–8035). http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdfhttp://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdfhttp://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf. Curran Associates, Inc. Romera, E., Alvarez, J. M., Bergasa, L. M., & Arroyo, R. (2017). Erfnet: efficient residual factorized convnet for real-time semantic segmentation. IEEE Transactions on Intelligent Transportation Systems, 19(1), 263–272. Saqib, M., Daud Khan, S., Sharma, N., & Blumenstein, M. (2017). A study on detecting drones using deep convolutional neural networks. In 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS) (pp. 1–5). Shelhamer, E., Long, J., & Darrell, T. (2017). Fully convolutional networks for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 640–651. https://doi.org/10.1109/TPAMI.2016.2572683. Shimoda, M., Sada, Y., & Nakahara, H. (2019). Filter-wise pruning approach to FPGA implementation of fully convolutional network for semantic segmentation. In Applied reconfigurable computing - 15th international symposium, ARC 2019, Darmstadt, Germany, April 9–11, 2019, Proceedings. https://doi.org/10.1007/978-3-030-17227-5_26(pp. 371–386). Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Tokui, S., Oono, K., Hido, S., & Clayton, J. (2015). Chainer: a next-generation open source framework for deep learning. In Proceedings of workshop on machine learning systems (LearningSys) in the twenty-ninth annual conference on neural information processing systems (NIPS). http://learningsys.org/papers/LearningSys_2015_paper_33.pdf. Wang, J., Yuan, Z., Liu, R., Yang, H., & Liu, Y. (2019). An n-way group association architecture and sparse data group association load balancing algorithm for sparse cnn accelerators. In Proceedings of the 24th Asia and South Pacific design automation conference, ASPDAC ’19. https://doi.org/10.1145/3287624.3287626. http://doi.acm.org/10.1145/3287624.3287626 (pp. 329–334). New York: ACM. Wong, H.T.H. A superscalar out-of-order x86 soft processor for fpga. Ph.D. thesis. Wu, B., Wan, A., Yue, X., Jin, P., Zhao, S., Golmant, N., Gholaminejad, A., Gonzalez, J., & Keutzer, K. (2018). Shift: a zero flop, zero parameter alternative to spatial convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 9127–9135). Xiang, Y., Schmidt, T., Narayanan, V., & Fox, D. (2017). Posecnn: a convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv:1711.00199. Yang, Y., Huang, Q., Wu, B., Zhang, T., Ma, L., Gambardella, G., Blott, M., Lavagno, L., Vissers, K., Wawrzynek, J., & et al. (2019). Synetgy: algorithm-hardware co-design for convnet accelerators on embedded fpgas. In Proceedings of the 2019 ACM/SIGDA international symposium on field-programmable gate arrays (pp. 23–32). Yu, J., Lukefahr, A., Palframan, D., Dasika, G., Das, R., & Mahlke, S. (2017). Scalpel: customizing dnn pruning to the underlying hardware parallelism. In Proceedings of the 44th annual international symposium on computer architecture, ISCA ’17. (pp. 548–560). New York: ACM https://doi.org/10.1145/3079856.3080215. Yuan, Z., Yue, J., Yang, H., Wang, Z., Li, J., Yang, Y., Guo, Q., Li, X., Chang, M., Yang, H., & Liu, Y. (2018). Sticker: a 0.41-62.1 tops/w 8bit neural network processor with multi-sparsity compatible convolution arrays and online tuning acceleration for fully connected layers. In 2018 IEEE symposium on VLSI circuits. https://doi.org/10.1109/VLSIC.2018.8502404 (pp. 33–34). Zhang, S., Du, Z., Zhang, L., Lan, H., Liu, S., Li, L., Guo, Q., Chen, T., & Chen, Y. (2016). Cambricon-x: an accelerator for sparse neural networks. In 2016 49th Annual IEEE/ACM international symposium on microarchitecture (MICRO). https://doi.org/10.1109/MICRO.2016.7783723 (pp. 1–12). Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In CVPR. https://doi.org/10.1109/CVPR.2017.660 (pp. 6230–6239). Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In 2017 IEEE conference on computer vision and pattern recognition (CVPR). https://doi.org/10.1109/CVPR.2017.660 (pp. 6230–6239). Zhao, H., Qi, X., Shen, X., Shi, J., & Jia, J. (2018). ICNet for real-time semantic segmentation on high-resolution images. In ECCV. Zhou, X., Du, Z., Guo, Q., Liu, S., Liu, C., Wang, C., Zhou, X., Li, L., Chen, T., & Chen, Y. (2018). Cambricon-s: addressing irregularity in sparse neural networks through a cooperative software/hardware approach. In 2018 51st Annual IEEE/ACM international symposium on microarchitecture (MICRO). https://doi.org/10.1109/MICRO.2018.00011 (pp. 15–28). Zhu, M., & Gupta, S. (2017). To prune, or not to prune: exploring the efficacy of pruning for model compression. CoRR arXiv:1710.01878.