AUC optimization for deep learning-based voice activity detection

EURASIP Journal on Audio, Speech, and Music Processing - Tập 2022 - Trang 1-12 - 2022

Xiao-Lei Zhang^1,2, Menglong Xu^1,2

¹Research & Development Institute of Northwestern Polytechnical University in Shenzhen, Shenzhen, China

²School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China

Tóm tắt

Voice activity detection (VAD) based on deep neural networks (DNN) have demonstrated good performance in adverse acoustic environments. Current DNN-based VAD optimizes a surrogate function, e.g., minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works on-the-fly with a dynamic decision threshold, and the receiver operating characteristic (ROC) curve is a global evaluation metric for VAD at all possible decision thresholds. In this paper, we propose to maximize the area under the ROC curve (MaxAUC) by DNN, which can maximize the performance of VAD in terms of the entire ROC curve. However, the objective of the AUC maximization is nondifferentiable. To overcome this difficulty, we relax the nondifferentiable loss function to two differentiable approximation functions—sigmoid loss and hinge loss. To study the effectiveness of the proposed MaxAUC-DNN VAD, we take either a standard feedforward neural network or a bidirectional long short-term memory network as the DNN model with either the state-of-the-art multi-resolution cochleagram or short-term Fourier transform as the acoustic feature. We conducted noise-independent training to all comparison methods. Experimental results show that taking AUC as the optimization objective results in higher performance than the common objectives of the minimum squared error and minimum cross-entropy. The experimental conclusion is consistent across different DNN structures, acoustic features, noise scenarios, training sets, and languages.

Tài liệu tham khảo

R. Tucker, Tucker, Voice activity detection using a periodicity measure. IEE Proc. I (Commun. Speech Vis.). 139(4), 377–380 (1992)

J.-C. Junqua, H. Wakita, in Acoustics, Speech, and Signal Processing, 1989. ICASSP-89., 1989 International Conference On. A comparative study of cepstral lifters and distance measures for all pole models of speech in noise (IEEE, 1989), pp. 476–479

E. Nemer, R. Goubran, S. Mahmoud, Robust voice activity detection using higher-order statistics in the LPC residual domain. IEEE Trans. Speech Audio Process. 9(3), 217–231 (2001)

J. Sohn, N.S. Kim, W. Sung, A statistical model-based voice activity detection. IEEE Signal Process. Lett. 6(1), 1–3 (1999)

J. Ramírez, J.C. Segura, C. Benítez, L. García, A. Rubio, Statistical voice activity detection using a multiple observation likelihood ratio test. IEEE Signal Process. Lett. 12(10), 689–692 (2005)

J.-H. Chang, N.S. Kim, Voice activity detection based on complex laplacian model. Electron. Lett. 39(7), 632–634 (2003)

J.W. Shin, J.-H. Chang, N.S. Kim, Statistical modeling of speech signals based on generalized gamma distribution. IEEE Signal Process. Lett. 12(3), 258–261 (2005)

J.H. Chang, N.S. Kim, S.K. Mitra, Voice activity detection based on multiple statistical models. IEEE Trans. Signal Process. 54(6), 1965–1976 (2006)

J. Padrell, D. Macho, C. Nadeu, in Acoustics, Speech, and Signal Processing, 2005. Proceedings.(ICASSP’05). IEEE International Conference On. Robust speech activity detection using lda applied to ff parameters, vol. 1 (IEEE, 2005), p. 557

J. Wu, X.L. Zhang, Efficient multiple kernel support vector machine based voice activity detection. IEEE Signal Process. Lett. 18(8), 466–499 (2011)

D. Dov, R. Talmon, I. Cohen, Multimodal kernel method for activity detection of sound sources. IEEE/ACM Trans. Audio Speech Lang. Process. 25(6), 1322–1334 (2017)

P. Teng, Y. Jia, Voice activity detection via noise reducing using non-negative sparse coding. IEEE Signal Process. Lett. 20(5), 475–478 (2013)

S.-W. Deng, J.-Q. Han, Statistical voice activity detection based on sparse representation over learned dictionary. Digit. Signal Process. 23(4), 1228–1232 (2013)

X.-L. Zhang, J. Wu, Deep belief networks based voice activity detection. IEEE Trans. Audio Speech Lang. Process. 21(4), 697–710 (2013)

X.-L. Zhang, J. Wu, in the 38th IEEE International Conference on Acoustic, Speech, and Signal Processing. Denoising deep neural networks based voice activity detection (2013), pp. 853–857

T. Hughes, K. Mierle, in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. Recurrent neural networks for voice activity detection (2013). pp. 7378–7382

F. Eyben, F. Weninger, S. Squartini, B. Schuller, in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. Real-life voice activity detection with lstm recurrent neural networks and an application to hollywood movies (IEEE, 2013) pp. 483–487

X.-L. Zhang, D. Wang, Boosting contextual information for deep neural network based voice activity detection. IEEE/ACM Trans. Audio Speech Lang. Process. 24(2), 252–264 (2016)

I. Hwang, H.-M. Park, J.-H. Chang, Ensemble of deep neural networks using acoustic environment classification for statistical model-based voice activity detection. Comput. Speech Lang. 38, 1–12 (2016)

Q. Wang, J. Du, X. Bao, Z.-R. Wang, L.-R. Dai, C.-H. Lee, In: Sixteenth Annual Conference of the International Speech Communication Association. A universal vad based on jointly trained deep neural networks (2015)

L. Wang, K. Phapatanaburi, Z. Go, S. Nakagawa, M. Iwahashi, J. Dang, in Proceedings of ICME. Limiting numerical precision of neural networks to achieve real-time voice activity detection (2018), pp. 1087–1092

Y. Tachioka, in Proceedings of ICASSP. Limiting numerical precision of neural networks to achieve real-time voice activity detection (2018), pp. 2236–2240

Y. Tachioka, in Proceedings of ICASSP. Dnn-based voice activity detection using auxiliary speech models in noisy environments (2018). pp. 5529–5533

W.A. Jassim, N. Harte, in Proceedings of ICASSP. Voice activity detection using neurograms (2018), pp. 5524–5528

Y. Jung, Y. Kim, Y. Choi, H. Kim, in Interspeech. Joint learning using denoising variational autoencoders for voice activity detection (2018), pp. 1210–1214

T. Xu, H. Zhang, X. Zhang, in 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). Joint training rescnn-based voice activity detection with speech enhancement (IEEE, 2019), pp. 1157–1162

G.W. Lee, H.K. Kim, Multi-task learning u-net for single-channel speech enhancement and mask-based voice activity detection. Appl. Sci. 10(9), 3230 (2020)

Y. Zhuang, S. Tong, M. Yin, Y. Qian, K. Yu, in 2016 10th International Symposium on Chinese Spoken Language Processing (ISCSLP). Multi-task joint-learning for robust voice activity detection (IEEE, 2016), pp. 1–5

X. Tan, X.-L. Zhang, in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Speech enhancement aided end-to-end multi-task learning for voice activity detection (IEEE, 2021), pp. 6823–6827

Y. Chen, S. Wang, Y. Qian, K. Yu, End-to-end speaker-dependent voice activity detection. arXiv preprint arXiv:2009.09906 (2020)

Z.-C. Fan, Z. Bai, X.-L. Zhang, S. Rahardja, J. Chen, in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Auc optimization for deep learning based voice activity detection (IEEE, 2019), pp. 6760–6764

H.B. Mann, D.R. Whitney, On a test of whether one of two random variables is stochastically larger than the other. Ann. Math. Stat. 50–60 (1947)

X.-L. Zhang, D. Wang, Boosting contextual information for deep neural network based voice activity detection. IEEE/ACM Trans Audio Speech Lang. Process. 24(2), 252–264 (2015)

Scholar Hub - Công cụ hỗ trợ trích dẫn và phân tích khoa học Việt Nam

Về chúng tôi

Scholar Hub là công cụ hỗ trợ trích dẫn và phân tích các bài báo, công bố khoa học Việt Nam. Công cụ trợ giúp người nghiên cứu, tạp chí, đơn vị nghiên cứu tra cứu, phân tích và thống kê dữ liệu nghiên cứu khoa học tại Việt Nam và quốc tế.
ScholarHub KHÔNG đăng thông tin tổng hợp, KHÔNG đăng lại nội dung từ các trang báo chí Việt Nam hoặc trang thông tin điện tử khác tại Việt Nam.

Thông tin, cập nhật

Đăng ký Tạp chí tham gia vào Scholar Hub

Phản hồi ý kiến về Scholar Hub

Bài viết, nội dung cập nhật

Chủ đề khoa học

Website liên kết

Hệ thống CSDL Khoa học & Công nghệ

Phần mềm kiểm tra trùng lặp Kiểm Tra Tài Liệu

Phần mềm xuất bản tạp chí điện tử VOJS

Nền tảng trắc nghiệm và đề thi đa lĩnh vực LetQA