Phương pháp phát hiện điểm cuối speech thích ứng trong môi trường SNR thấp

International Journal of Speech Technology - Tập 20 - Trang 651-658 - 2017
Linhui Sun1,2, Min Su1, Zhenzhen Yang2
1College of Telecommunications, Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing, China
2Key Lab of Broadband Wireless Communication and Sensor Network Technology, Ministry of Education, Nanjing University of Posts and Telecommunications, Nanjing, China

Tóm tắt

Việc phát hiện điểm cuối của tín hiệu nói đã cho thấy sự thành công trong nhận diện và nâng cao chất lượng tín hiệu nói. Tuy nhiên, các phương pháp phát hiện điểm cuối truyền thống thường mất hiệu quả trong các môi trường có tỷ lệ tín hiệu trên tiếng ồn (SNR) thấp hoặc trong các môi trường có tiếng ồn không ổn định. Để cải thiện độ chính xác của việc phát hiện điểm cuối trong môi trường SNR thấp, một phương pháp phát hiện điểm cuối dựa trên thuật toán thích ứng để điều chỉnh ngưỡng được đề xuất trong bài báo này. Phép trừ phổ từ việc ước lượng phổ đa đầu thu được thực hiện để nâng cao chất lượng tín hiệu nói. Trong quá trình phát hiện, khoảng cách cepstral của hệ số cepstrum tần số Mel (MFCC) được sử dụng và các ngưỡng được điều chỉnh một cách thích ứng theo các môi trường khác nhau. Các thí nghiệm mô phỏng cho thấy, trong các môi trường tiếng ồn khác nhau với các SNR khác nhau, thuật toán của chúng tôi có độ chính xác phát hiện điểm cuối tốt hơn so với các thuật toán phát hiện khác. Ngoài ra, thuật toán cũng thể hiện sự mạnh mẽ khi hoạt động trong các môi trường SNR thấp.

Từ khóa

#phát hiện điểm cuối #thuật toán thích ứng #tín hiệu nói #cepstrum tần số Mel #khoảng cách cepstral #môi trường tiếng ồn

Tài liệu tham khảo

Bou-Ghazale, S. E., & Assaleh, K. (2002). A robust endpoint detection of speech for noisy environments with application to automatic speech recognition. 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing, Orlando, FL, USA, pp. IV-3808-IV-3811. Dong, H. (2016). An improved speech endpoint detection algorithm in low SNR environment. Computer Technology and Development, 3, 71–74. Han, Z. H. & Wang, J. (2015). Research on speech endpoint detection in low signal-to-noise ratios, The 27th Chinese Control and Decision Conference (2015 CCDC), Qingdao, pp. 3635–3639. Ion, V. & Haeb-Umbach, R. (2008). A novel uncertainty decoding rule with applications to transmission error robust speech recognition. IEEE Transactions on Audio Speech and Language Processing, 16(5), 1047–1060. Ji, Z. H., Yang, H., Li, R. & Jin, Y. (2016). A speech endpoint detection algorithm based on short-term autocorrelation and zero-crossing. Electronic Science & Technology, 9, 52–55. Jin, L. & Cheng, J. (2010). An Improved Speech Endpoint Detection Based on Spectral Subtraction and Adaptive Sub-band Spectral Entropy, 2010 International Conference on Intelligent Computation Technology and Automation, Changsha, pp. 591–594. Li, R., Hu, C. H., & Yu, J. (2013). Improvement of speech endpoint detection algorithm based on spectral entropy. Journal of Wuhan University of Technology, 7, 134–139. Song, Z. H. (2013). Application of Matlab in analysis and synthesis of speech signa. Beijing: Beijing university of aeronautics and astronautics press. Walden, A., Percival, D. B., & McCoy, E. J. (1998). Spectrum estimation by wavelet thresholding of multitaper estimators. IEEE Transactions on Signal Processing, 46(12), 3153–3165. Wu, P., Zhao, G., & Zou, M. (2008). Modified spectral subtraction based on multitaper spectrum estimation. Modern Electronics Technique, 12, 150–152. Yin, Q., Wu, H., & Zhao, L. (2008). Research on Endpoint Detection Method of Noisy Speech Signal[A]. Sichuan Institute of Acoustics, Shanxi Institute of Acoustics, Hubei Institute of Acoustics, Chinese Academy of Acoustics Ultrasound Electronics Committee. 2008: 4. Zhang, C., Zeng, X. & Wang, S. (2012). Endpoint detection based on spectral variance of critical band. Technical Acoustics, 2012, (02): 204–208. Zhang, H., & Hu, H.(2010). An endpoint detection algorithm based on MFCC and spectral entropy using BP NN, 2010 2nd International Conference on Signal Processing Systems, Dalian, 2010, pp. V2-509-V2-513. Zhang, Y., Wang, K., & Yan, B. (2016). Speech endpoint detection algorithm with low signal-to-noise based on improved conventional spectral entropy, 2016 12th World Congress on Intelligent Control and Automation (WCICA), Guilin, pp. 3307–3311.