Investigations on the combination of four algorithms to increase the noise robustness of a DSR front-end for real world car data

B. Andrassy1, F. Hilger2, C. Beaugeant1
1Siemens AG, Munich, Germany
2Lehrnuhl für Informatik VI, RWnl Aaehen-University of Technology, Aachen, Germany

Tóm tắt

This paper shows how the noise robustness of a MFCC feature extraction front-end can be improved by integrating four noise robustness algorithms:a spectral attenuation, a noise level normalisation, a cepstral mean normalization and a frame dropping algorithm. The algorithms were tested separately and in varying combinations on three real world car data sets with different amounts of mismatch between the training and the testing conditions. It was shown that although the algorithms partly have similar effects none of them is completely redundant. Every algorithm can contribute to a further improvement of the recognition results so the best results can be achieved by a combination of all four of them. A relative reduction of the word error rate of up to 57% is achieved.

Từ khóa

#Noise robustness #Noise level #Attenuation #Speech recognition #Speech enhancement #Filter bank #Integrated circuit noise #Testing #Distributed computing #Wiener filter

Tài liệu tham khảo

hirsch, 2000, The Aurora Experimental Framework for the Performance Evaluation of Speech Recognition Systems under Noisy Conditions, ASR2000 -Intemational Workshop on Automatic Speech Recognition, 181 hilger, 2000, Noise Level Normalization and Reference Adaptation for Robust Speech Recognition, Proc ASR2000 Int Workshop on Automatic Speech Recognition, 64 lindberg, 2001, Danish SpeechDat-Car Database for ETSI STQ Aurora Advanced DSR, documentation on the Danish Aurora Project Database CD-ROMs moreno, 2000, SpeechDat Car: Spanish Selected Subset, Ver.1, Universitat Politecnica de Catalunya Spain netsch, 2001, Description and Baseline Results for the Subset of the SpeechDat-Car German Database used for ETSI STQ Aurora WI008 Advanced Front-End Evaluation, documentation on the German Aurora Project Database CD-ROMs lim, 1983, Speech enhancement Prentice-Hall Signal Processing Series Alan V Oppenheim 2000, ETSI ES 201 108 vl.1.2 Distributed Speech Recognition; Front-end Feature Extraction Algorithm; Compression Algorithm