Speech recognition using advanced HMM2 features

K. Weber1,2, S. Bengio1, H. Bourlard1,2
1Dalle Molle Institute for Perceptual Artificial Intelligence, Martigny, Switzerland
2EPFL-Swiss Federal Institute of Technology, Lausanne, Lausanne, Switzerland

Tóm tắt

HMM2 is a particular hidden Markov model where state emission probabilities of the temporal (primary) HMM are modeled through (secondary) state-dependent frequency-based HMMs (see Weber, K. et al., Proc. ICSGP, vol.III, p.147-50, 2000). As we show in another paper (see Weber et al., Proc. Eurospeech, Sep. 2001), a secondary HMM can also be used to extract robust ASR features. Here, we further investigate this novel approach towards using a full HMM2 as feature extractor, working in the spectral domain, and extracting robust formant-like features for a standard ASR system. HMM2 performs a nonlinear, state-dependent frequency warping, and it is shown that the resulting frequency segmentation actually contains particularly discriminant features. To improve the HMM2 system further, we complement the initial spectral energy vectors with frequency information. Finally, adding temporal information to the HMM2 feature vector yields further improvements. These conclusions are experimentally validated on the Numbers95 database, where word error rates of 15%, using only a 4-dimensional feature vector (3 formant-like parameters and one time index) were obtained.

Từ khóa

#Speech recognition #Hidden Markov models #Frequency #Feature extraction #Robustness #Automatic speech recognition #Data mining #Spatial databases #Error analysis #Indexes

Tài liệu tham khảo

nadeu, 1999, On the Filter-bank-based Parameterization Front-End for Robust HMM Speech Recognition, Proc Robust'99, 235 samaria, 1994, Face Recognition Using Hidden Markov Models weber, 2000, HMM2- a novel approach to HMM emission probability estimation, Proc ICSLP, iii, 147 weber, 2001, HMM2- Extraction of Formant Structures and their Use for Robust ASR, Proc EUROSPEECH weber, 2001, A Pragmatic View of the Application of HMM2 for ASR, IDIAP-RR 01–23 young, 1995, The HTK Book, Cambridge University eickeler, 1999, High Performance Face Recognition Using Pseudo 2D-Hidden Markov Models, European Control Conference (ECC) 10.1016/0165-1684(92)90112-A holmes, 2000, Segmental HMMs: Modelling Dynamics and Underlying Structure for Automatic Speech Recognition, IMA Workshop Mathematical Foundations of Speech Processing and Recognition 10.1109/ICASSP.1998.674352 10.1109/TASSP.1986.1164908 konig, 1994, Modeling Dynamics in Connectionist Speech Recognition - The Time Index Model, ICSI TR-94–012 cole, 1995, New Telephone Speech Corpora at CSLU, Proc EUROSPEECH, 1, 821 bengio, 2000, An EM Algorithm for HMMs with Emission Distributions Represented by HMMs, IDIAP-RR 00–11 10.1109/ICASSP.1993.319752