MPE-based discriminative linear transforms for speaker adaptation

Computer Speech & Language - Tập 22 - Trang 256-272 - 2008
Lan Wang1, Philip C. Woodland1
1Cambridge University Engineering Department, Trumpington Street, Cambridge CB2 1PZ, UK

Tài liệu tham khảo

Anastasakos, T., Balakrishnan, S.V., 1998. The use of confidence measures in unsupervised adaptation of speech recognizers. In: Proceedings of the ICSLP’98. Sydney, Australia, pp. 2303–2306. Evermann, G., Woodland, P.C., 2000. Large vocabulary decoding and confidence estimation using word posterior probabilities. In: Proceedings of the ICASSP’00. Istanbul, Turkey, pp. 1655–1659. Gales, 1998, Maximum likelihood linear transformations for HMM-based speech recognition, Computer Speech and Language, 12, 75, 10.1006/csla.1998.0043 Gales, 1996, Mean and variance adaptation within the MLLR framework, Computer Speech and Language, 10, 10.1006/csla.1996.0013 Gunawardana, A., Byrne, W., 2001. Discriminative speaker adaptation with conditional maximum likelihood linear regression. In: Proceedings of the Eurospeech’01, pp. 1203–1206. Gunawardana, 2002 Kumar, N., 1997. Investigation of Silicon-Auditory Models and Generalization of Linear Discriminant Analysis for Improved Speech Recognition. Ph.D. Dissertation, John Hopkins University, USA. Leggetter, C.J., Woodland, P.C., 1994. Speaker Adaptation of HMM Using Linear Regression. CUED/F-INFENG/TR181, Department of Engineering, University of Cambridge. Leggetter, 1995, Maximum likelihood linear regression for speaker adaptation of continuous Hidden Markov models, Computer Speech and Language, 9, 171, 10.1006/csla.1995.0010 Leggetter, C.J., Woodland, P.C., 1995b. Flexible Speater adaptation for large uocabulary speech recognition. In: Proceedings of the Eulope Conference on Speech Communication and Technology, vol. 2. Madrid, Spain, pp. 1155–1158. Mangu, 2000, Finding consensus in speech recognition: word error minimization and other applications of confusion network, Computer Speech and Language, 14, 373, 10.1006/csla.2000.0152 McDonough, J., Waibel, A., 2003. Maximum mutual information speaker adaptive training with semi-tied covariance matrices. In: Proceedings of the ICASSP’03. Hong Kong, China, pp. 128–131. McDonough, J., Schaaf, T., Waibel, A., 2002. On maximum mutual information speaker-adapted training. In: Proceedings of the ICASSP’02. Florida, USA, pp. 601–604. Povey, D., 2004. Discriminative training for large vocabulary speech recognition. Ph.D. Dissertation, Department of Engineering, University of Cambridge, UK, 2004. Povey, D., Woodland, P.C., 2002. Minimum phone error and i-smoothing for improved discriminative training. In: Proceedings of the ICASSP’02, vol. 1. Florida, USA, pp. 105–108. Uebel, L.F., Woodland, P.C., 2001. Speaker adaptation using lattice-based MLLR. In: Proceedings of ISCA ITRW Adaptation Methods for Automatic Speech Recognition. Sophia-Antipolis, France, pp. 57–60. Uebel, L.F., Woodland, P.C., 2001. Discriminative linear transforms for speaker adaptation. In: Proceedings of ISCA ITRW Adaptation Methods for Automatic Speech Recognition. Sophia-Antipolis, France, pp. 61–63. Wang, L., 2006. Discriminative linear transforms for adaptation and adaptive training. Ph.D. Dissertation, Department of Engineering, University of Cambridge, UK. Woodland, 2002, Large scale discriminative training of hidden markov models for speech recognition, Computer Speech and Language, 16, 25, 10.1006/csla.2001.0182 Woodland, P.C., Gales, M., Hain, T. et al., 2003. CU-HTK STT Systems for RT03. Rich Transcription Workshop 2003, Boston, USA, 2003. Zeppenfeld, T., Finke, M., Ries, K., Westphal, M., Waibel, A., 1997. Recognition of conversational telephone speech using the JANUS speech engine. In: Proceedings of the ICASSP’97. Munich, Germany, pp. 1815–1818.