Chan, W., Jaitly, N., Le, Q., and Vinyals, O., Listen, attend and spell: A neural network for large vocabulary conversational speech recognition, IEEE International Conference on Acoustics, Speech and Signal Processing, Shanghai, 2016, pp. 4960–4964.
Wang, Y., Li, J. and Gong, Y., Small-footprint high-performance deep neural network-based speech recognition using split-VQ, IEEE International Conference on Acoustics, Speech and Signal Processing, 2015, pp. 4984–4988.
Wu, C., Karanasou, P., Gales, M.J.F., and Sim K.C., Stimulated deep neural network for speech recognition, in Interspeech, San Francisco, 2016, pp. 400–404.
Graves, A., Mohamed, A.R. and Hinton, G., Speech recognition with deep recurrent neural networks, IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, 2013, pp. 6645–6649.
Salvador, S.W. and Weber, F.V., US Patent 9 153 231, 2015.
Cai, J., Li, F., Zhang, Y., and Liu, Y., Research on multi-base depth neural network speech recognition, Advanced Information Technology, Electronic and Automation Control Conference, Chongqing, 2017, pp. 1540–1544.
Chorowski, J., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y., Attention-based models for speech recognition, Comput. Sci., 2015, vol. 10, no. 4, pp. 429–439.
Miao, Y., Gowayyed, M., and Metze, F., EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding, Automatic Speech Recognition & Understanding, Scottsdale, 2015, pp. 167–174.
Schwarz, A., Huemmer, C., Maas, R. and Kellermann, W., Spatial diffuseness features for DNN-based speech recognition in noisy and reverberant environments, IEEE International Conference on Acoustics, Speech and Signal Processing, 2015, pp. 4380–4384.
Kipyatkova, I., Experimenting with hybrid TDNN/HMM acoustic models for Russian speech recognition, Speech and Computer: 19th International Conference, 2017, pp. 362–369.
Yoshioka, T., Karita, S. and Nakatani, T., Far-field speech recognition using CNN-DNN-HMM with convolution in time’, IEEE International Conference on Acoustics, Speech and Signal Processing, Brisbane, 2015, pp. 4360–4364.
Wang, Y., Bao, F., Zhang, H. and Gao, G.L., Research on Mongolian speech recognition based on FSMN, Natural Language Processing and Chinese Computing, 2017, pp. 243–254.
Alam, M.J., Gupta, V., Kenny, P., and Dumouchel, P., Speech recognition in reverberant and noisy environments employing multiple feature extractors and i-vector speaker adaptation’, EURASIP J. Adv. Signal Process., 2015, vol. 2015, no. 1, p. 50.
Brayda, L., Wellekens, C., and Omologo, M., N-best parallel maximum likelihood beamformers for robust speech recognition, Signal Processing Conference, Florence, 2015, pp. 1–4.
Ali, A., Zhang, Y., Cardinal, P., Dahak, N., Vogel, S., and Glass, J.R., A complete KALDI recipe for building Arabic speech recognition systems, 2014 Spoken Language Technology Workshop, South Lake Tahoe, NV, 2015, pp. 525–529.