On the use of deep feedforward neural networks for automatic language identification

Computer Speech & Language - Tập 40 - Trang 46-59 - 2016
Ignacio Lopez-Moreno1, Javier Gonzalez-Dominguez2, David Martinez3, Oldřich Plchot4, Joaquin Gonzalez-Rodriguez2, Pedro J. Moreno1
1Google, Inc., New York, USA
2ATVS-Biometric Recognition Group, Universidad Autonoma de Madrid, Madrid, Spain
3I3A, Zaragoza, Spain
4Brno University of Technology, Brno, Czech Republic

Tài liệu tham khảo

Brummer, 2010 Brummer, 2012, Description and analysis of the Brno276 System for LRE2011, 216 Brümmer Brümmer, 2006, On calibration of language recognition scores Ciresan, D., Meier, U., Gambardella, L., Schmidhuber, J., 2010. Deep Big Simple Neural Nets Excel on Handwritten Digit Recognition, CoRR abs/1003.0358. Cole, 1989, Language identification with neural networks: a feasibility study, 525 Davis, 1980, Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences, IEEE Trans. Acoust. Speech Signal Process, 28, 357, 10.1109/TASSP.1980.1163420 Dean, 2012, Large scale distributed deep networks, 1232 Dehak, 2011, Language recognition via i-vectors and dimensionality reduction, 857 Dehak, 2011, Front-end factor analysis for speaker verification, IEEE Trans. Audio Speech Lang. Process, 19, 788, 10.1109/TASL.2010.2064307 Ferrer, 2010, A comparison of approaches for modeling prosodic features in speaker recognition, 4414 Fontaine, 1997, Nonlinear discriminant analysis for improved speech recognition Gonzalez-Dominguez, 2014, A real-time end-to-end multilingual speech recognition architecture, IEEE J. Sel. Top. Signal Process, 1 Gonzalez-Dominguez, 2015, Frame-by-frame language identification in short utterances using deep neural networks, Neural Netw, 64, 49, 10.1016/j.neunet.2014.08.006 Grézl, 2007, Probabilistic and bottle-neck features for LVCSR of meetings, 757 Grézl, 2009, Investigation into bottle-neck features for meeting speech recognition, 2947 Hermansky, 1994, RASTA processing of speech, IEEE Trans. Speech Audio Process, 2, 578, 10.1109/89.326616 Hermansky, 2000, Tandem connectionist feature extraction for conventional HMM systems, 1635 Hinton, 2012, Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups, IEEE Signal Process. Mag, 29, 82, 10.1109/MSP.2012.2205597 Jiang, 2014, Deep bottleneck features for spoken language identification, PLoS ONE, 9, e100795, 10.1371/journal.pone.0100795 Kenny Kenny, 2008, A study of interspeaker variability in speaker verification, IEEE Trans. Audio Speech Lang. Process, 16, 980, 10.1109/TASL.2008.925147 Leena, 2005, Neural network classifiers for language identification using phonotactic and prosodic features, 404 Li, 2013, Spoken language recognition: from fundamentals to practice, P. IEEE, 101, 1136, 10.1109/JPROC.2012.2237151 Lopez-Moreno, I., Gonzalez-Dominguez, J., Plchot, O., Martinez, D., Gonzalez-Rodriguez, J., Moreno, P., 2014. Automatic language identification using deep neural networks. Acoustics, Speech, and Signal Processing, IEEE International Conference on. Martínez, 2011, I3A language recognition system description for NIST LRE 2011 Martínez, 2013, Dysarthria intelligibility assessment in a factor analysis total variability space, 2133 Martinez, 2011, Language recognition in iVectors space, 861 Matĕjka, 2014, Neural network bottleneck features for language identification, 299 McCree, 2014, Multiclass discriminative training of i-vector language recognition Mohamed, 2012, Acoustic modeling using deep belief networks, IEEE Trans. Audio Speech Lang. Process, 20, 14, 10.1109/TASL.2011.2109382 Mohamed, 2012, Understanding how deep belief networks perform acoustic modelling, 4273 Montavon, 2009, Deep learning for spoken language identification Muthusamy, 1994, Reviewing automatic language identification, IEEE Signal Process. Mag, 11, 33, 10.1109/79.317925 NIST Reynolds, 2000, Speaker verification using adapted Gaussian mixture models, Dig. Sig. Process, 10, 19, 10.1006/dspr.1999.0361 Reynolds, 2003, The SuperSID project: exploiting high-level information for high-accuracy speaker recognition, vol. 4, 784 Richardson Sturim, 2011, The MIT LL 2010 speaker recognition evaluation system: scalable language-independent speaker recognition, 5272 Torres-Carrasquillo, 2002, Approaches to language identification using Gaussian mixture models and shifted delta cepstral features, vol. 1, 89 Welling, 1999, Improved methods for vocal tract normalization, 761 Xia, 2012, Using i-vector space model for emotion recognition, 2230 Yu, 2011, Deep learning and its applications to signal and information processing [exploratory DSP], IEEE Signal Process. Mag, 28, 145, 10.1109/MSP.2010.939038 Zeiler, 2013, On rectified linear units for speech processing