Multi-modal simultaneous machine translation fusion of image information

Journal of Engineering Research - Tập 11 - Trang 100085 - 2023
Yan Huang1, Zhanyang Wanga1, TianYuan Zhang1, Chun Xu2, Hui Lianga1
1College of Software Engineering, Zhengzhou University of Light Industry, Zhengzhou, Henan, China
2College of Computer, Xinjiang University of Finance and Economics, Urumqi, Xinjiang, China

Tài liệu tham khảo

Young, 2014, From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions, TACL, 2, 67, 10.1162/tacl_a_00166 A. Grissom II, H. He, J. Boyd-Graber, et al., Don’t until the final verb wait: Reinforcement learning for simultaneous machine translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 2014, pp. 1342–52. Y. Oda, G. Neubig, S. Sakti, et al., Syntax-based simultaneous translation through prediction of unseen syntactic constituents, in: Proceedings of the 53th Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, vol. 1, Beijing, China, 2015, pp. 198–207. M. Ma, L. Huang, H. Xiong, et al., STACL: simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Firenze, Italy, 2019, pp. 3025–36. Seeber, 2019, Simultaneous interpreting, Routledge Handb. Interpret., 91 B. Zheng, R. Zheng, M. Ma, et al., Simultaneous translation with flexible policy via restricted imitation learning, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Firenze, Italy, 2019, pp. 5816–22. A. Alinejad, M. Siahbani, A. Sarkar, Prediction improves simultaneous neural machine translation, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 2018, pp. 3022–7. S. Matsubara, K. Iwashima, N. Kawaguchi, et al., Simultaneous Japanese-English interpretation based on early prediction of English verb, in: Proceedings of the Fourth Symposium on Natural Language Processing, Chiang Mai, Thailand, 2000, pp. 268–73. J. Gu, G. Neubig, K. Cho, et al., Learning to translate in real-time with neural machine translation, in: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Valencia, Spain, 2017, pp. 1053–62. I. Calixto, Q. Liu, N. Campbell, Doubly-attentive decoder for multi-modal neural machine translation, in: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Vancouver, Canada, 2017, pp. 1913–24. J. Libovicky, J. Helcl, Attention strategies for multi-source sequence-to-sequence learning, in: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, vol. 2, Vancouver, Canada, 2017, pp. 196–202. I. Calixto, Q. Liu, N. Campbell, Incorporating global visual features into attention-based neural machine translation, in: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark, 2017, pp. 992–1003. D. Elliott, A. Kádár, Imagination improves multimodal translation, in: Proceedings of the Eighth International Joint Conference on Natural Language Processing, vol. 1, Taipei, Taiwan, 2017, pp. 130–41. M. Zhou, R. Cheng, Y.J. Lee, et al., Visual attention grounding neural model for multimodal machine translation, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 2018, pp. 3643–53. L. Specia, S. Frank, K. Sima’An, et al., A shared task on multimodal machine translation and cross-lingual image description, in: Proceedings of the First Conference on Machine Translation, vol. 2, Berlin, Germany, 2016, pp. 543–53. C. Lala, L. Specia, Multimodal lexical translation, in: Proceedings of the Eleventh International Conference on Language Resources and Evaluation, Miyazaki, Japan, 2018, pp. 3810–7. S. Gella, D. Elliott, F. Keller, Cross-lingual visual verb sense disambiguation, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, Hong Kong, China, 2019, pp. 1998–2004. O. Caglayan, P. Madhyastha, L. Specia, et al., Probing the need for visual context in multimodal machine translation, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, Minneapolis, USA, 2019, pp. 4159–70. J. Hitschler, S. Schamoni, S. Riezler, Multimodal pivots for image caption translation, in: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Berlin, Germany, 2016, pp. 2399–409. R. Sennrich, B. Haddow, A. Birch, Neural Machine Translation of Rare Words with Subword Units, Computer Science, 2015. Vaswani, 2017, Attention is all you need, Adv. Neural Inf. Process. Syst., 30, 5998 Y. Miao, P. Blunsom, L. Specia, A generative framework for simultaneous machine translation, in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2016, pp. 6697–706. (Online). A. Alinejad, Shavarani, S. H, A. Sarkar, Translation-based supervision for policy generation in simultaneous neural machine translation, in: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 1734–44. (Online). K. Papineni, S. Roukos, T. Ward, et al., BLEU: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, USA, 2022, pp. 311–318. I. Provilkov, D. Emelianenko, E. Voita, BPE-dropout: simple and effective subword regularization, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 1882–92. (Online).