Collaborative steering of microphone array and video camera toward multi-lingual tele-conference through speech-to-speech translation
Tóm tắt
It is very important for multilingual teleconferencing through speech-to-speech translation to capture distant-talking speech with high quality. In addition, the speaker image is also needed to realize a natural communication in such a conference. A microphone array is an ideal candidate for capturing distant-talking speech. Uttered speech can be enhanced and speaker images can be captured by steering a microphone array and a video camera in the speaker direction. However, to realize automatic steering, it is necessary to localize the talker. To overcome this problem, we propose collaborative steering of the microphone array and the video camera in real-time for a multilingual teleconference through speech-to-speech translation. We conducted experiments in a real room environment. The speaker localization rate (i.e., speaker image capturing rate) was 97.7%, speech recognition rate was 90.0%, and TOEIC score was 530/spl sim/540 points, subject to locating the speaker at a 2.0 meter distance from the microphone array.
Từ khóa
#Microphone arrays #Collaboration #Cameras #Teleconferencing #Speech synthesis #Direction of arrival estimation #Loudspeakers #Acoustic noise #Working environment noise #Natural languagesTài liệu tham khảo
10.1109/TAP.1982.1142739
flanagan, 1993, Spatially selective sound capture for speech and audio processing, Speech Communication, 13, 207, 10.1016/0167-6393(93)90072-S
10.1109/ICASSP.2000.859144
10.1109/TASSP.1976.1162830
sugaya, 2000, Evaluation of the ATR-MATRIX Speech Translation System With a Pair Comparison Method Between the System and Humans, Proc ICSLP2000, 1105
nakamura, 2001, A Speech Translation System Applied to a Real-World Task/Domain and Its Evaluation Using Real-World Speech Data, IEICE Trans Inf & Syst, e84 d, 142
10.1007/978-1-4612-3632-0
10.1121/1.392786
