English
Related papers

Related papers: Deep Multimodal Learning for Audio-Visual Speech R…

200 papers

Stream fusion, also known as system combination, is a common technique in automatic speech recognition for traditional hybrid hidden Markov model approaches, yet mostly unexplored for modern deep neural network end-to-end model…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-15 Timo Lohrenz , Zhengyang Li , Tim Fingscheidt

The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-19 Peter Wu , Bohan Yu , Kevin Scheck , Alan W Black , Aditi S. Krishnapriyan , Irene Y. Chen , Tanja Schultz , Shinji Watanabe , Gopala K. Anumanchipalli

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

Utilizing the sensor characteristics of the audio, visible camera, and thermal camera, the robustness of person recognition can be enhanced. Existing multimodal person recognition frameworks are primarily formulated assuming that multimodal…

Multimedia · Computer Science 2022-10-25 Vijay John , Yasutomo Kawanishi

Neural speech separation has made remarkable progress and its integration with automatic speech recognition (ASR) is an important direction towards realizing multi-speaker ASR. This work provides an insightful investigation of speech…

Speech enhancement can potentially benefit from the visual information from the target speaker, such as lip movement and facial expressions, because the visual aspect of speech is essentially unaffected by acoustic environment. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-24 Xinmeng Xu , Jianjun Hao

Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performance of tasks like…

Sound · Computer Science 2024-04-24 Donghuo Zeng , Yanan Wang , Kazushi Ikeda , Yi Yu

Visual speech recognition (VSR) is the task of recognizing spoken language from video input only, without any audio. VSR has many applications as an assistive technology, especially if it could be deployed in mobile devices and embedded…

Computation and Language · Computer Science 2019-06-06 Nilay Shrivastava , Astitwa Saxena , Yaman Kumar , Rajiv Ratn Shah , Debanjan Mahata , Amanda Stent

We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known that ensemble methods…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-04 Marin Jezidžić , Matej Mihelčić

Multimodal deep learning methods capture synergistic features from multiple modalities and have the potential to improve accuracy for stress detection compared to unimodal methods. However, this accuracy gain typically comes from high…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Morteza Bodaghi , Majid Hosseini , Raju Gottumukkala

Self-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed out the strength of integrating SSLR with single-channel…

Sound · Computer Science 2022-10-20 Yoshiki Masuyama , Xuankai Chang , Samuele Cornell , Shinji Watanabe , Nobutaka Ono

Distance Metric Learning (DML) has typically dominated the audio-visual speaker verification problem space, owing to strong performance in new and unseen classes. In our work, we explored multitask learning techniques to further enhance…

Sound · Computer Science 2024-09-25 Anith Selvakumar , Homa Fashandi

Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to learn representations from unpaired modalities remains…

Sound · Computer Science 2025-04-15 Yasar Abbas Ur Rehman , Kin Wai Lau , Yuyang Xie , Ma Lan , JiaJun Shen

Robust semantic perception for autonomous vehicles relies on effectively combining multiple sensors with complementary strengths and weaknesses. State-of-the-art sensor fusion approaches to semantic perception often treat sensor data…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tim Broedermannn , Christos Sakaridis , Luigi Piccinelli , Wim Abbeloos , Luc Van Gool

Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or…

Sound · Computer Science 2025-07-11 Sidong Zhang , Shiv Shankar , Trang Nguyen , Andrea Fanelli , Madalina Fiterau

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the…

With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data fusion and the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Wei Dai , Dequan Zheng , Feng Yu , Yanrong Zhang , Yaohui Hou

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

Recently, many novel techniques have been introduced to deal with spoofing attacks, and achieve promising countermeasure (CM) performances. However, these works only take the stand-alone CM models into account. Nowadays, a spoofing aware…

Sound · Computer Science 2022-03-30 Haibin Wu , Lingwei Meng , Jiawen Kang , Jinchao Li , Xu Li , Xixin Wu , Hung-yi Lee , Helen Meng