English
Related papers

Related papers: SAV-SE: Scene-aware Audio-Visual Speech Enhancemen…

200 papers

Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowledge, such as…

Computation and Language · Computer Science 2026-01-26 Yiming Rong , Yixin Zhang , Ziyi Wang , Deyang Jiang , Yunlong Zhao , Haoran Wu , Shiyu Zhou , Bo Xu

Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by image techniques, where these two views are similarly…

Automatic speaker verification (ASV) technology is recently finding its way to end-user applications for secure access to personal data, smart services or physical facilities. Similar to other biometric technologies, speaker verification is…

Sound · Computer Science 2016-09-16 Cemal Hanilci , Tomi Kinnunen , Md Sahidullah , Aleksandr Sizov

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

Sound · Computer Science 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal…

Sound · Computer Science 2023-06-02 Juan F. Montesinos , Daniel Michelsanti , Gloria Haro , Zheng-Hua Tan , Jesper Jensen

Advances in automatic speaker verification (ASV) promote research into the formulation of spoofing detection systems for real-world applications. The performance of ASV systems can be degraded severely by multiple types of spoofing attacks,…

Sound · Computer Science 2024-08-27 Zhenyu Wang , John H. L. Hansen

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up without compromising…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Sungwoo Cho , Se-Young Yun

Speaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve the system's…

Sound · Computer Science 2024-04-09 He Wang , Pengcheng Guo , Pan Zhou , Lei Xie

Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on exploring downstream…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Muyi Sun , Yixuan Wang , Hong Wang , Chen Su , Man Zhang , Xingqun Qi , Qi Li , Zhenan Sun

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-19 Ziyi Ni , Minglun Han , Feilong Chen , Linghui Meng , Jing Shi , Pin Lv , Bo Xu

In this paper, we are interested in unsupervised (unknown noise) audio-visual speech enhancement based on variational autoencoders (VAEs), where the probability distribution of clean speech spectra is simulated using an encoder-decoder…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-10 Mostafa Sadeghi , Xavier Alameda-Pineda

Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time video-to-audio…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Benjamin Klein , Kazi Ruslan Rahman , Sanchita Ghose

In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squared Error (MSE) are widely used; however, they often correlate poorly with perceptual quality and…

Sound · Computer Science 2026-03-18 Chih-Ning Chen , Jen-Cheng Hou , Hsin-Min Wang , Shao-Yi Chien , Yu Tsao , Fan-Gang Zeng

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 George Sterpu , Christian Saam , Naomi Harte

Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted…

Sound · Computer Science 2025-08-25 Wei Wang , Wangyou Zhang , Chenda Li , Jiatong Shi , Shinji Watanabe , Yanmin Qian

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomings; the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Asmar Nadeem , Adrian Hilton , Robert Dawes , Graham Thomas , Armin Mustafa
‹ Prev 1 8 9 10 Next ›