English
Related papers

Related papers: Audio-Visual Speech Codecs: Rethinking Audio-Visua…

200 papers

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Yao Shi , Yunfei Xu , Hongbin Suo , Yulong Wan , Haifeng Liu

Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook essential para-linguistic cues, including emotion and…

Human-Computer Interaction · Computer Science 2025-06-23 Qixin Wang , Songtao Zhou , Zeyu Jin , Chenglin Guo , Shikun Sun , Xiaoyu Qin

Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in situations where audio signals are unavailable or degraded.…

Sound · Computer Science 2026-02-03 Jaejun Lee , Yoori Oh , Kyogu Lee

Automatic speech recognition (ASR) has gained remarkable successes thanks to recent advances of deep learning, but it usually degrades significantly under real-world noisy conditions. Recent works introduce speech enhancement (SE) as…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-19 Yuchen Hu , Chen Chen , Qiushi Zhu , Eng Siong Chng

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

Enhancing speech signal quality in adverse acoustic environments is a persistent challenge in speech processing. Existing deep learning based enhancement methods often struggle to effectively remove background noise and reverberation in…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Heming Wang , Meng Yu , Hao Zhang , Chunlei Zhang , Zhongweiyang Xu , Muqiao Yang , Yixuan Zhang , Dong Yu

In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio quality. While current models are capable of generating…

Sound · Computer Science 2026-05-14 Yicheng Gu , Junan Zhang , Chaoren Wang , Jerry Li , Zhizheng Wu , Lauri Juvela

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only approaches at low…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Ahsan Adeel , Mandar Gogate , Amir Hussain

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only…

Computation and Language · Computer Science 2020-02-19 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Neural audio codecs are at the core of modern conversational speech technologies, converting continuous speech into sequences of discrete tokens that can be processed by LLMs. However, existing codecs typically operate at fixed frame rates,…

Machine Learning · Computer Science 2026-02-05 Luca Della Libera , Cem Subakan , Mirco Ravanelli

This paper introduces an augmented reality (AR) captioning framework designed to support Deaf and Hard of Hearing (DHH) learners in STEM classrooms by integrating non-verbal emotional cues into live transcriptions. Unlike conventional…

Human-Computer Interaction · Computer Science 2025-04-29 Sunday David Ubur

Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR have improved performance in noisy environments compared to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-29 Zhaofeng Lin , Naomi Harte

Existing deep learning based speech enhancement mainly employ a data-driven approach, which leverage large amounts of data with a variety of noise types to achieve noise removal from noisy signal. However, the high dependence on the data…

Sound · Computer Science 2024-01-24 Huaying Xue , Xiulian Peng , Yan Lu

Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the speaker's mouth. The use of the video signal requires…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Dmitriy Serdyuk , Otavio Braga , Olivier Siohan

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial…

Image and Video Processing · Electrical Eng. & Systems 2025-06-17 Riku Takahashi , Ryugo Morita , Jinjia Zhou

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Daniel Michelsanti , Olga Slizovskaia , Gloria Haro , Emilia Gómez , Zheng-Hua Tan , Jesper Jensen