English
Related papers

Related papers: ReVISE: Self-Supervised Speech Resynthesis with Vi…

200 papers

This brief literature review studies the problem of audiovisual speech synthesis, which is the problem of generating an animated talking head given a text as input. Due to the high complexity of this problem, we approach it as the…

Sound · Computer Science 2021-03-09 Efthymios Georgiou , Athanasios Katsamanis

Supervised speech enhancement relies on parallel databases of degraded speech signals and their clean reference signals during training. This setting prohibits the use of real-world degraded speech data that may better represent the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-22 Yangyang Xia , Buye Xu , Anurag Kumar

Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-28 Yi Ren , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

Machine learning models are known to perpetuate and even amplify the biases present in the data. However, these data biases frequently do not become apparent until after the models are deployed. Our work tackles this issue and enables the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Angelina Wang , Alexander Liu , Ryan Zhang , Anat Kleiman , Leslie Kim , Dora Zhao , Iroha Shirai , Arvind Narayanan , Olga Russakovsky

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only…

Computation and Language · Computer Science 2020-02-19 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Jiazhi Guan , Zhiliang Xu , Hang Zhou , Kaisiyuan Wang , Shengyi He , Zhanwang Zhang , Borong Liang , Haocheng Feng , Errui Ding , Jingtuo Liu , Jingdong Wang , Youjian Zhao , Ziwei Liu

This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-25 Mingyang Zhang , Yi Zhou , Yi Ren , Chen Zhang , Xiang Yin , Haizhou Li

Audio-visual speech recognition (AVSR) can effectively and significantly improve the recognition rates of small-vocabulary systems, compared to their audio-only counterparts. For large-vocabulary systems, however, there are still many…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-13 Wentao Yu , Steffen Zeiler , Dorothea Kolossa

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods…

Sound · Computer Science 2025-10-13 Ke Xue , Rongfei Fan , Lixin , Dawei Zhao , Chao Zhu , Han Hu

Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, such as lip movements and facial expressions. This task is notably challenging due to the absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Matthew Kit Khinn Teng , Haibo Zhang , Takeshi Saitoh

Speech resynthesis is a generic task for which we want to synthesize audio with another audio as input, which finds applications for media monitors and journalists.Among different tasks addressed by speech resynthesis, voice conversion…

Sound · Computer Science 2024-08-07 Thibault Gaudier , Marie Tahon , Anthony Larcher , Yannick Estève

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with model cascading.…

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper…

Sound · Computer Science 2026-02-03 Zhipeng Chen , Xinheng Wang , Lun Xie , Haijie Yuan , Hang Pan

This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from…

Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Yihan Wu , Yifan Peng , Yichen Lu , Xuankai Chang , Ruihua Song , Shinji Watanabe

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

Sound · Computer Science 2025-06-03 Nabarun Goswami , Tatsuya Harada

In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during listening and harming downstream tasks performance. To…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-03 Hsin-Tien Chiang , Hao Zhang , Yong Xu , Meng Yu , Dong Yu

We present a neural analysis and synthesis (NANSY) framework that can manipulate voice, pitch, and speed of an arbitrary speech signal. Most of the previous works have focused on using information bottleneck to disentangle analysis features…

Sound · Computer Science 2021-10-29 Hyeong-Seok Choi , Juheon Lee , Wansoo Kim , Jie Hwan Lee , Hoon Heo , Kyogu Lee

Speech enhancement improves speech quality and promotes the performance of various downstream tasks. However, most current speech enhancement work was mainly devoted to improving the performance of downstream automatic speech recognition…

Sound · Computer Science 2022-09-16 Jianrong Wang , Xiaomin Li , Xuewei Li , Mei Yu , Qiang Fang , Li Liu

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li
‹ Prev 1 4 5 6 7 8 10 Next ›