中文
相关论文

相关论文: Audio-Visual Speech Codecs: Rethinking Audio-Visua…

200 篇论文

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Audio-visual feature synchronization for real-time speech enhancement in hearing aids represents a progressive approach to improving speech intelligibility and user experience, particularly in strong noisy backgrounds. This approach…

音频与语音处理 · 电气工程与系统科学 2025-08-28 Nasir Saleem , Mandar Gogate , Kia Dashtipour , Adeel Hussain , Usman Anwar , Adewale Adetomi , Tughrul Arslan , Amir Hussain

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Humans are adept at leveraging visual cues from lip movements for recognizing speech in adverse listening conditions. Audio-Visual Speech Recognition (AVSR) models follow similar approach to achieve robust speech recognition in noisy…

音频与语音处理 · 电气工程与系统科学 2024-05-24 Maxime Burchi , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg , Radu Timofte

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the unique speaking styles of…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Weizhi Zhong , Jichang Li , Yinqi Cai , Ming Li , Feng Gao , Liang Lin , Guanbin Li

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Yihan Wu , Yichen Lu , Yifan Peng , Xihua Wang , Ruihua Song , Shinji Watanabe

Acoustical mismatch among training and testing phases degrades outstandingly speech recognition results. This problem has limited the development of real-world nonspecific applications, as testing conditions are highly variant or even…

声音 · 计算机科学 2013-05-13 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

Recent talking head synthesis works typically adopt speech features extracted from large-scale pre-trained acoustic models. However, the intrinsic many-to-many relationship between speech and lip motion causes phoneme-viseme alignment…

图形学 · 计算机科学 2025-10-16 Yihuan Huang , Jiajun Liu , Yanzhen Ren , Jun Xue , Wuyang Liu , Zongkun Sun

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip…

人工智能 · 计算机科学 2024-06-19 Young Jin Ahn , Jungwoo Park , Sangha Park , Jonghyun Choi , Kee-Eung Kim

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from…

End-to-end Automatic Speech Recognition (ASR) systems based on neural networks have seen large improvements in recent years. The availability of large scale hand-labeled datasets and sufficient computing resources made it possible to train…

计算机视觉与模式识别 · 计算机科学 2023-01-05 Maxime Burchi , Radu Timofte

Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various…

声音 · 计算机科学 2025-11-03 Jiarong Du , Zhan Jin , Peijun Yang , Juan Liu , Zhuo Li , Xin Liu , Ming Li

The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which…

声音 · 计算机科学 2024-09-30 Youngjae Kim , Yejin Jeon , Gary Geunbae Lee

Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Wei-Cheng Tseng , David Harwath

Lip sync is a fundamental audio-visual task. However, existing lip sync methods fall short of being robust in the wild. One important cause could be distracting factors on the visual input side, making extracting lip motion information…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Chun Wang

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Zexu Pan , Wupeng Wang , Shengkui Zhao , Chong Zhang , Kun Zhou , Yukun Ma , Bin Ma

This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Yusheng Tian , Jingyu Li , Tan Lee

Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. However, audio-visual…

声音 · 计算机科学 2024-08-13 HyoJung Han , Mohamed Anwar , Juan Pino , Wei-Ning Hsu , Marine Carpuat , Bowen Shi , Changhan Wang