中文
相关论文

相关论文: Visual Speech Enhancement Without A Real Visual St…

200 篇论文

Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap…

计算与语言 · 计算机科学 2025-11-26 Wesley Bian , Xiaofeng Lin , Guang Cheng

Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements. This makes the lip reading models…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Minsu Kim , Hyunjun Kim , Yong Man Ro

Recent advances in diffusion models have led to significant progress in audio-driven lip synchronization. However, existing methods typically rely on constrained audio-visual alignment priors or multi-stage learning of intermediate…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Junxian Ma , Shiwen Wang , Jian Yang , Junyi Hu , Jian Liang , Guosheng Lin , Jingbo chen , Kai Li , Yu Meng

Generating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion…

In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancement such as the L2…

音频与语音处理 · 电气工程与系统科学 2019-07-08 Suwon Shon , Hao Tang , James Glass

Text-to-speech(TTS) has undergone remarkable improvements in performance, particularly with the advent of Denoising Diffusion Probabilistic Models (DDPMs). However, the perceived quality of audio depends not solely on its content, pitch,…

音频与语音处理 · 电气工程与系统科学 2024-04-23 Huadai Liu , Rongjie Huang , Xuan Lin , Wenqiang Xu , Maozong Zheng , Hong Chen , Jinzheng He , Zhou Zhao

Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape, color, texture),…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Yu Li , Feng Xue , Shujie Li , Jinrui Zhang , Shuang Yang , Dan Guo , Richang Hong

High-quality AI-powered video dubbing demands precise audio-lip synchronization, high-fidelity visual generation, and faithful preservation of identity and background. Most existing methods rely on a mask-based training strategy, where the…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Xindi Zhang , Dechao Meng , Steven Xiao , Qi Wang , Peng Zhang , Bang Zhang

The optimization of a wavelet-based algorithm to improve speech intelligibility along with the full data set and results are reported. The discrete-time speech signal is split into frequency sub-bands via a multi-level discrete wavelet…

声音 · 计算机科学 2022-07-25 Tianqu Kang , Anh-Dung Dinh , Binghong Wang , Tianyuan Du , Yijia Chen , Kevin Chau

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

声音 · 计算机科学 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

We aim to edit the lip movements in talking video according to the given speech while preserving the personal identity and visual details. The task can be decomposed into two sub-problems: (1) speech-driven lip motion generation and (2)…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Runyi Yu , Tianyu He , Ailing Zhang , Yuchi Wang , Junliang Guo , Xu Tan , Chang Liu , Jie Chen , Jiang Bian

Recent studies have shown impressive performance in Lip-to-speech synthesis that aims to reconstruct speech from visual information alone. However, they have been suffering from synthesizing accurate speech in the wild, due to insufficient…

声音 · 计算机科学 2023-02-20 Minsu Kim , Joanna Hong , Yong Man Ro

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods…

声音 · 计算机科学 2025-10-13 Ke Xue , Rongfei Fan , Lixin , Dawei Zhao , Chao Zhu , Han Hu

Achieving robust speech separation for overlapping speakers in various acoustic environments with noise and reverberation remains an open challenge. Although existing datasets are available to train separators for specific scenarios, they…

声音 · 计算机科学 2024-08-30 Ke Chen , Jiaqi Su , Taylor Berg-Kirkpatrick , Shlomo Dubnov , Zeyu Jin

We describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-speech facial movements…

音频与语音处理 · 电气工程与系统科学 2020-05-29 Ahmed Hussen Abdelaziz , Barry-John Theobald , Paul Dixon , Reinhard Knothe , Nicholas Apostoloff , Sachin Kajareker

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has demonstrated…

音频与语音处理 · 电气工程与系统科学 2025-02-11 Jing-Xuan Zhang , Tingzhi Mao , Longjiang Guo , Jin Li , Lichen Zhang

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would…

计算与语言 · 计算机科学 2018-02-16 Kaizhi Qian , Yang Zhang , Shiyu Chang , Xuesong Yang , Dinei Florencio , Mark Hasegawa-Johnson

This paper introduces a novel approach to Visual Forced Alignment (VFA), aiming to accurately synchronize utterances with corresponding lip movements, without relying on audio cues. We propose a novel VFA approach that integrates a local…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Yi He , Lei Yang , Shilin Wang

This paper studies the task of speech reconstruction from ultrasound tongue images and optical lip videos recorded in a silent speaking mode, where people only activate their intra-oral and extra-oral articulators without producing sound.…

音频与语音处理 · 电气工程与系统科学 2023-04-13 Rui-Chen Zheng , Yang Ai , Zhen-Hua Ling

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

计算机视觉与模式识别 · 计算机科学 2019-09-04 Chenhao Wang
‹ 上一页 1 8 9 10 下一页 ›