中文
相关论文

相关论文: End-to-End Lip Synchronisation Based on Pattern Cl…

200 篇论文

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead…

图形学 · 计算机科学 2022-12-09 Zhentao Yu , Zixin Yin , Deyu Zhou , Duomin Wang , Finn Wong , Baoyuan Wang

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Xu Wang , Shengeng Tang , Fei Wang , Lechao Cheng , Dan Guo , Feng Xue , Richang Hong

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

声音 · 计算机科学 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

We present a deep learning framework for real-time speech-driven 3D facial animation from just raw waveforms. Our deep neural network directly maps an input sequence of speech audio to a series of micro facial action unit activations and…

计算机视觉与模式识别 · 计算机科学 2017-12-11 Hai X. Pham , Yuting Wang , Vladimir Pavlovic

Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent layers to tackle voice activity detection directly from the…

音频与语音处理 · 电气工程与系统科学 2020-05-27 Marvin Lavechin , Marie-Philippe Gill , Ruben Bousbib , Hervé Bredin , Leibny Paola Garcia-Perera

Time-frequency masking or spectrum prediction computed via short symmetric windows are commonly used in low-latency deep neural network (DNN) based source separation. In this paper, we propose the usage of an asymmetric analysis-synthesis…

音频与语音处理 · 电气工程与系统科学 2021-06-23 Shanshan Wang , Gaurav Naithani , Archontis Politis , Tuomas Virtanen

The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence. This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses,…

计算与语言 · 计算机科学 2022-04-07 Yogesh Virkar , Marcello Federico , Robert Enyedi , Roberto Barra-Chicote

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with…

人工智能 · 计算机科学 2026-01-07 Zeyu Ling , Xiaodong Gu , Jiangnan Tang , Changqing Zou

Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate the likeness of voices and faces following contrastive…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Chong Peng , Liqiang He , Dan Su

In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units(CPUs) and Graphics Processing Units (GPUs).…

In this paper, we propose a novel end-to-end neural-network-based speaker diarization method. Unlike most existing methods, our proposed method does not have separate modules for extraction and clustering of speaker representations.…

音频与语音处理 · 电气工程与系统科学 2019-09-16 Yusuke Fujita , Naoyuki Kanda , Shota Horiguchi , Kenji Nagamatsu , Shinji Watanabe

We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Connectionist Temporal…

计算与语言 · 计算机科学 2016-11-30 Chris Lengerich , Awni Hannun

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Deep learning techniques have achieved specific results in recording device source identification. The recording device source features include spatial information and certain temporal information. However, most recording device source…

声音 · 计算机科学 2022-12-06 Chunyan Zeng , Dongliang Zhu , Zhifeng Wang , Minghu Wu , Wei Xiong , Nan Zhao

This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identification for monaural multi-talker audio. Firstly, we…

音频与语音处理 · 电气工程与系统科学 2021-04-07 Naoyuki Kanda , Guoli Ye , Yashesh Gaur , Xiaofei Wang , Zhong Meng , Zhuo Chen , Takuya Yoshioka

We propose a multitask training method for attention-based end-to-end speech recognition models. We regularize the decoder in a listen, attend, and spell model by multitask training it on both audio-text and text-only data. Trained on the…

计算与语言 · 计算机科学 2021-06-15 Peidong Wang , Tara N. Sainath , Ron J. Weiss

Acoustic Echo Cancellation (AEC) whose aim is to suppress the echo originated from acoustic coupling between loudspeakers and microphones, plays a key role in voice interaction. Linear adaptive filter (AF) is always used for handling this…

声音 · 计算机科学 2021-06-01 Lu Ma , Song Yang , Yaguang Gong , Xintian Wang , Zhongqin Wu

The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the…

多媒体 · 计算机科学 2025-12-01 Yuyue Wang , Xin Cheng , Yihan Wu , Xihua Wang , Jinchuan Tian , Ruihua Song

In this paper, we investigate a deep learning approach for speech denoising through an efficient ensemble of specialist neural networks. By splitting up the speech denoising task into non-overlapping subproblems and introducing a…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Aswin Sivaraman , Minje Kim