中文
相关论文

相关论文: Audio-Driven Talking Face Video Generation with Dy…

200 篇论文

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

计算机视觉与模式识别 · 计算机科学 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Audio-driven talking head generation is a core component of digital avatars, and 3D Gaussian Splatting has shown strong performance in real-time rendering of high-fidelity talking heads. However, achieving precise control over fine-grained…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Shaoyang Xie , Xiaofeng Cong , Baosheng Yu , Zhipeng Gui , Jie Gui , Yuan Yan Tang , James Tin-Yau Kwok

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

Neural audio codecs are at the core of modern conversational speech technologies, converting continuous speech into sequences of discrete tokens that can be processed by LLMs. However, existing codecs typically operate at fixed frame rates,…

机器学习 · 计算机科学 2026-02-05 Luca Della Libera , Cem Subakan , Mirco Ravanelli

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an…

音频与语音处理 · 电气工程与系统科学 2021-07-23 Sefik Emre Eskimez , You Zhang , Zhiyao Duan

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

计算机视觉与模式识别 · 计算机科学 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

Deep neural networks face many problems in the field of hyperspectral image classification, lack of effective utilization of spatial spectral information, gradient disappearance and overfitting as the model depth increases. In order to…

计算机视觉与模式识别 · 计算机科学 2023-07-14 Guandong Li

Talking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end,…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Fa-Ting Hong , Zunnan Xu , Zixiang Zhou , Jun Zhou , Xiu Li , Qin Lin , Qinglin Lu , Dan Xu

With the advance of deep learning technology, automatic video generation from audio or text has become an emerging and promising research topic. In this paper, we present a novel approach to synthesize video from the text. The method builds…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Sibo Zhang , Jiahong Yuan , Miao Liao , Liangjun Zhang

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the image domain remains…

计算机视觉与模式识别 · 计算机科学 2022-12-06 Xian Liu , Qianyi Wu , Hang Zhou , Yuanqi Du , Wayne Wu , Dahua Lin , Ziwei Liu

Generating realistic, dyadic talking head video requires ultra-low latency. Existing chunk-based methods require full non-causal context windows, introducing significant delays. This high latency critically prevents the immediate,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Bohong Chen , Haiyang Liu

Fully connected layer is an essential component of Convolutional Neural Networks (CNNs), which demonstrates its efficiency in computer vision tasks. The CNN process usually starts with convolution and pooling layers that first break down…

计算机视觉与模式识别 · 计算机科学 2020-09-24 M. Amine Mahmoudi , Aladine Chetouani , Fatma Boufera , Hedi Tabia

Cross-lingual voice conversion (CLVC) is a quite challenging task since the source and target speakers speak different languages. This paper proposes a CLVC framework based on bottleneck features and deep neural network (DNN). In the…

音频与语音处理 · 电气工程与系统科学 2019-11-12 M Kiran Reddy , K Sreenivasa Rao

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

计算机视觉与模式识别 · 计算机科学 2017-01-10 Ariel Ephrat , Shmuel Peleg

The objective of this paper is a neural network model that controls the pose and expression of a given face, using another face or modality (e.g. audio). This model can then be used for lightweight, sophisticated video and image editing. We…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Olivia Wiles , A. Sophia Koepke , Andrew Zisserman

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content.…

计算机视觉与模式识别 · 计算机科学 2023-08-28 Ziqiao Peng , Haoyu Wu , Zhenbo Song , Hao Xu , Xiangyu Zhu , Jun He , Hongyan Liu , Zhaoxin Fan

The most recent deep neural network (DNN) models exhibit impressive denoising performance in the time-frequency (T-F) magnitude domain. However, the phase is also a critical component of the speech signal that is easily overlooked. In this…

音频与语音处理 · 电气工程与系统科学 2021-06-10 Lu Zhang , Mingjiang Wang , Zehua Zhang , Xuyi Zhuang

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti