中文
相关论文

相关论文: Speech2Video Synthesis with 3D Skeleton Regulariza…

200 篇论文

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic…

多媒体 · 计算机科学 2022-04-07 Liyang Chen , Zhiyong Wu , Jun Ling , Runnan Li , Xu Tan , Sheng Zhao

Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this research was to study the automatic generation of high-quality…

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks.…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Dongchan Min , Minyoung Song , Eunji Ko , Sung Ju Hwang

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial…

Video-to-speech synthesis (also known as lip-to-speech) refers to the translation of silent lip movements into the corresponding audio. This task has received an increasing amount of attention due to its self-supervised nature (i.e., can be…

This paper addresses the problem of generating whole-body motion from speech. Despite great successes, prior methods still struggle to produce reasonable and diverse whole-body motions from speech. This is due to their reliance on…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Jinsong Zhang , Minjie Zhu , Yuxiang Zhang , Yebin Liu , Kun Li

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Sridhar S , Nithin A , Shakeel Rifath , Vasantha Raj

Synthesizing realistic co-speech gestures is an important and yet unsolved problem for creating believable motions that can drive a humanoid robot to interact and communicate with human users. Such capability will improve the impressions of…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Shuhong Lu , Youngwoo Yoon , Andrew Feng

In this paper, we propose a novel lip-to-speech generative adversarial network, Visual Context Attentional GAN (VCA-GAN), which can jointly model local and global lip movements during speech synthesis. Specifically, the proposed VCA-GAN…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Minsu Kim , Joanna Hong , Yong Man Ro

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

计算机视觉与模式识别 · 计算机科学 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to…

人机交互 · 计算机科学 2021-08-27 Siyang Wang , Simon Alexanderson , Joakim Gustafson , Jonas Beskow , Gustav Eje Henter , Éva Székely

This paper presents a new method for 3D action recognition with skeleton sequences (i.e., 3D trajectories of human skeleton joints). The proposed method first transforms each skeleton sequence into three clips each consisting of several…

计算机视觉与模式识别 · 计算机科学 2017-11-21 Qiuhong Ke , Mohammed Bennamoun , Senjian An , Ferdous Sohel , Farid Boussaid

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we…

计算机视觉与模式识别 · 计算机科学 2019-05-09 Daniel Cudeiro , Timo Bolkart , Cassidy Laidlaw , Anurag Ranjan , Michael J. Black

Different people have different facial expressions while speaking emotionally. A realistic facial animation system should consider such identity-specific speaking styles and facial idiosyncrasies to achieve high-degree of naturalness and…

人工智能 · 计算机科学 2023-10-27 Elif Bozkurt

Speechreading is the task of inferring phonetic information from visually observed articulatory facial movements, and is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a…

计算机视觉与模式识别 · 计算机科学 2017-08-31 Ariel Ephrat , Tavi Halperin , Shmuel Peleg

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

计算机视觉与模式识别 · 计算机科学 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

This paper presents a novel system that enables intelligent robots to exhibit realistic body gestures while communicating with humans. The proposed system consists of a listening model and a speaking model used in corresponding…

计算机视觉与模式识别 · 计算机科学 2019-11-18 Minjie Hua , Fuyuan Shi , Yibing Nan , Kai Wang , Hao Chen , Shiguo Lian

We present a method to edit a target portrait footage by taking a sequence of audio as input to synthesize a photo-realistic video. This method is unique because it is highly dynamic. It does not assume a person-specific rendering network…

计算机视觉与模式识别 · 计算机科学 2020-01-16 Linsen Song , Wayne Wu , Chen Qian , Ran He , Chen Change Loy