English
Related papers

Related papers: Audio-driven Neural Gesture Reenactment with Video…

200 papers

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emotion is one of the key…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Xingqun Qi , Chen Liu , Lincheng Li , Jie Hou , Haoran Xin , Xin Yu

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Ji-Hoon Kim , Jeongsoo Choi , Jaehun Kim , Chaeyoung Jung , Joon Son Chung

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

In this paper, we propose a learning-based method to compose a video-story from a group of video clips that describe an activity or experience. We learn the coherence between video clips from real videos via the Recurrent Neural Network…

Computer Vision and Pattern Recognition · Computer Science 2018-02-01 Guangyu Zhong , Yi-Hsuan Tsai , Sifei Liu , Zhixun Su , Ming-Hsuan Yang

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-11 Shivam Mehta , Ruibo Tu , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

We propose technology to enable a new medium of expression, where video elements can be looped, merged, and triggered, interactively. Like audio, video is easy to sample from the real world but hard to segment into clean reusable elements.…

Human-Computer Interaction · Computer Science 2017-05-23 Corneliu Ilisescu , Halil Aytac Kanaci , Matteo Romagnoli , Neill D. F. Campbell , Gabriel J. Brostow

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a…

Human-Computer Interaction · Computer Science 2023-05-19 Sicheng Yang , Zhiyong Wu , Minglei Li , Zhensong Zhang , Lei Hao , Weihong Bao , Haolin Zhuang

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with their matching…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Thomas Hummel , Otniel-Bogdan Mercea , A. Sophia Koepke , Zeynep Akata

We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME), a new mesh-level…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Haiyang Liu , Zihao Zhu , Giorgio Becherini , Yichen Peng , Mingyang Su , You Zhou , Xuefei Zhe , Naoya Iwamoto , Bo Zheng , Michael J. Black

Generating 3D human gestures and speech from a text script is critical for creating realistic talking avatars. One solution is to leverage separate pipelines for text-to-speech (TTS) and speech-to-gesture (STG), but this approach suffers…

Multimedia · Computer Science 2024-09-26 Zixin Guo , Jian Zhang

Upsampling videos of human activity is an interesting yet challenging task with many potential applications ranging from gaming to entertainment and sports broadcasting. The main difficulty in synthesizing video frames in this setting stems…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Hsuan-I Ho , Xu Chen , Jie Song , Otmar Hilliges

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

Automatic gesture generation from speech generally relies on implicit modelling of the nondeterministic speech-gesture relationship and can result in averaged motion lacking defined form. Here, we propose a database-driven approach of…

Human-Computer Interaction · Computer Science 2021-03-05 Ylva Ferstl , Michael Neff , Rachel McDonnell

For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it…

Graphics · Computer Science 2020-09-07 Youngwoo Yoon , Bok Cha , Joo-Haeng Lee , Minsu Jang , Jaeyeon Lee , Jaehong Kim , Geehyuk Lee

In this paper, we propose a novel machine learning architecture for facial reenactment. In particular, contrary to the model-based approaches or recent frame-based methods that use Deep Convolutional Neural Networks (DCNNs) to generate…

Computer Vision and Pattern Recognition · Computer Science 2020-05-25 Mohammad Rami Koujan , Michail Christos Doukas , Anastasios Roussos , Stefanos Zafeiriou

The goal of video summarization is to select keyframes that are visually diverse and can represent a whole story of an input video. State-of-the-art approaches for video summarization have mostly regarded the task as a frame-wise keyframe…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Jungin Park , Jiyoung Lee , Ig-Jae Kim , Kwanghoon Sohn

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hongye Cheng , Tianyu Wang , Guangsi Shi , Zexing Zhao , Yanwei Fu