English
Related papers

Related papers: Text-based Editing of Talking-head Video

200 papers

We describe our novel deep learning approach for driving animated faces using both acoustic and visual information. In particular, speech-related facial movements are generated using audiovisual information, and non-speech facial movements…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-29 Ahmed Hussen Abdelaziz , Barry-John Theobald , Paul Dixon , Reinhard Knothe , Nicholas Apostoloff , Sachin Kajareker

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Panwen Hu , Nan Xiao , Feifei Li , Yongquan Chen , Rui Huang

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting…

Multimedia · Computer Science 2026-01-16 Diqiong Jiang , Kai Zhu , Dan Song , Jian Chang , Chenglizhao Chen , Zhenyu Wu

The generation of natural and high-quality speech from text is a challenging problem in the field of natural language processing. In addition to speech generation, speech editing is also a crucial task, which requires the seamless and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-11 Antonios Alexos , Pierre Baldi

Although automatically animating audio-driven talking heads has recently received growing interest, previous efforts have mainly concentrated on achieving lip synchronization with the audio, neglecting two crucial elements for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Shuai Tan , Bin Ji , Ye Pan

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Kim Sung-Bin , Lee Chae-Yeon , Gihun Son , Oh Hyun-Bin , Janghoon Ju , Suekyeong Nam , Tae-Hyun Oh

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Since the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to…

Artificial Intelligence · Computer Science 2022-10-14 Aolan Sun , Xulong Zhang , Tiandong Ling , Jianzong Wang , Ning Cheng , Jing Xiao

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Eyal Molad , Eliahu Horwitz , Dani Valevski , Alex Rav Acha , Yossi Matias , Yael Pritch , Yaniv Leviathan , Yedid Hoshen

We propose an end-to-end lecture video generation system that can generate realistic and complete lecture videos directly from annotated slides, instructor's reference voice and instructor's reference portrait video. Our system is primarily…

Multimedia · Computer Science 2022-09-20 Wenbin Wang , Yang Song , Sanjay Jha

Text-based video editing has recently attracted considerable interest in changing the style or replacing the objects with a similar structure. Beyond this, we demonstrate that properties such as shape, size, location, motion, etc., can also…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yue Ma , Xiaodong Cun , Sen Liang , Jinbo Xing , Yingqing He , Chenyang Qi , Siran Chen , Qifeng Chen

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip. In this paper, for the sake of both furthering this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-10 Lijie Fan , Wenbing Huang , Chuang Gan , Junzhou Huang , Boqing Gong

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

The text-driven image and video diffusion models have achieved unprecedented success in generating realistic and diverse content. Recently, the editing and variation of existing images and videos in diffusion-based generative models have…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Yuyang Zhao , Enze Xie , Lanqing Hong , Zhenguo Li , Gim Hee Lee

In this paper, we propose a novel approach to convert given speech audio to a photo-realistic speaking video of a specific person, where the output video has synchronized, realistic, and expressive rich body dynamics. We achieve this by…

Computer Vision and Pattern Recognition · Computer Science 2020-10-12 Miao Liao , Sibo Zhang , Peng Wang , Hao Zhu , Xinxin Zuo , Ruigang Yang

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often require fine-tuning for…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 He Feng , Donglin Di , Yongjia Ma , Wei Chen , Tonghua Su

How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural…

Computer Vision and Pattern Recognition · Computer Science 2019-05-24 Tae-Hyun Oh , Tali Dekel , Changil Kim , Inbar Mosseri , William T. Freeman , Michael Rubinstein , Wojciech Matusik

In this paper we present VideoSET, a method for Video Summary Evaluation through Text that can evaluate how well a video summary is able to retain the semantic information contained in its original video. We observe that semantics is most…

Computer Vision and Pattern Recognition · Computer Science 2014-06-24 Serena Yeung , Alireza Fathi , Li Fei-Fei

Realistic talking-head video generation is critical for virtual avatars, film production, and interactive systems. Current methods struggle with nuanced emotional expressions due to the lack of fine-grained emotion control. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Jiayi Lyu , Leigang Qu , Wenjing Zhang , Hanyu Jiang , Kai Liu , Zhenglin Zhou , Xiaobo Xia , Jian Xue , Tat-Seng Chua

We introduce a novel text-to-pose video editing method, ReimaginedAct. While existing video editing tasks are limited to changes in attributes, backgrounds, and styles, our method aims to predict open-ended human action changes in video.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Lan Wang , Vishnu Boddeti , Sernam Lim
‹ Prev 1 8 9 10 Next ›