English
Related papers

Related papers: Perceptual Conversational Head Generation with Reg…

200 papers

High-fidelity digital humans are increasingly used in interactive applications, yet achieving both visual realism and real-time responsiveness remains a major challenge. We present a high-fidelity, real-time conversational digital human…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hongbin Huang , Junwei Li , Tianxin Xie , Zhuang Li , Cekai Weng , Yaodong Yang , Yue Luo , Li Liu , Jing Tang , Zhijing Shao , Zeyu Wang

Generating talking person portraits with arbitrary speech audio is a crucial problem in the field of digital human and metaverse. A modern talking face generation method is expected to achieve the goals of generalized audio-lip…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Zhenhui Ye , Jinzheng He , Ziyue Jiang , Rongjie Huang , Jiawei Huang , Jinglin Liu , Yi Ren , Xiang Yin , Zejun Ma , Zhou Zhao

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness and realism.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yuzhe Weng , Haotian Wang , Yuanhong Yu , Jun Du , Shan He , Xiaoyan Wu , Haoran Xu

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper introduces JoyGen, a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Qili Wang , Dajiang Wu , Zihang Xu , Junshi Huang , Jun Lv

Given one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The…

Multimedia · Computer Science 2023-02-28 Jianrong Wang , Yaxin Zhao , Li Liu , Hongkai Fan , Tianyi Xu , Qi Li , Sen Li

Recent "Thinking with Video" approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reasoning artifacts. Even strong VGMs, however, exhibit two recurring failure modes on…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Joowon Kim , Seungho Shin , Joonhyung Park , Eunho Yang

Multimodal headline utilizes both video frames and transcripts to generate the natural language title of the videos. Due to a lack of large-scale, manually annotated data, the task of annotating grounded headlines for video is labor…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Lingfeng Qiao , Chen Wu , Ye Liu , Haoyuan Peng , Di Yin , Bo Ren

Speech-driven 3D facial animation is challenging due to the complex geometry of human faces and the limited availability of 3D audio-visual data. Prior works typically focus on learning phoneme-level features of short audio windows with…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

Self-supervised pre-training, such as BERT, MASS and BART, has emerged as a powerful technique for natural language understanding and generation. Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train…

Computation and Language · Computer Science 2020-09-22 Bin Bi , Chenliang Li , Chen Wu , Ming Yan , Wei Wang , Songfang Huang , Fei Huang , Luo Si

Recent video and language pretraining frameworks lack the ability to generate sentences. We present Multimodal Video Generative Pretraining (MV-GPT), a new pretraining framework for learning from unlabelled videos which can be effectively…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Paul Hongsuck Seo , Arsha Nagrani , Anurag Arnab , Cordelia Schmid

Longitudinal Dialogues (LD) are the most challenging type of conversation for human-machine dialogue systems. LDs include the recollections of events, personal thoughts, and emotions specific to each individual in a sparse sequence of…

Computation and Language · Computer Science 2023-11-01 Seyed Mahed Mousavi , Simone Caldarella , Giuseppe Riccardi

This paper summarises the design of the candidate ED for the Challenge on Learned Image Compression 2024. This candidate aims at providing an anchor based on conventional coding technologies to the learning-based approaches mostly targeted…

Image and Video Processing · Electrical Eng. & Systems 2024-01-05 Pierrick Philippe , Théo Ladune , Stéphane Davenet , Thomas Leguay

Traditionally, video conferencing is a widely adopted solution for telecommunication, but a lack of immersiveness comes inherently due to the 2D nature of facial representation. The integration of Virtual Reality (VR) in a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-03 Surabhi Gupta , Ashwath Shetty , Avinash Sharma

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Xudong Lin , Gedas Bertasius , Jue Wang , Shih-Fu Chang , Devi Parikh , Lorenzo Torresani

Real-time interactive video-chat portraits have been increasingly recognized as the future trend, particularly due to the remarkable progress made in text and voice chat technologies. However, existing methods primarily focus on real-time…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jinwei Qi , Chaonan Ji , Sheng Xu , Peng Zhang , Bang Zhang , Liefeng Bo

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

Multiple studies in the past have shown that there is a strong correlation between human vocal characteristics and facial features. However, existing approaches generate faces simply from voice, without exploring the set of features that…

Computer Vision and Pattern Recognition · Computer Science 2021-07-19 Hao Liang , Lulan Yu , Guikang Xu , Bhiksha Raj , Rita Singh

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Tanzila Rahman , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Shweta Mahajan , Leonid Sigal

Over the past few years there is a growing interest in the learning-based self driving system. To ensure safety, such systems are first developed and validated in simulators before being deployed in the real world. However, most of the…

Robotics · Computer Science 2021-03-15 Quanyi Li , Zhenghao Peng , Qihang Zhang , Chunxiao Liu , Bolei Zhou

Current generative video models excel at producing novel content from text and image prompts, but leave a critical gap in editing existing pre-recorded videos, where minor alterations to the spoken script require preserving motion, temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 John Flynn , Wolfgang Paier , Dimitar Dinev , Sam Nhut Nguyen , Hayk Poghosyan , Manuel Toribio , Sandipan Banerjee , Guy Gafni
‹ Prev 1 8 9 10 Next ›