English
Related papers

Related papers: InfinityHuman: Towards Long-Term Audio-Driven Huma…

200 papers

Recent advancements in human video generation and animation tasks, driven by diffusion models, have achieved significant progress. However, expressive and realistic human animation remains challenging due to the trade-off between motion…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Chao Liang , Jianwen Jiang , Wang Liao , Jiaqi Yang , Zerong zheng , Weihong Zeng , Han Liang

We present InfiniDreamer, a novel framework for arbitrarily long human motion generation. InfiniDreamer addresses the limitations of current motion generation methods, which are typically restricted to short sequences due to the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Wenjie Zhuo , Fan Ma , Hehe Fan

In this paper, we present a video-based learning framework for animating personalized 3D talking faces from audio. We introduce two training-time data normalizations that significantly improve data sample efficiency. First, we isolate and…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Avisek Lahiri , Vivek Kwatra , Christian Frueh , John Lewis , Chris Bregler

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Alireza Javanmardi , Pragati Jaiswal , Tewodros Amberbir Habtegebrial , Christen Millerdurai , Shaoxiang Wang , Alain Pagani , Didier Stricker

Lip synchronization is the task of aligning a speaker's lip movements in video with corresponding speech audio, and it is essential for creating realistic, expressive video content. However, existing methods often rely on reference frames…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Ziqiao Peng , Jiwen Liu , Haoxian Zhang , Xiaoqiang Liu , Songlin Tang , Pengfei Wan , Di Zhang , Hongyan Liu , Jun He

Current 3D human animation methods struggle to achieve photorealism: kinematics-based approaches lack non-rigid dynamics (e.g., clothing dynamics), while methods that leverage video diffusion priors can synthesize non-rigid motion but…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Qi Sun , Can Wang , Jiaxiang Shang , Yingchun Liu , Jing Liao

Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Jiahao Cui , Hui Li , Yao Yao , Hao Zhu , Hanlin Shang , Kaihui Cheng , Hang Zhou , Siyu Zhu , Jingdong Wang

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input…

Computer Vision and Pattern Recognition · Computer Science 2023-01-03 Balamurugan Thambiraja , Ikhsanul Habibie , Sadegh Aliakbarian , Darren Cosker , Christian Theobalt , Justus Thies

Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for contemporary diffusion-based models, which often suffer from identity drift and are limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Jingxuan He , Busheng Su , Finn Wong

Video dubbing aims to synthesize realistic, lip-synced videos from a reference video and a driving audio signal. Although existing methods can accurately generate mouth shapes driven by audio, they often fail to preserve identity-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Runzhen Liu , Qinjie Lin , Yunfei Liu , Lijian Lin , Ye Zhu , Yu Li , Chuhua Xian , Fa-Ting Hong

Predicting future human behavior from an input human video is a useful task for applications such as autonomous driving and robotics. While most previous works predict a single future, multiple futures with different behavior can…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Naoya Fushishita , Antonio Tejero-de-Pablos , Yusuke Mukuta , Tatsuya Harada

Spatio-temporal coherency is a major challenge in synthesizing high quality videos, particularly in synthesizing human videos that contain rich global and local deformations. To resolve this challenge, previous approaches have resorted to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Yaohui Wang , Xin Ma , Xinyuan Chen , Cunjian Chen , Antitza Dantcheva , Bo Dai , Yu Qiao

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the motion of a source person. The problem is inherently…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Sifan Wu , Zhenguang Liu , Beibei Zhang , Roger Zimmermann , Zhongjie Ba , Xiaosong Zhang , Kui Ren

Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermediate temporal anchors…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Junyoung Seo , Rodrigo Mira , Alexandros Haliassos , Stella Bounareli , Honglie Chen , Linh Tran , Seungryong Kim , Zoe Landgraf , Jie Shen

In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermediate representations…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Huawei Wei , Zejun Yang , Zhisheng Wang

Human image animation involves generating a video from a static image by following a specified pose sequence. Current approaches typically adopt a multi-stage pipeline that separately learns appearance and motion, which often leads to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Qilin Wang , Zhengkai Jiang , Chengming Xu , Jiangning Zhang , Yabiao Wang , Xinyi Zhang , Yun Cao , Weijian Cao , Chengjie Wang , Yanwei Fu

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse character types-realistic humans, full-body figures, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Hongwei Yi , Tian Ye , Shitong Shao , Xuancheng Yang , Jiantong Zhao , Hanzhong Guo , Terrance Wang , Qingyu Yin , Zeke Xie , Lei Zhu , Wei Li , Michael Lingelbach , Daquan Zhou

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuming Jiang , Shuai Yang , Tong Liang Koh , Wayne Wu , Chen Change Loy , Ziwei Liu

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zijun Wang , Panwen Hu , Jing Wang , Terry Jingchen Zhang , Yuhao Cheng , Long Chen , Yiqiang Yan , Zutao Jiang , Hanhui Li , Xiaodan Liang