中文
相关论文

相关论文: LEO: Generative Latent Image Animator for Human Vi…

200 篇论文

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate videos at novel…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Yudong Jin , Sida Peng , Xuan Wang , Tao Xie , Zhen Xu , Yifan Yang , Yujun Shen , Hujun Bao , Xiaowei Zhou

While recent years have witnessed great progress on using diffusion models for video generation, most of them are simple extensions of image generation frameworks, which fail to explicitly consider one of the key differences between videos…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Jingyun Liang , Yuchen Fan , Kai Zhang , Radu Timofte , Luc Van Gool , Rakesh Ranjan

Long video generation has gained increasing attention due to its widespread applications in fields such as entertainment and simulation. Despite advances, synthesizing temporally coherent and visually compelling long sequences remains a…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Jiahao Chen , Hangjie Yuan , Yichen Qian , Jingyun Liang , Jiazheng Xing , Pengwei Liu , Weihua Chen , Fan Wang , Bing Su

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant challenges in handling…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Xiaojing Zhong , Xinyi Huang , Xiaofeng Yang , Guosheng Lin , Qingyao Wu

Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for contemporary diffusion-based models, which often suffer from identity drift and are limited to…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Jingxuan He , Busheng Su , Finn Wong

Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, synchronized exo-ego data inherently introduces substantial…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Mohammad Mahdi , Nedko Savov , Danda Pani Paudel , Luc Van Gool

Conditional image-to-video (cI2V) generation aims to synthesize a new plausible video starting from an image (e.g., a person's face) and a condition (e.g., an action class label like smile). The key challenge of the cI2V task lies in the…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Haomiao Ni , Changhao Shi , Kai Li , Sharon X. Huang , Martin Renqiang Min

Digital human motion synthesis is a vibrant research field with applications in movies, AR/VR, and video games. Whereas methods were proposed to generate natural and realistic human motions, most only focus on modeling humans and largely…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Quanzhou Li , Jingbo Wang , Chen Change Loy , Bo Dai

Human motion synthesis conditioned on textual input has gained significant attention in recent years due to its potential applications in various domains such as gaming, film production, and virtual reality. Conditioned Motion synthesis…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Avinash Amballa , Gayathri Akkinapalli , Vinitra Muralikrishnan

We describe a new spatio-temporal video autoencoder, based on a classic spatial image autoencoder and a novel nested temporal autoencoder. The temporal encoder is represented by a differentiable visual memory composed of convolutional long…

机器学习 · 计算机科学 2016-09-02 Viorica Patraucean , Ankur Handa , Roberto Cipolla

The success of deep learning models has led to their adaptation and adoption by prominent video understanding methods. The majority of these approaches encode features in a joint space-time modality for which the inner workings and learned…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Alexandros Stergiou , Nikos Deligiannis

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions.…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Xiaodi Li , Pan Xie , Yi Ren , Qijun Gan , Chen Zhang , Fangyuan Kong , Xiang Yin , Bingyue Peng , Zehuan Yuan

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Lingteng Qiu , Xiaodong Gu , Peihao Li , Qi Zuo , Weichao Shen , Junfei Zhang , Kejie Qiu , Weihao Yuan , Guanying Chen , Zilong Dong , Liefeng Bo

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Xiangyang Luo , Xiaozhe Xin , Tao Feng , Xu Guo , Meiguang Jin , Junfeng Ma

Recovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and…

计算机视觉与模式识别 · 计算机科学 2021-08-25 Siwei Zhang , Yan Zhang , Federica Bogo , Marc Pollefeys , Siyu Tang

AI-generated content has attracted lots of attention recently, but photo-realistic video synthesis is still challenging. Although many attempts using GANs and autoregressive models have been made in this area, the visual quality and length…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Yingqing He , Tianyu Yang , Yong Zhang , Ying Shan , Qifeng Chen

Computational imaging methods increasingly rely on powerful generative diffusion models to tackle challenging image restoration tasks. In particular, state-of-the-art zero-shot image inverse solvers leverage distilled text-to-image latent…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Alessio Spagnoletti , Andrés Almansa , Marcelo Pereyra

Creating plausible virtual actors from images of real actors remains one of the key challenges in computer vision and computer graphics. Marker-less human motion estimation and shape modeling from images in the wild bring this challenge to…

计算机视觉与模式识别 · 计算机科学 2020-01-22 Thiago L. Gomes , Renato Martins , João Ferreira , Erickson R. Nascimento

Videos express highly structured spatio-temporal patterns of visual data. A video can be thought of as being governed by two factors: (i) temporally invariant (e.g., person identity), or slowly varying (e.g., activity), attribute-induced…

计算机视觉与模式识别 · 计算机科学 2018-03-26 Jiawei He , Andreas Lehrmann , Joseph Marino , Greg Mori , Leonid Sigal

Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Runze Zhang , Guoguang Du , Xiaochuan Li , Qi Jia , Liang Jin , Lu Liu , Jingjing Wang , Cong Xu , Zhenhua Guo , Yaqian Zhao , Xiaoli Gong , Rengang Li , Baoyu Fan