中文
相关论文

相关论文: Zero-shot High-fidelity and Pose-controllable Char…

200 篇论文

Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While deep learning-based methods have recently shown superior accuracy compared to traditional…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Dominik Kuczkowski , Laura Ruotsalainen

Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Mallikarjun B. R. , Fei Yin , Vikram Voleti , Nikita Drobyshev , Maksim Lapin , Aaryaman Vasishta , Varun Jampani

Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. In content creation workflows, precise and simultaneous control over camera motion, object motion, and…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Sixiao Zheng , Zimian Peng , Yanpeng Zhou , Yi Zhu , Hang Xu , Xiangru Huang , Yanwei Fu

Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the state-of-the-art of vid2vid has advanced significantly,…

计算机视觉与模式识别 · 计算机科学 2019-10-29 Ting-Chun Wang , Ming-Yu Liu , Andrew Tao , Guilin Liu , Jan Kautz , Bryan Catanzaro

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Mingyang Wu , Ashirbad Mishra , Soumik Dey , Shuo Xing , Naveen Ravipati , Hansi Wu , Binbin Li , Zhengzhong Tu

Pose diversity is an inherent representative characteristic of 2D images. Due to the 3D to 2D projection mechanism, there is evident content discrepancy among distinct pose images. This is the main obstacle bothering pose transformation…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Yuelong Li , Tengfei Xiao , Lei Geng , Jianming Wang

Advancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Zichen Tang , Hongyu Yang , Hanchen Zhang , Jiaxin Chen , Di Huang

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

Recent large-scale text-to-image generation models have made significant improvements in the quality, realism, and diversity of the synthesized images and enable users to control the created content through language. However, the…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Samaneh Azadi , Thomas Hayes , Akbar Shah , Guan Pang , Devi Parikh , Sonal Gupta

Diffusion models have made significant advances in generating high-quality images, but their application to video generation has remained challenging due to the complexity of temporal motion. Zero-shot video editing offers a solution by…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Xirui Li , Chao Ma , Xiaokang Yang , Ming-Hsuan Yang

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for applications that demand…

计算机视觉与模式识别 · 计算机科学 2025-01-27 Yujian Liu , Shidang Xu , Jing Guo , Dingbin Wang , Zairan Wang , Xianfeng Tan , Xiaoli Liu

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Min-Jung Kim , Jeongho Kim , Hoiyeong Jin , Junha Hyung , Jaegul Choo

Current approaches to pose generation rely heavily on intermediate representations, either through two-stage pipelines with quantization or autoregressive models that accumulate errors during inference. This fundamental limitation leads to…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Yayuan Li , Filippos Bellos , Jason Corso

Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely designed for realistic videos-struggle to evaluate animation-style generation with its…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Leyi Wu , Pengjun Fang , Kai Sun , Yazhou Xing , Yinwei Wu , Songsong Wang , Ziqi Huang , Dan Zhou , Yingqing He , Ying-Cong Chen , Qifeng Chen

Identity-preserving text-to-video generation (IPT2V) empowers users to produce diverse and imaginative videos with consistent human facial identity. Despite recent progress, existing methods often suffer from significant identity distortion…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Yuanzhi Wang , Xuhua Ren , Jiaxiang Cheng , Bing Ma , Kai Yu , Sen Liang , Wenyue Li , Tianxiang Zheng , Qinglin Lu , Zhen Cui

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Youxin Pang , Yong Zhang , Weize Quan , Yanbo Fan , Xiaodong Cun , Ying Shan , Dong-ming Yan

Traditional methods for constructing high-quality, personalized head avatars from monocular videos demand extensive face captures and training time, posing a significant challenge for scalability. This paper introduces a novel approach to…

计算机视觉与模式识别 · 计算机科学 2024-02-20 Zhixuan Yu , Ziqian Bai , Abhimitra Meka , Feitong Tan , Qiangeng Xu , Rohit Pandey , Sean Fanello , Hyun Soo Park , Yinda Zhang

In this paper, we introduce PoseCrafter, a one-shot method for personalized video generation following the control of flexible poses. Built upon Stable Diffusion and ControlNet, we carefully design an inference process to produce…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yong Zhong , Min Zhao , Zebin You , Xiaofeng Yu , Changwang Zhang , Chongxuan Li

Automated pose correction remains a significant challenge in AI-driven fitness systems, despite extensive research in activity recognition. This work presents PosePilot, a novel system that integrates pose recognition with real-time…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Rushiraj Gadhvi , Priyansh Desai , Siddharth