中文
相关论文

相关论文: ActCam: Zero-Shot Joint Camera and 3D Motion Contr…

200 篇论文

The proliferation of video content demands efficient and flexible neural network based approaches for generating new video content. In this paper, we propose a novel approach that combines zero-shot text-to-video generation with ControlNet…

计算机视觉与模式识别 · 计算机科学 2023-05-11 Rohan Dhesikan , Vignesh Rajmohan

The diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Chenyang Qi , Xiaodong Cun , Yong Zhang , Chenyang Lei , Xintao Wang , Ying Shan , Qifeng Chen

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

We propose FreeSim, a camera simulation method for autonomous driving. FreeSim emphasizes high-quality rendering from viewpoints beyond the recorded ego trajectories. In such viewpoints, previous methods have unacceptable degradation…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Lue Fan , Hao Zhang , Qitai Wang , Hongsheng Li , Zhaoxiang Zhang

We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Xuanchi Ren , Tianchang Shen , Jiahui Huang , Huan Ling , Yifan Lu , Merlin Nimier-David , Thomas Müller , Alexander Keller , Sanja Fidler , Jun Gao

Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Edurne Bernal-Berdun , Ana Serrano , Belen Masia , Matheus Gadelha , Yannick Hold-Geoffroy , Xin Sun , Diego Gutierrez

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes (e.g., appearance,…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Jiwen Yu , Xiaodong Cun , Chenyang Qi , Yong Zhang , Xintao Wang , Ying Shan , Jian Zhang

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Hyelin Nam , Hyojun Go , Byeongjun Park , Byung-Hoon Kim , Hyungjin Chung

Treating human motion and camera trajectory generation separately overlooks a core principle of cinematography: the tight interplay between actor performance and camera work in the screen space. In this paper, we are the first to cast this…

图形学 · 计算机科学 2026-04-02 Robin Courant , Xi Wang , David Loiseaux , Marc Christie , Vicky Kalogeiton

This paper presents a method that allows users to design cinematic video shots in the context of image-to-video generation. Shot design, a critical aspect of filmmaking, involves meticulously planning both camera movements and object…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Jinbo Xing , Long Mai , Cusuh Ham , Jiahui Huang , Aniruddha Mahapatra , Chi-Wing Fu , Tien-Tsin Wong , Feng Liu

Motions in a video primarily consist of camera motion, induced by camera movement, and object motion, resulting from object movement. Accurate control of both camera and object motion is essential for video generation. However, existing…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Zhouxia Wang , Ziyang Yuan , Xintao Wang , Tianshui Chen , Menghan Xia , Ping Luo , Ying Shan

We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fields, providing…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xinyao Liao , Xianfang Zeng , Liao Wang , Gang Yu , Guosheng Lin , Chi Zhang

We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct the source character itself but to reinterpret its motion as a plausible and editable…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Liuhan Chen , Lei Zhong , Jiewei Wang , Qin Shuai , Li Yuan , Leidong Fan , Qing Li , Kanglin Liu

Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent editing of subjects and scenes remains a critical challenge. Recent attempts to incorporate richer…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Fengyuan Yang , Luying Huang , Jiazhi Guan , Quanwei Yang , Dongwei Pan , Jianglin Fu , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou , Angela Yao

Long-term video generation and prediction remain challenging tasks in computer vision, particularly in partially observable scenarios where cameras are mounted on moving platforms. The interaction between observed image frames and the…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Meenakshi Sarkar , Debasish Ghose

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Qin Guo , Tianyu Yang , Xuanhua He , Fei Shen , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Dan Xu

Human activity videos involve rich, varied interactions between people and objects. In this paper we develop methods for generating such videos -- making progress toward addressing the important, open problem of video generation in complex…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Megha Nawhal , Mengyao Zhai , Andreas Lehrmann , Leonid Sigal , Greg Mori

We propose a zero-shot approach for generating consistent videos of animated characters based on Text-to-Image (T2I) diffusion models. Existing Text-to-Video (T2V) methods are expensive to train and require large-scale video datasets to…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Abdelrahman Eldesokey , Peter Wonka

Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video-to-4D methods typically rely on manually annotated camera poses, which are labor-intensive and brittle…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Dongyue Lu , Ao Liang , Tianxin Huang , Xiao Fu , Yuyang Zhao , Baorui Ma , Liang Pan , Wei Yin , Lingdong Kong , Wei Tsang Ooi , Ziwei Liu

Recent advancements in generative models have provided promising solutions for synthesizing realistic driving videos, which are crucial for training autonomous driving perception models. However, existing approaches often struggle with…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Wei Wu , Xi Guo , Weixuan Tang , Tingxuan Huang , Chiyu Wang , Dongyue Chen , Chenjing Ding