English
Related papers

Related papers: EPiC: Efficient Video Camera Control Learning with…

200 papers

Visual localization is the task of estimating a 6-DoF camera pose of a query image within a provided 3D reference map. Thanks to recent advances in various 3D sensors, 3D point clouds are becoming a more accurate and affordable option for…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Minjung Kim , Junseo Koo , Gunhee Kim

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Chenjie Cao , Jingkai Zhou , Shikai Li , Jingyun Liang , Chaohui Yu , Fan Wang , Xiangyang Xue , Yanwei Fu

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Dejia Xu , Yifan Jiang , Chen Huang , Liangchen Song , Thorsten Gernoth , Liangliang Cao , Zhangyang Wang , Hao Tang

Filmmaking and animation production often require sophisticated techniques for coordinating camera transitions and object movements, typically involving labor-intensive real-world capturing. Despite advancements in generative AI for video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Yaowei Li , Xintao Wang , Zhaoyang Zhang , Zhouxia Wang , Ziyang Yuan , Liangbin Xie , Yuexian Zou , Ying Shan

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Shaoteng Liu , Yuechen Zhang , Wenbo Li , Zhe Lin , Jiaya Jia

Motion-controllable video generation is crucial for egocentric applications in virtual reality and embodied AI. However, existing methods often struggle to achieve 3D-consistent fine-grained hand articulation. By adopting on 2D trajectories…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Chenyangguang Zhang , Botao Ye , Boqi Chen , Alexandros Delitzas , Fangjinhua Wang , Marc Pollefeys , Xi Wang

Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. Compared to generic video synthesis, egocentric generation is…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Qiao Gu , Lingni Ma , Adam W Harley , Richard Newcombe , Florian Shkurti , Julian Straub

Video generation technologies are developing rapidly and have broad potential applications. Among these technologies, camera control is crucial for generating professional-quality videos that accurately meet user expectations. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Wanquan Feng , Jiawei Liu , Pengqi Tu , Tianhao Qi , Mingzhen Sun , Tianxiang Ma , Songtao Zhao , Siyu Zhou , Qian He

Following the advancements in text-guided image generation technology exemplified by Stable Diffusion, video generation is gaining increased attention in the academic community. However, relying solely on text guidance for video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Haoyu Zhao , Yuanfan Guo , Jianhua Han , Hang Xu , Xiaodan Liang

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera pose annotations, which are both data-intensive and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zhenghong Zhou , Jie An , Jiebo Luo

Implicit Neural Point Cloud (INPC) is a recent hybrid representation that combines the expressiveness of neural fields with the efficiency of point-based rendering, achieving state-of-the-art image quality in novel view synthesis. However,…

Although natural language instructions offer an intuitive way to guide automated image editing, deep-learning models often struggle to achieve high-quality results, largely due to the difficulty of creating large, high-quality training…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Sherry X. Chen , Misha Sra , Pradeep Sen

Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt to learn these…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Guojun Lei , Chi Wang , Yikai Wang , Hong Li , Ying Song , Weiwei Xu

Recent text-to-image (T2I) generators can synthesize realistic images, but still struggle with compositional prompts involving multiple objects, counts, attributes, and relations. We introduce EPIC (Efficient Predicate-Guided Inference-Time…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Sunung Mun , Sunghyun Cho , Jungseul Ok

We propose FlowAnchor, a training-free framework for stable and efficient inversion-free, flow-based video editing. Inversion-free editing methods have recently shown impressive efficiency and structure preservation in images by directly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Ze Chen , Lan Chen , Yuanhang Li , Qi Mao

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Zeqi Xiao , Wenqi Ouyang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most existing approaches that adapt CLIP for action recognition have primarily focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Hyo Jin Jon , Longbin Jin , Eun Yi Kim

Incomplete point clouds captured by 3D sensors often result in the loss of both geometric and semantic information. Most existing point cloud completion methods are built on rotation-variant frameworks trained with data in canonical poses,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zhifan Ni , Eckehard Steinbach

Text-guided generative diffusion models unlock powerful image creation and editing tools. While these have been extended to video generation, current approaches that edit the content of existing footage while retaining structure require…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Patrick Esser , Johnathan Chiu , Parmida Atighehchian , Jonathan Granskog , Anastasis Germanidis

The pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models. However, pretraining such a strong model is difficult in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Xiaoshui Huang , Zhou Huang , Sheng Li , Wentao Qu , Tong He , Yuenan Hou , Yifan Zuo , Wanli Ouyang