English
Related papers

Related papers: CamCo: Camera-Controllable 3D-Consistent Image-to-…

200 papers

This paper presents DualCamCtrl, a novel end-to-end diffusion model for camera-controlled video generation. Recent works have advanced this field by representing camera poses as ray-based conditions, yet they often lack sufficient scene…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Hongfei Zhang , Kanghao Chen , Zixin Zhang , Harold Haodong Chen , Yuanhuiyi Lyu , Yuqi Zhang , Shuai Yang , Kun Zhou , Yingcong Chen

Image generation and editing have seen a great deal of advancements with the rise of large-scale diffusion models that allow user control of different modalities such as text, mask, depth maps, etc. However, controlled editing of videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 AmirHossein Zamani , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Qiucheng Wu , Handong Zhao , Zhixin Shu , Jing Shi , Yang Zhang , Shiyu Chang

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

Recent progress in generative diffusion models has greatly advanced text-to-video generation. While text-to-video models trained on large-scale, diverse datasets can produce varied outputs, these generations often deviate from user…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Runtao Liu , Haoyu Wu , Zheng Ziqiang , Chen Wei , Yingqing He , Renjie Pi , Qifeng Chen

Image animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xin Ma , Yaohui Wang , Genyun Jia , Xinyuan Chen , Tien-Tsin Wong , Cunjian Chen

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Weixi Feng , Chao Liu , Sifei Liu , William Yang Wang , Arash Vahdat , Weili Nie

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yisu Zhang , Chenjie Cao , Tengfei Wang , Xuhui Zuo , Junta Wu , Jianke Zhu , Chunchao Guo

Text-to-video generation has shown promising results. However, by taking only natural languages as input, users often face difficulties in providing detailed information to precisely control the model's output. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Hsin-Ping Huang , Yu-Chuan Su , Deqing Sun , Lu Jiang , Xuhui Jia , Yukun Zhu , Ming-Hsuan Yang

Camera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Chenjie Cao , Jingkai Zhou , Shikai Li , Jingyun Liang , Chaohui Yu , Fan Wang , Xiangyang Xue , Yanwei Fu

Diffusion models have demonstrated impressive performance in generating high-quality videos from text prompts or images. However, precise control over the video generation process, such as camera manipulation or content editing, remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Zekai Gu , Rui Yan , Jiahao Lu , Peng Li , Zhiyang Dou , Chenyang Si , Zhen Dong , Qifeng Liu , Cheng Lin , Ziwei Liu , Wenping Wang , Yuan Liu

Motion-controllable video generation is crucial for egocentric applications in virtual reality and embodied AI. However, existing methods often struggle to achieve 3D-consistent fine-grained hand articulation. By adopting on 2D trajectories…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Chenyangguang Zhang , Botao Ye , Boqi Chen , Alexandros Delitzas , Fangjinhua Wang , Marc Pollefeys , Xi Wang

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a specific camera setting such as creating different fields of view using a 24mm lens versus a 70mm lens, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yu Yuan , Xijun Wang , Yichen Sheng , Prateek Chennuri , Xingguang Zhang , Stanley Chan

Recent years have seen significant advancements in 3D generation. While methods like score distillation achieve impressive results, they often require extensive per-scene optimization, which limits their time efficiency. On the other hand,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Chen Wang , Jiatao Gu , Xiaoxiao Long , Yuan Liu , Lingjie Liu

State-of-the-art diffusion models can generate highly realistic images based on various conditioning like text, segmentation, and depth. However, an essential aspect often overlooked is the specific camera geometry used during image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Andrey Voynov , Amir Hertz , Moab Arar , Shlomi Fruchter , Daniel Cohen-Or

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Qihang Zhang , Shuangfei Zhai , Miguel Angel Bautista , Kevin Miao , Alexander Toshev , Joshua Susskind , Jiatao Gu

Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Léopold Maillard , Tom Durand , Adrien Ramanana Rahary , Maks Ovsjanikov

Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and temporal control. We introduce a 4D-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yiming Wang , Qihang Zhang , Shengqu Cai , Tong Wu , Jan Ackermann , Zhengfei Kuang , Yang Zheng , Frano Rajič , Siyu Tang , Gordon Wetzstein
‹ Prev 1 3 4 5 6 7 10 Next ›