中文
相关论文

相关论文: CamLit: Unified Video Diffusion with Explicit Came…

200 篇论文

We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful harmonization). Unlike images, acquiring labeled data for…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Jun Myeong Choi , Jae Shin Yoon , Luchao Qi , Roni Sengupta , Joon-Young Lee

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Jiabin Luo , Junhui Lin , Zeyu Zhang , Biao Wu , Meng Fang , Ling Chen , Hao Tang

Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tackcle the mentioned…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Junpeng Jiang , Gangyi Hong , Lijun Zhou , Enhui Ma , Hengtong Hu , Xia Zhou , Jie Xiang , Fan Liu , Kaicheng Yu , Haiyang Sun , Kun Zhan , Peng Jia , Miao Zhang

Novel view synthesis from a single image requires inferring occluded regions of objects and scenes whilst simultaneously maintaining semantic and physical consistency with the input. Existing approaches condition neural radiance fields…

计算机视觉与模式识别 · 计算机科学 2023-02-21 Jiatao Gu , Alex Trevithick , Kai-En Lin , Josh Susskind , Christian Theobalt , Lingjie Liu , Ravi Ramamoorthi

Recent advances in diffusion-based generative models have established a new paradigm for image and video relighting. However, extending these capabilities to 4D relighting remains challenging, due primarily to the scarcity of paired 4D…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Zhenghuang Wu , Kang Chen , Zeyu Zhang , Hao Tang

In recent years, diffusion models have emerged as the most powerful approach in image synthesis. However, applying these models directly to video synthesis presents challenges, as it often leads to noticeable flickering contents. Although…

计算机视觉与模式识别 · 计算机科学 2023-08-11 Zhongjie Duan , Lizhou You , Chengyu Wang , Cen Chen , Ziheng Wu , Weining Qian , Jun Huang

We present FloVD, a novel video diffusion model for camera-controllable video generation. FloVD leverages optical flow to represent the motions of the camera and moving objects. This approach offers two key benefits. Since optical flow can…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Wonjoon Jin , Qi Dai , Chong Luo , Seung-Hwan Baek , Sunghyun Cho

Single-image relighting is a challenging task that involves reasoning about the complex interplay between geometry, materials, and lighting. Many prior methods either support only specific categories of images, such as portraits, or require…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Haian Jin , Yuan Li , Fujun Luan , Yuanbo Xiangli , Sai Bi , Kai Zhang , Zexiang Xu , Jin Sun , Noah Snavely

Recent single-image relighting methods, powered by advanced generative models, have achieved impressive photorealism on synthetic benchmarks. However, their effectiveness in the complex visual landscape of the real world remains largely…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Lezhong Wang , Mehmet Onurcan Kaya , Siavash Bigdeli , Jeppe Revall Frisvad

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urgently required.…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Xiaofan Li , Yifu Zhang , Xiaoqing Ye

Humans have the remarkable ability to construct consistent mental models of an environment, even under limited or varying levels of illumination. We wish to endow robots with this same capability. In this paper, we tackle the challenge of…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Tianyi Zhang , Kaining Huang , Weiming Zhi , Matthew Johnson-Roberson

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

We introduce a scalable framework for novel view synthesis from RGB-D images with largely incomplete scene coverage. While generative neural approaches have demonstrated spectacular results on 2D images, they have not yet achieved similar…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Zuoyue Li , Tianxing Fan , Zhenqiang Li , Zhaopeng Cui , Yoichi Sato , Marc Pollefeys , Martin R. Oswald

Event cameras offer significant advantages for low-light video enhancement, primarily due to their high dynamic range. Current research, however, is severely limited by the absence of large-scale, real-world, and spatio-temporally aligned…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Kanghao Chen , Guoqiang Liang , Hangyu Li , Yunfan Lu , Lin Wang

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Dejia Xu , Yifan Jiang , Chen Huang , Liangchen Song , Thorsten Gernoth , Liangliang Cao , Zhangyang Wang , Hao Tang

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Yuxuan Xue , Ruofan Liang , Egor Zakharov , Timur Bagautdinov , Chen Cao , Giljoo Nam , Shunsuke Saito , Gerard Pons-Moll , Javier Romero

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

We introduce a 3D-aware diffusion model, ZeroNVS, for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds, we propose new techniques to address challenges…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Kyle Sargent , Zizhang Li , Tanmay Shah , Charles Herrmann , Hong-Xing Yu , Yunzhi Zhang , Eric Ryan Chan , Dmitry Lagun , Li Fei-Fei , Deqing Sun , Jiajun Wu

Given a set of images of a scene, the re-rendering of this scene from novel views and lighting conditions is an important and challenging problem in Computer Vision and Graphics. On the one hand, most existing works in Computer Vision…

计算机视觉与模式识别 · 计算机科学 2022-07-28 Linjie Lyu , Ayush Tewari , Thomas Leimkuehler , Marc Habermann , Christian Theobalt

Novel View Synthesis (NVS) for street scenes play a critical role in the autonomous driving simulation. The current mainstream technique to achieve it is neural rendering, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Zhongrui Yu , Haoran Wang , Jinze Yang , Hanzhang Wang , Zeke Xie , Yunfeng Cai , Jiale Cao , Zhong Ji , Mingming Sun