English
Related papers

Related papers: DrivingScene: A Multi-Task Online Feed-Forward 3D …

200 papers

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

While existing feed-forward Gaussian splatting models offer computational efficiency and can generalize to sparse view settings, their performance is fundamentally constrained by relying on a single forward pass for inference. We propose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Haofei Xu , Daniel Barath , Andreas Geiger , Marc Pollefeys

We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Zhaoxi Chen , Tianqi Liu , Long Zhuo , Jiawei Ren , Zeng Tao , He Zhu , Fangzhou Hong , Liang Pan , Ziwei Liu

The increasing demand for virtual reality applications has highlighted the significance of crafting immersive 3D assets. We present a text-to-3D 360$^{\circ}$ scene generation pipeline that facilitates the creation of comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Shijie Zhou , Zhiwen Fan , Dejia Xu , Haoran Chang , Pradyumna Chari , Tejas Bharadwaj , Suya You , Zhangyang Wang , Achuta Kadambi

We propose GaussCtrl, a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS and edits them by using a pre-trained 2D diffusion model…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jing Wu , Jia-Wang Bian , Xinghui Li , Guangrun Wang , Ian Reid , Philip Torr , Victor Adrian Prisacariu

We introduce AnySplat, a feed forward network for novel view synthesis from uncalibrated image collections. In contrast to traditional neural rendering pipelines that demand known camera poses and per scene optimization, or recent feed…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Lihan Jiang , Yucheng Mao , Linning Xu , Tao Lu , Kerui Ren , Yichen Jin , Xudong Xu , Mulin Yu , Jiangmiao Pang , Feng Zhao , Dahua Lin , Bo Dai

This paper presents an effective solution for view extrapolation in autonomous driving scenarios. Recent approaches focus on generating shifted novel view images from given viewpoints using diffusion models. However, these methods heavily…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yuang Jia , Jinlong Wang , Jiayi Zhao , Chunlam Li , Shunzhou Wang , Wei Gao

Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Honggyu An , Jaewoo Jung , Mungyeom Kim , Chaehyun Kim , Minkyeong Jeon , Jisang Han , Kazumi Fukuda , Takuya Narihira , Hyuna Ko , Junsu Kim , Sunghwan Hong , Yuki Mitsufuji , Seungryong Kim

Recently, the integration of the efficient feed-forward scheme into 3D Gaussian Splatting (3DGS) has been actively explored. However, most existing methods focus on sparse view reconstruction of small regions and cannot produce eligible…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yunsong Wang , Tianxin Huang , Hanlin Chen , Gim Hee Lee

We present InstantStyleGaussian, an innovative 3D style transfer method based on the 3D Gaussian Splatting (3DGS) scene representation. By inputting a target-style image, it quickly generates new 3D GS scenes. Our method operates on…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Xin-Yi Yu , Jun-Xin Yu , Li-Bo Zhou , Yan Wei , Lin-Lin Ou

Synthesizing consistent and photorealistic 3D scenes is an open problem in computer vision. Video diffusion models generate impressive videos but cannot directly synthesize 3D representations, i.e., lack 3D consistency in the generated…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Katja Schwarz , Norman Mueller , Peter Kontschieder

Reconstructing 3D scenes from sparse viewpoints is a long-standing challenge with wide applications. Recent advances in feed-forward 3D Gaussian sparse-view reconstruction methods provide an efficient solution for real-time novel view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Yang Xiao , Guoan Xu , Qiang Wu , Wenjing Jia

3D Gaussian Splatting achieves high-fidelity novel view synthesis, but its application to online long-sequence scenarios is still limited. Existing methods either rely on slow per-scene optimization or fail to provide efficient incremental…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Guichen Huang , Ruoyu Wang , Xiangjun Gao , Che Sun , Yuwei Wu , Shenghua Gao , Yunde Jia

Dynamic scene reconstruction is a long-term challenge in the field of 3D vision. Recently, the emergence of 3D Gaussian Splatting has provided new insights into this problem. Although subsequent efforts rapidly extend static 3D Gaussian to…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Ruijie Zhu , Yanzhe Liang , Hanzhi Chang , Jiacheng Deng , Jiahao Lu , Wenfei Yang , Tianzhu Zhang , Yongdong Zhang

The online reconstruction of dynamic scenes from multi-view streaming videos faces significant challenges in training, rendering and storage efficiency. Harnessing superior learning speed and real-time rendering capabilities, 3D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Qiankun Gao , Jiarui Meng , Chengxiang Wen , Jie Chen , Jian Zhang

How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hanwen Liang , Junli Cao , Vidit Goel , Guocheng Qian , Sergei Korolev , Demetri Terzopoulos , Konstantinos N. Plataniotis , Sergey Tulyakov , Jian Ren

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chenhan Jiang , Yu Chen , Qingwen Zhang , Jifei Song , Songcen Xu , Dit-Yan Yeung , Jiankang Deng

Recent advances in diffusion models have revolutionized 2D and 3D content creation, yet generating photorealistic dynamic 4D scenes remains a significant challenge. Existing dynamic 4D generation methods typically rely on distilling…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Vinayak Gupta , Yunze Man , Yu-Xiong Wang

Evaluating and training autonomous driving systems require diverse and scalable corner cases. However, most existing scene generation methods lack controllability, accuracy, and versatility, resulting in unsatisfactory generation results.…

Robotics · Computer Science 2024-10-11 Sheng Wang , Ge Sun , Fulong Ma , Tianshuai Hu , Qiang Qin , Yongkang Song , Lei Zhu , Junwei Liang

Reconstructing and predicting dynamic 3D scenes from multi-view videos is a foundational task for robotics, AR/VR, and digital twins. Recent physics-informed Gaussian Splatting methods achieve impressive future frame extrapolation but lack…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Denis Gridusov , Maxim Popov , Sergey Kolyubin