English
Related papers

Related papers: Regulating Intermediate 3D Features for Vision-Cen…

200 papers

We introduce DiffPhysCam, a differentiable camera simulator designed to support robotics and embodied AI applications by enabling gradient-based optimization in visual perception pipelines. Generating synthetic images that closely mimic…

Graphics · Computer Science 2025-08-13 Bo-Hsun Chen , Nevindu M. Batagoda , Dan Negrut

Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these modalities face a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xuyang Chen , Conglang Zhang , Chuanheng Fu , Zihao Yang , Kaixuan Zhou , Yizhi Zhang , Jianan He , Yanfeng Zhang , Mingwei Sun , Zengmao Wang , Zhen Dong , Xiaoxiao Long , Liqiu Meng

Reconstructing 3D objects from images is inherently an ill-posed problem due to ambiguities in geometry, appearance, and topology. This paper introduces collaborative inverse rendering with persistent homology priors, a novel strategy that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xiang Gao , Xinmu Wang , Yuanpeng Liu , Yue Wang , Junqi Huang , Wei Chen , Xianfeng Gu

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge this gap often suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yujie Wei , Xinyu Liu , Shiwei Zhang , Hangjie Yuan , Jinbo Xing , Zhekai Chen , Xiang Wang , Haonan Qiu , Rui Zhao , Yutong Feng , Ruihang Chu , Yingya Zhang , Yike Guo , Xihui Liu , Hongming Shan

Unsigned distance functions (UDFs) have been a vital representation for open surfaces. With different differentiable renderers, current methods are able to train neural networks to infer a UDF by minimizing the rendering errors with the UDF…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Wenyuan Zhang , Chunsheng Wang , Kanle Shi , Yu-Shen Liu , Zhizhong Han

While weakly supervised multi-view face reconstruction (MVR) is garnering increased attention, one critical issue still remains open: how to effectively interact and fuse multiple image information to reconstruct high-precision 3D models.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Weiguang Zhao , Chaolong Yang , Jianan Ye , Rui Zhang , Yuyao Yan , Xi Yang , Bin Dong , Amir Hussain , Kaizhu Huang

We present a technique to synthesize and analyze volume-rendered images using generative models. We use the Generative Adversarial Network (GAN) framework to compute a model from a large collection of volume renderings, conditioned on (1)…

Graphics · Computer Science 2019-07-18 Matthew Berger , Jixian Li , Joshua A. Levine

High-fidelity and controllable 3D simulation is essential for addressing the long-tail data scarcity in Autonomous Driving (AD), yet existing methods struggle to simultaneously achieve photorealistic rendering and interactive traffic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Zhiyuan Liu , Daocheng Fu , Pinlong Cai , Lening Wang , Ying Liu , Yilong Ren , Botian Shi , Jianqiang Wang

3D semantic occupancy prediction aims to forecast detailed geometric and semantic information of the surrounding environment for autonomous vehicles (AVs) using onboard surround-view cameras. Existing methods primarily focus on intricate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Cross-modal systems trained on 2D visual inputs are presented with a dimensional shift when processing 3D scenes. An in-scene camera bridges the dimensionality gap but requires learning a control module. We introduce a new method that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Jason Armitage , Rico Sennnrich

3D occupancy, an advanced perception technology for driving scenarios, represents the entire scene without distinguishing between foreground and background by quantifying the physical space into a grid map. The widely adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Jinke Li , Xiao He , Chonghua Zhou , Xiaoqiang Cheng , Yang Wen , Dan Zhang

Transfer Function (TF) generation is a fundamental problem in Direct Volume Rendering (DVR). A TF maps voxels to color and opacity values to reveal inner structures. Existing TF tools are complex and unintuitive for the users who are more…

Graphics · Computer Science 2017-12-01 Naimul Khan , Riadh Ksantini , Ling Guan

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Yunuo Chen , Junli Cao , Vidit Goel , Sergei Korolev , Chenfanfu Jiang , Jian Ren , Sergey Tulyakov , Anil Kag

We propose DistillNeRF, a self-supervised learning framework addressing the challenge of understanding 3D environments from limited 2D observations in outdoor autonomous driving scenes. Our method is a generalizable feedforward model that…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Letian Wang , Seung Wook Kim , Jiawei Yang , Cunjun Yu , Boris Ivanovic , Steven L. Waslander , Yue Wang , Sanja Fidler , Marco Pavone , Peter Karkus

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Providing a depth-rich Virtual Reality (VR) experience to users without causing discomfort remains to be a challenge with today's commercially available head-mounted displays (HMDs), which enforce strict measures on stereoscopic camera…

Graphics · Computer Science 2019-11-12 Emre Avan , Ufuk Celikcan , Tolga K. Capin , Hasmet Gurcay

Obtaining high-quality 3D reconstructions of room-scale scenes is of paramount importance for upcoming applications in AR or VR. These range from mixed reality applications for teleconferencing, virtual measuring, virtual room planing, to…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 Dejan Azinović , Ricardo Martin-Brualla , Dan B Goldman , Matthias Nießner , Justus Thies

This paper aims to address the challenge of reconstructing long volumetric videos from multi-view RGB videos. Recent dynamic view synthesis methods leverage powerful 4D representations, like feature grids or point cloud sequences, to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Zhen Xu , Yinghao Xu , Zhiyuan Yu , Sida Peng , Jiaming Sun , Hujun Bao , Xiaowei Zhou

Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sparse action supervision, which underutilizes their powerful scene understanding and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Xiaodong Mei , Diankun Zhang , Hongwei Xie , Guang Chen , Hangjun Ye , Dan Xu

Recent advancements in diffusion models have significantly enhanced the data synthesis with 2D control. Yet, precise 3D control in street view generation, crucial for 3D perception tasks, remains elusive. Specifically, utilizing Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Ruiyuan Gao , Kai Chen , Enze Xie , Lanqing Hong , Zhenguo Li , Dit-Yan Yeung , Qiang Xu