中文
相关论文

相关论文: Towards Visual Foundational Models of Physical Sce…

200 篇论文

Although neural radiance fields (NeRF) have shown impressive advances for novel view synthesis, most methods typically require multiple input images of the same scene with accurate camera poses. In this work, we seek to substantially reduce…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Kai-En Lin , Lin Yen-Chen , Wei-Sheng Lai , Tsung-Yi Lin , Yi-Chang Shih , Ravi Ramamoorthi

The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to integrate the strengths of both approaches. In this work, we propose a methodology that…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Azmi Haider , Dan Rosenbaum

How do we know that a kitchen is a kitchen by looking? Relatively little is known about how we conceptualize and categorize different visual environments. Traditional models of visual perception posit that scene categorization is achieved…

神经元与认知 · 定量生物学 2014-11-20 Michelle R. Greene , Christopher Baldassano , Andre Esteva , Diane M. Beck , Li Fei-Fei

In this work, we aim to detect the changes caused by object variations in a scene represented by the neural radiance fields (NeRFs). Given an arbitrary view and two sets of scene images captured at different timestamps, we can predict the…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Rui Huang , Binbin Jiang , Qingyi Zhao , William Wang , Yuxiang Zhang , Qing Guo

Implicit neural representation has demonstrated promising results in 3D reconstruction on various scenes. However, existing approaches either struggle to model fast-moving objects or are incapable of handling large-scale camera ego-motions…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Tianchen Deng , Yanbo Wang , Yejia Liu , Chenpeng Su , Jingchuan Wang , Danwei Wang , Shao-Yuan Lo , Weidong Chen

Effective environment perception is crucial for enabling downstream robotic applications. Individual robotic agents often face occlusion and limited visibility issues, whereas multi-agent systems can offer a more comprehensive mapping of…

机器人学 · 计算机科学 2024-10-01 Hongrui Zhao , Boris Ivanovic , Negar Mehr

Diffusion Probabilistic Field (DPF) models the distribution of continuous functions defined over metric spaces. While DPF shows great potential for unifying data generation of various modalities including images, videos, and 3D geometry, it…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Kangfu Mei , Mo Zhou , Vishal M. Patel

Implicit representations such as Neural Radiance Fields (NeRF) have been shown to be very effective at novel view synthesis. However, these models typically require manual and careful human data collection for training. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Pierre Marza , Laetitia Matignon , Olivier Simonin , Dhruv Batra , Christian Wolf , Devendra Singh Chaplot

Neural radiance fields (NeRFs) have emerged as a prominent pre-training paradigm for vision-centric autonomous driving, which enhances 3D geometry and appearance understanding in a fully self-supervised manner. To apply NeRF-based…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Hyeonjun Jeong , Juyeb Shin , Dongsuk Kum

General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Hsiao-Yu Tung , Mingyu Ding , Zhenfang Chen , Daniel Bear , Chuang Gan , Joshua B. Tenenbaum , Daniel LK Yamins , Judith E Fan , Kevin A. Smith

Diffusion probabilistic models have quickly become a major approach for generative modeling of images, 3D geometry, video and other domains. However, to adapt diffusion generative modeling to these domains the denoising network needs to be…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Peiye Zhuang , Samira Abnar , Jiatao Gu , Alex Schwing , Joshua M. Susskind , Miguel Ángel Bautista

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Yaniv Nikankin , Niv Haim , Michal Irani

We propose a framework for learning neural scene representations directly from images, without 3D supervision. Our key insight is that 3D structure can be imposed by ensuring that the learned representation transforms like a real 3D scene.…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Emilien Dupont , Miguel Angel Bautista , Alex Colburn , Aditya Sankar , Carlos Guestrin , Josh Susskind , Qi Shan

Modelling individual objects in a scene as Neural Radiance Fields (NeRFs) provides an alternative geometric scene representation that may benefit downstream robotics tasks such as scene understanding and object manipulation. However, we…

机器人学 · 计算机科学 2022-10-10 Jad Abou-Chakra , Feras Dayoub , Niko Sünderhauf

Complex visual scenes that are composed of multiple objects, each with attributes, such as object name, location, pose, color, etc., are challenging to describe in order to train neural networks. Usually,deep learning networks are trained…

神经与进化计算 · 计算机科学 2023-03-27 E. Paxon Frady , Spencer Kent , Quinn Tran , Pentti Kanerva , Bruno A. Olshausen , Friedrich T. Sommer

Visual illusions in humans arise when interpreting out-of-distribution stimuli: if the observer is adapted to certain statistics, perception of outliers deviates from reality. Recent studies have shown that artificial neural networks (ANNs)…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Alex Gomez-Villa , Kai Wang , Alejandro C. Parraga , Bartlomiej Twardowski , Jesus Malo , Javier Vazquez-Corral , Joost van de Weijer

Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Nithesh Chandher Karthikeyan , Jonas Unger , Gabriel Eilertsen

We propose a novel visual re-localization method based on direct matching between the implicit 3D descriptors and the 2D image with transformer. A conditional neural radiance field(NeRF) is chosen as the 3D scene representation in our…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Jianlin Liu , Qiang Nie , Yong Liu , Chengjie Wang

Generating unbounded 3D scenes is crucial for large-scale scene understanding and simulation. Urban scenes, unlike natural landscapes, consist of various complex man-made objects and structures such as roads, traffic signs, vehicles, and…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Junge Zhang , Qihang Zhang , Li Zhang , Ramana Rao Kompella , Gaowen Liu , Bolei Zhou

Internal representations are crucial for understanding deep neural networks, such as their properties and reasoning patterns, but remain difficult to interpret. While mapping from feature space to input space aids in interpreting the…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Nils Neukirch , Johanna Vielhaben , Nils Strodthoff