中文
相关论文

相关论文: PanoDreamer: Consistent Text to 360-Degree Scene G…

200 篇论文

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

We introduce 3inGAN, an unconditional 3D generative model trained from 2D images of a single self-similar 3D scene. Such a model can be used to produce 3D "remixes" of a given scene, by mapping spatial latent codes into a 3D volumetric…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Animesh Karnewar , Oliver Wang , Tobias Ritschel , Niloy Mitra

Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Qiao Yu , Xianzhi Li , Yuan Tang , Xu Han , Jinfeng Xu , Long Hu , Min Chen

Panoramic image enables deeper understanding and more holistic perception of $360^\circ$ surrounding environment, which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made…

计算机视觉与模式识别 · 计算机科学 2023-06-06 Yuan Dong , Chuan Fang , Liefeng Bo , Zilong Dong , Ping Tan

We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face key challenges: 2D-lifting approaches suffer from geometric…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Sikuang Li , Chen Yang , Jiemin Fang , Taoran Yi , Jia Lu , Jiazhong Cen , Lingxi Xie , Wei Shen , Qi Tian

3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reasoning. Traditional supervised models leverage explicit 3D geometry but exhibit limited…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Seongmin Jung , Seongho Choi , Gunwoo Jeon , Minsu Cho , Jongwoo Lim

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

机器人学 · 计算机科学 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

We propose a systematic learning-based approach to the generation of massive quantities of synthetic 3D scenes and arbitrary numbers of photorealistic 2D images thereof, with associated ground truth information, for the purposes of…

计算机视觉与模式识别 · 计算机科学 2018-06-21 Chenfanfu Jiang , Siyuan Qi , Yixin Zhu , Siyuan Huang , Jenny Lin , Lap-Fai Yu , Demetri Terzopoulos , Song-Chun Zhu

Vision-and-Language Navigation (VLN) requires the agent to follow language instructions to navigate through 3D environments. One main challenge in VLN is the limited availability of photorealistic training environments, which makes it hard…

计算机视觉与模式识别 · 计算机科学 2023-05-31 Jialu Li , Mohit Bansal

Generating high-resolution images with generative models has recently been made widely accessible by leveraging diffusion models pre-trained on large-scale datasets. Various techniques, such as MultiDiffusion and SyncDiffusion, have further…

计算机视觉与模式识别 · 计算机科学 2025-01-08 Stanislav Frolov , Brian B. Moser , Andreas Dengel

Continual learning refers to the ability of humans and animals to incrementally learn over time in a given environment. Trying to simulate this learning process in machines is a challenging task, also due to the inherent difficulty in…

计算机视觉与模式识别 · 计算机科学 2021-09-17 Enrico Meloni , Alessandro Betti , Lapo Faggi , Simone Marullo , Matteo Tiezzi , Stefano Melacci

Synthesizing photo-realistic visual observations from an ego vehicle's driving trajectory is a critical step towards scalable training of self-driving models. Reconstruction-based methods create 3D scenes from driving logs and synthesize…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Jiageng Mao , Boyi Li , Boris Ivanovic , Yuxiao Chen , Yan Wang , Yurong You , Chaowei Xiao , Danfei Xu , Marco Pavone , Yue Wang

This is a technical report on the 360-degree panoramic image generation task based on diffusion models. Unlike ordinary 2D images, 360-degree panoramic images capture the entire $360^\circ\times 180^\circ$ field of view. So the rightmost…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Mengyang Feng , Jinlin Liu , Miaomiao Cui , Xuansong Xie

Controllable generative models for images and videos have seen significant success, yet 3D scene generation, especially in unbounded scenarios like autonomous driving, remains underdeveloped. Existing methods lack flexible controllability…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Ruiyuan Gao , Kai Chen , Zhihao Li , Lanqing Hong , Zhenguo Li , Qiang Xu

We introduce "ImageDream," an innovative image-prompt, multi-view diffusion model for 3D object generation. ImageDream stands out for its ability to produce 3D models of higher quality compared to existing state-of-the-art,…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Peng Wang , Yichun Shi

3D Gaussian Splatting has demonstrated superior performance in rendering efficiency and quality, yet the generation of 3D Gaussians still remains a challenge without proper geometric priors. Existing methods have explored predicting point…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Weiqi Zhang , Junsheng Zhou , Haotian Geng , Kanle Shi , Shenkun Xu , Yi Fang , Yu-Shen Liu

Recently, 3D Gaussian Splatting (3DGS) has shown encouraging performance for open vocabulary scene understanding tasks. However, previous methods cannot distinguish 3D instance-level information, which usually predicts a heatmap between the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Hongjia Zhai , Hai Li , Zhenzhe Li , Xiaokun Pan , Yijia He , Guofeng Zhang

3D LiDAR scanners are playing an increasingly important role in autonomous driving as they can generate depth information of the environment. However, creating large 3D LiDAR point cloud datasets with point-level labels requires a…

计算机视觉与模式识别 · 计算机科学 2018-04-03 Xiangyu Yue , Bichen Wu , Sanjit A. Seshia , Kurt Keutzer , Alberto L. Sangiovanni-Vincentelli

Understanding geometric, semantic, and instance information in 3D scenes from sequential video data is essential for applications in robotics and augmented reality. However, existing Simultaneous Localization and Mapping (SLAM) methods…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Runnan Chen , Zhaoqing Wang , Jiepeng Wang , Yuexin Ma , Mingming Gong , Wenping Wang , Tongliang Liu

With the advent of portable 360{\deg} cameras, panorama has gained significant attention in applications like virtual reality (VR), virtual tours, robotics, and autonomous driving. As a result, wide-baseline panorama view synthesis has…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Cheng Zhang , Haofei Xu , Qianyi Wu , Camilo Cruz Gambardella , Dinh Phung , Jianfei Cai