中文
相关论文

相关论文: Veila: Panoramic LiDAR Generation from a Monocular…

200 篇论文

Layout design generation has recently gained significant attention due to its potential applications in various fields, including UI, graphic, and floor plan design. However, existing models face two main challenges that limits their…

人机交互 · 计算机科学 2024-05-24 Chin-Yi Cheng , Ruiqi Gao , Forrest Huang , Yang Li

We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle this challenging problem with a scalable yet powerful approach,…

Visual navigation, a fundamental challenge in mobile robotics, demands versatile policies to handle diverse environments. Classical methods leverage geometric solutions to minimize specific costs, offering adaptability to new scenarios but…

机器人学 · 计算机科学 2025-04-15 Yiming Zeng , Hao Ren , Shuhang Wang , Junlong Huang , Hui Cheng

Calibration is an essential prerequisite for the accurate data fusion of LiDAR and camera sensors. Traditional calibration techniques often require specific targets or suitable scenes to obtain reliable 2D-3D correspondences. To tackle the…

计算机视觉与模式识别 · 计算机科学 2025-01-29 Shujuan Huang , Chunyu Lin , Yao Zhao

Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusion vision-language models (DVLMs) have shown strong…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Shrinidhi Kumbhar , Haofu Liao , Srikar Appalaraju , Kunwar Yashraj Singh

Vision-based depth estimation is a key feature in autonomous systems, which often relies on a single camera or several independent ones. In such a monocular setup, dense depth is obtained with either additional input from one or several…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Florent Bartoccioni , Éloi Zablocki , Patrick Pérez , Matthieu Cord , Karteek Alahari

The complementary fusion of light detection and ranging (LiDAR) data and image data is a promising but challenging task for generating high-precision and high-density point clouds. This study proposes an innovative LiDAR-guided stereo…

计算机视觉与模式识别 · 计算机科学 2022-02-25 Yongjun Zhang , Siyuan Zou , Xinyi Liu , Xu Huang , Yi Wan , Yongxiang Yao

Mobile robots and autonomous vehicles rely on multi-modal sensor setups to perceive and understand their surroundings. Aside from cameras, LiDAR sensors represent a central component of state-of-the-art perception systems. In addition to…

计算机视觉与模式识别 · 计算机科学 2018-04-27 Florian Piewak , Peter Pinggera , Manuel Schäfer , David Peter , Beate Schwarz , Nick Schneider , David Pfeiffer , Markus Enzweiler , Marius Zöllner

Light detection and ranging (LiDAR) is a ubiquitous tool to provide precise spatial awareness in various perception environments. A bionic LiDAR that can mimic human-like vision by adaptively gazing at selected regions of interest within a…

While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally constrained by environmental variability. Specifically, camera sensors suffer from severe…

计算机视觉与模式识别 · 计算机科学 2026-05-18 A. Enes Doruk , Abdelaziz Hussein , Hasan F. Ates

Vision-language-action (VLA) models have shown strong generalization for robotic action prediction through large-scale vision-language pretraining. However, most existing models rely solely on RGB cameras, limiting their perception and,…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Heyu Guo , Shanmu Wang , Ruichun Ma , Shiqi Jiang , Yasaman Ghasempour , Omid Abari , Baining Guo , Lili Qiu

Efficiently predicting motion plans directly from vision remains a fundamental challenge in robotics, where planning typically requires explicit goal specification and task-specific design. Recent vision-language-action (VLA) models infer…

LiDARs and cameras are the two main sensors that are planned to be included in many announced autonomous vehicles prototypes. Each of the two provides a unique form of data from a different perspective to the surrounding environment. In…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Amr S. Mohamed , Ali Abdelkader , Mohamed Anany , Omar El-Behady , Muhammad Faisal , Asser Hangal , Hesham M. Eraqi , Mohamed N. Moustafa

In this work, we propose a novel approach, namely WeatherDG, that can generate realistic, weather-diverse, and driving-screen images based on the cooperation of two foundation models, i.e, Stable Diffusion (SD) and Large Language Model…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Chenghao Qian , Yuhu Guo , Yuhong Mo , Wenjing Li

Holographic wave-shaping has found numerous applications across the physical sciences, especially since the development of digital spatial-light modulators (SLMs). A key challenge in digital holography consists in finding optimal hologram…

图像与视频处理 · 电气工程与系统科学 2019-11-05 Jannes Gladrow

We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Abdelrahman Eldesokey , Aleksandar Cvejic , Bernard Ghanem , Peter Wonka

The rapid progress of Multimodal Large Language Models (MLLMs) has unlocked the potential for enhanced 3D scene understanding and spatial reasoning. A recent line of work explores learning spatial reasoning directly from multi-view images,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Kanghee Lee , Injae Lee , Minseok Kwak , Jungi Hong , Kwonyoung Ryu , Jaesik Park

Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Yuzhen Li , Min Liu , Zhaoyang Li , Yuan Bian , Xueping Wang , Erbo Zhai , Yaonan Wang

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Shufan Li , Konstantinos Kallidromitis , Hritik Bansal , Akash Gokul , Yusuke Kato , Kazuki Kozuka , Jason Kuen , Zhe Lin , Kai-Wei Chang , Aditya Grover

We introduce MVControl, a novel neural network architecture that enhances existing pre-trained multi-view 2D diffusion models by incorporating additional input conditions, e.g. edge maps. Our approach enables the generation of controllable…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Zhiqi Li , Yiming Chen , Lingzhe Zhao , Peidong Liu
‹ 上一页 1 8 9 10 下一页 ›