中文
相关论文

相关论文: BEVDiffuser: Plug-and-Play Diffusion Model for BEV…

200 篇论文

An accurate initial heading angle is essential for efficient and safe navigation across diverse domains. Unlike magnetometers, gyroscopes can provide accurate heading reference independent of the magnetic disturbances in a process known as…

机器人学 · 计算机科学 2025-07-30 Gershy Ben-Arie , Daniel Engelsman , Rotem Dror , Itzik Klein

Localization in GNSS-denied and GNSS-degraded environments is a challenge for the safe widespread deployment of autonomous vehicles. Such GNSS-challenged environments require alternative methods for robust localization. In this work, we…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Shounak Sural , Ragunathan Rajkumar

Talk2BEV is a large vision-language model (LVLM) interface for bird's-eye view (BEV) maps in autonomous driving contexts. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set…

This paper introduces TopoDiffuser, a diffusion-based framework for multimodal trajectory prediction that incorporates topometric maps to generate accurate, diverse, and road-compliant future motion forecasts. By embedding structural cues…

机器人学 · 计算机科学 2025-08-04 Zehui Xu , Junhui Wang , Yongliang Shi , Chao Gao , Guyue Zhou

Denoising diffusion models trained at web-scale have revolutionized image generation. The application of these tools to engineering design is an intriguing possibility, but is currently limited by their inability to parse and enforce…

机器学习 · 计算机科学 2023-06-19 Nikos Arechiga , Frank Permenter , Binyang Song , Chenyang Yuan

The goal of speech enhancement (SE) is to eliminate the background interference from the noisy speech signal. Generative models such as diffusion models (DM) have been applied to the task of SE because of better generalization in unseen…

声音 · 计算机科学 2023-09-06 Wen Wang , Dongchao Yang , Qichen Ye , Bowen Cao , Yuexian Zou

Existing traffic simulation models often fall short in capturing the intricacies of real-world scenarios, particularly the interactive behaviors among multiple traffic participants, thereby limiting their utility in the evaluation and…

机器人学 · 计算机科学 2026-02-03 Zhiyu Huang , Zixu Zhang , Ameya Vaidya , Yuxiao Chen , Chen Lv , Jaime Fernández Fisac

The application of vision-based multi-view environmental perception system has been increasingly recognized in autonomous driving technology, especially the BEV-based models. Current state-of-the-art solutions primarily encode image…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Di Wu , Feng Yang , Benlian Xu , Pan Liao , Wenhui Zhao , Dingwen Zhang

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. However, we argue…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Zeming Chen , Hang Zhao

Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by…

机器人学 · 计算机科学 2025-12-09 Xiujian Liang , Jiacheng Liu , Mingyang Sun , Qichen He , Cewu Lu , Jianhua Sun

Bird's-eye-view (BEV) is a powerful and widely adopted representation for road scenes that captures surrounding objects and their spatial locations, along with overall context in the scene. In this work, we focus on bird's eye semantic…

计算机视觉与模式识别 · 计算机科学 2020-06-24 Mong H. Ng , Kaahan Radia , Jianfei Chen , Dequan Wang , Ionel Gog , Joseph E. Gonzalez

Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Shuhong Zheng , Zhipeng Bao , Ruoyu Zhao , Martial Hebert , Yu-Xiong Wang

We present FlightDiffusion, a diffusion-model-based framework for training autonomous drones from first-person view (FPV) video. Our model generates realistic video sequences from a single frame, enriched with corresponding action spaces to…

Previous raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Yufei Wang , Yi Yu , Wenhan Yang , Lanqing Guo , Lap-Pui Chau , Alex C. Kot , Bihan Wen

Goal-driven mobile robot navigation in map-less environments requires effective state representations for reliable decision-making. Inspired by the favorable properties of Bird's-Eye View (BEV) in point clouds for visual perception, this…

机器人学 · 计算机科学 2024-09-04 Jiahao Jiang , Yuxiang Yang , Yingqi Deng , Chenlong Ma , Jing Zhang

Existing approaches to drone visual geo-localization predominantly adopt the image-based setting, where a single drone-view snapshot is matched with images from other platforms. Such task formulation, however, underutilizes the inherent…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Hao Ju , Shaofei Huang , Si Liu , Zhedong Zheng

Multi-modal sensor fusion in Bird's Eye View (BEV) representation has become the leading approach for 3D object detection. However, existing methods often rely on depth estimators or transformer encoders to transform image features into BEV…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Yongjin Lee , Hyeon-Mun Jeong , Yurim Jeon , Sanghyun Kim

Training deep neural networks has become a common approach for addressing image restoration problems. An alternative for training a "task-specific" network for each observation model is to use pretrained deep denoisers for imposing only the…

图像与视频处理 · 电气工程与系统科学 2024-04-16 Tomer Garber , Tom Tirer

Diffusion models have shown exceptional scaling properties in the image synthesis domain, and initial attempts have shown similar benefits for applying diffusion to unconditional text synthesis. Denoising diffusion models attempt to…

音频与语音处理 · 电气工程与系统科学 2022-10-17 Matthew Baas , Kevin Eloff , Herman Kamper

Discrete diffusion models have emerged as a promising direction for vision-language tasks, offering bidirectional context modeling and theoretical parallelization. However, their practical application is severely hindered by a…

计算与语言 · 计算机科学 2025-10-24 Yatai Ji , Teng Wang , Yuying Ge , Zhiheng Liu , Sidi Yang , Ying Shan , Ping Luo
‹ 上一页 1 8 9 10 下一页 ›