English
Related papers

Related papers: Driving with DINO: Vision Foundation Features as a…

200 papers

We often aim to generate images that are both photorealistic and 3D-consistent, adhering to precise geometry, material, and viewpoint controls. Typically, this is achieved by fine-tuning an image generator, pre-trained on billions of real…

Graphics · Computer Science 2026-05-15 Ido Sobol , Kihyuk Sohn , Yoav Blum , Egor Zakharov , Max Bluvstein , Andrea Vedaldi , Or Litany

Diffusion models have fundamentally transformed the field of generative models, making the assessment of similarity between customized model outputs and reference inputs critically important. However, traditional perceptual similarity…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Yiren Song , Xiaokang Liu , Mike Zheng Shou

Simulation frameworks have been key enablers for the development and validation of autonomous driving systems. However, existing methods struggle to comprehensively address the autonomy-oriented requirements of balancing: (i) dynamical…

Robotics · Computer Science 2026-02-23 Tanmay Vilas Samak , Chinmay Vilas Samak , Bing Li , Venkat Krovi

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

A fundamental bottleneck in Novel View Synthesis (NVS) for autonomous driving is the inherent supervision gap on novel trajectories: models are tasked with synthesizing unseen views during inference, yet lack ground truth images for these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Hongbo Lu , Liang Yao , Chenghao He , Fan Liu , Wenlong Liao , Tao He , Pai Peng

Generative diffusion models for end-to-end autonomous driving often suffer from mode collapse, tending to generate conservative and homogeneous behaviors. While DiffusionDrive employs predefined anchors representing different driving…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jialv Zou , Shaoyu Chen , Bencheng Liao , Zhiyu Zheng , Yuehao Song , Lefei Zhang , Qian Zhang , Wenyu Liu , Xinggang Wang

Velocity-model building is a fundamental component of seismic imaging, yet it remains a challenging inverse problem due to limited data coverage, nonlinearity, and the need to integrate heterogeneous information such as well logs. We…

Geophysics · Physics 2026-03-03 Francesco Brandolin , Tariq Alkhalifah

LiDAR-based semantic segmentation is a key component for autonomous mobile robots, yet large-scale annotation of LiDAR point clouds is prohibitively expensive and time-consuming. Although simulators can provide labeled synthetic data,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Tomoya Miyawaki , Kazuto Nakashima , Yumi Iwashita , Ryo Kurazume

We present FlightDiffusion, a diffusion-model-based framework for training autonomous drones from first-person view (FPV) video. Our model generates realistic video sequences from a single frame, enriched with corresponding action spaces to…

The Vision-Language Foundation Model has recently shown outstanding performance in various perception learning tasks. The outstanding performance of the vision-language model mainly relies on large-scale pre-training datasets and different…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Thanh-Dat Truong , Xin Li , Bhiksha Raj , Jackson Cothren , Khoa Luu

Accurate and high-fidelity driving scene reconstruction demands the effective utilization of comprehensive scene information as conditional inputs. Existing methods predominantly rely on 3D bounding boxes and BEV road maps for foreground…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Zhao Yang , Zezhong Qian , Xiaofan Li , Weixiang Xu , Gongpeng Zhao , Ruohong Yu , Lingsi Zhu , Longjun Liu

Rapid autonomous traversal of unstructured terrain is essential for scenarios such as disaster response, search and rescue, or planetary exploration. As a vehicle navigates at the limit of its capabilities over extreme terrain, its dynamics…

Robotics · Computer Science 2024-12-03 Jason Gibson , Anoushka Alavilli , Erica Tevere , Evangelos A. Theodorou , Patrick Spieler

Detectors often suffer from degraded performance, primarily due to the distributional gap between the source and target domains. This issue is especially evident in single-source domains with limited data, as models tend to rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Mingbo Hong , Feng Liu , Caroline Gevaert , George Vosselman , Hao Cheng

Collaborative perception plays a crucial role in enhancing environmental understanding by expanding the perceptual range and improving robustness against sensor failures, which primarily involves collaborative 3D detection and tracking…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Xunjie He , Christina Dao Wen Lee , Meiling Wang , Chengran Yuan , Zefan Huang , Yufeng Yue , Marcelo H. Ang

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language…

Robotics · Computer Science 2025-09-25 Pengxiang Li , Yinan Zheng , Yue Wang , Huimin Wang , Hang Zhao , Jingjing Liu , Xianyuan Zhan , Kun Zhan , Xianpeng Lang

Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often struggle in high-dynamic- range scenes or high-speed scenarios due to motion blur and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Mustafa Sakhaia , Kaung Sithua , Min Khant Soe Okea , Maciej Wielgosza

Generating realistic and diverse road scenarios is essential for autonomous vehicle testing and validation. Nevertheless, owing to the complexity and variability of real-world road environments, creating authentic and varied scenarios for…

Robotics · Computer Science 2024-11-15 Junjie Zhou , Lin Wang , Qiang Meng , Xiaofan Wang

Open-vocabulary object detectors such as Grounding DINO are trained on vast and diverse data, achieving remarkable performance on challenging datasets. Due to that, it is unclear where to find their limitations, which is of major concern…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Annika Mütze , Sadia Ilyas , Christian Dörpelkus , Matthias Rottmann

Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based approaches excel in semantic generalization, they frequently lack the fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Haoxi Zeng , Qiankun Liu , Yi Bin , Haiyue Zhang , Yujuan Ding , Guoqing Wang , Deqiang Ouyang , Heng Tao Shen

Recently, the diffusion model has emerged as a powerful generative technique for robotic policy learning, capable of modeling multi-mode action distributions. Leveraging its capability for end-to-end autonomous driving is a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Bencheng Liao , Shaoyu Chen , Haoran Yin , Bo Jiang , Cheng Wang , Sixu Yan , Xinbang Zhang , Xiangyu Li , Ying Zhang , Qian Zhang , Xinggang Wang