English
Related papers

Related papers: Unified Map Prior Encoder for Mapping and Planning

200 papers

Universal Multimodal Retrieval (UMR) aims to enable search across various modalities using a unified model, where queries and candidates can consist of pure text, images, or a combination of both. Previous work has attempted to adopt…

Computation and Language · Computer Science 2025-04-02 Xin Zhang , Yanzhao Zhang , Wen Xie , Mingxin Li , Ziqi Dai , Dingkun Long , Pengjun Xie , Meishan Zhang , Wenjie Li , Min Zhang

Bird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Yiwei Zhang , Jin Gao , Fudong Ge , Guan Luo , Bing Li , Zhaoxiang Zhang , Haibin Ling , Weiming Hu

Many recent inpainting works have achieved impressive results by leveraging Deep Neural Networks (DNNs) to model various prior information for image restoration. Unfortunately, the performance of these methods is largely limited by the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Chenjie Cao , Qiaole Dong , Yanwei Fu

In recent years, the field of machine learning has made phenomenal progress in the pursuit of simulating real-world data generation processes. One notable example of such success is the variational autoencoder (VAE). In this work, with a…

Machine Learning · Statistics 2021-12-30 Hwan Goh , Sheroze Sheriffdeen , Jonathan Wittmer , Tan Bui-Thanh

With the exponential growth of multimedia data, leveraging multimodal sensors presents a promising approach for improving accuracy in human activity recognition. Nevertheless, accurately identifying these activities using both video data…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Rex Liu , Xin Liu

Purpose: To introduce a combined machine learning (ML) and physics-based image reconstruction framework that enables navigator-free, highly accelerated multishot echo planar imaging (msEPI), and demonstrate its application in…

Visual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in leveraging high-quality pre-trained visual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yuan Gao , Chen Chen , Tianrong Chen , Jiatao Gu

4D millimeter-wave (MMW) radar, which provides both height information and dense point cloud data over 3D MMW radar, has become increasingly popular in 3D object detection. In recent years, radar-vision fusion models have demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Haocheng Zhao , Runwei Guan , Taoyu Wu , Ka Lok Man , Limin Yu , Yutao Yue

We propose and demonstrate a fast, robust method for using satellite images to localize an Unmanned Aerial Vehicle (UAV). Previous work using satellite images has large storage and computation costs and is unable to run in real time. In…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Mollie Bianchi , Timothy D. Barfoot

Image dehazing has witnessed significant advancements with the development of deep learning models. However, most existing methods focus solely on single-modal RGB features, neglecting the inherent correlation between scene depth and haze…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zengyuan Zuo , Junjun Jiang , Gang Wu , Xianming Liu

Masked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation strategies for MAE still remain open questions, different…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Kai Chen , Zhili Liu , Lanqing Hong , Hang Xu , Zhenguo Li , Dit-Yan Yeung

In the field of autonomous driving, Bird's-Eye-View (BEV) perception has attracted increasing attention in the community since it provides more comprehensive information compared with pinhole front-view images and panoramas. Traditional BEV…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Jiale Wei , Junwei Zheng , Ruiping Liu , Jie Hu , Jiaming Zhang , Rainer Stiefelhagen

Although Vision Transformers (ViTs) have recently advanced computer vision tasks significantly, an important real-world problem was overlooked: adapting to variable input resolutions. Typically, images are resized to a fixed resolution,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Wenzhuo Liu , Fei Zhu , Shijie Ma , Cheng-Lin Liu

Fast and efficient motion planning algorithms are crucial for many state-of-the-art robotics applications such as self-driving cars. Existing motion planning methods become ineffective as their computational complexity increases…

Robotics · Computer Science 2019-02-26 Ahmed H. Qureshi , Anthony Simeonov , Mayur J. Bency , Michael C. Yip

Over-the-air computation (AirComp) seamlessly integrates communication and computation by exploiting the waveform superposition property of multiple-access channels. Different from the existing works that focus on transceiver design of…

Signal Processing · Electrical Eng. & Systems 2021-06-02 Min Fu , Yong Zhou , Yuanming Shi , Wei Chen , Rui Zhang

Neural shape representation generally refers to representing 3D geometry using neural networks, e.g., computing a signed distance or occupancy value at a specific spatial position. In this paper we present a neural-network architecture…

Machine Learning · Computer Science 2024-08-22 Stefan Rhys Jeske , Jonathan Klein , Dominik L. Michels , Jan Bender

Using a deep autoencoder (DAE) for end-to-end communication in multiple-input multiple-output (MIMO) systems is a novel concept with significant potential. DAE-aided MIMO has been shown to outperform singular-value decomposition (SVD)-based…

Information Theory · Computer Science 2022-02-14 Xinliang Zhang , Mojtaba Vaezi , Timothy J. O'Shea

Building a multi-modality multi-task neural network toward accurate and robust performance is a de-facto standard in perception task of autonomous driving. However, leveraging such data from multiple sensors to jointly optimize the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Tengju Ye , Wei Jing , Chunyong Hu , Shikun Huang , Lingping Gao , Fangzhen Li , Jingke Wang , Ke Guo , Wencong Xiao , Weibo Mao , Hang Zheng , Kun Li , Junbo Chen , Kaicheng Yu

Mixture of Vision Encoders (MoVE) has emerged as a powerful approach to enhance the fine-grained visual understanding of multimodal large language models (MLLMs), improving their ability to handle tasks such as complex optical character…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Mozhgan Nasr Azadani , James Riddell , Sean Sedwards , Krzysztof Czarnecki

Hybrid Mamba-Transformer networks have recently garnered broad attention. These networks can leverage the scalability of Transformers while capitalizing on Mamba's strengths in long-context modeling and computational efficiency. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Yunze Liu , Li Yi