English
Related papers

Related papers: GoRela: Go Relative for Viewpoint-Invariant Motion…

200 papers

As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Zhuofan Zong , Bingqi Ma , Dazhong Shen , Guanglu Song , Hao Shao , Dongzhi Jiang , Hongsheng Li , Yu Liu

This paper studies the problem of multi-agent cooperative localization of a common reference coordinate frame in $\mathbb{R}^3$. Each agent in a system maintains a body-fixed coordinate frame and its actual \textit{frame transformation}…

Optimization and Control · Mathematics 2019-07-11 Quoc Van Tran , Hyo-Sung Ahn

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Jingjing Qian , Boyao Han , Chen Shi , Lei Xiao , Long Yang , Shaoshuai Shi , Li Jiang

Estimating vehicles' locations is one of the key components in intelligent traffic management systems (ITMSs) for increasing traffic scene awareness. Traditionally, stationary sensors have been employed in this regard. The development of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Elnaz Namazi , Rudolf Mester , Chaoru Lu , Jingyue Li

Accurately recognizing a revisited place is crucial for embodied agents to localize and navigate. This requires visual representations to be distinct, despite strong variations in camera viewpoint and scene appearance. Existing visual place…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Kartik Garg , Sai Shubodh Puligilla , Shishir Kolathaya , Madhava Krishna , Sourav Garg

In this paper, we explore the use of vehicle-to-vehicle (V2V) communication to improve the perception and motion forecasting performance of self-driving vehicles. By intelligently aggregating the information received from multiple nearby…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Tsun-Hsuan Wang , Sivabalan Manivasagam , Ming Liang , Bin Yang , Wenyuan Zeng , James Tu , Raquel Urtasun

Visual Teach-and-Repeat Navigation is a direct solution for mobile robot to be deployed in unknown environments. However, robust trajectory repeat navigation still remains challenged due to environmental changing and dynamic objects. In…

Robotics · Computer Science 2025-10-13 Jikai Wang , Yunqi Cheng , Kezhi Wang , Zonghai Chen

In autonomous driving, deep learning enabled motion prediction is a popular topic. A critical gap in traditional motion prediction methodologies lies in ensuring equivariance under Euclidean geometric transformations and maintaining…

Robotics · Computer Science 2025-08-05 Yuping Wang , Jier Chen

The robot exploration task has been widely studied with applications spanning from novel environment mapping to item delivery. For some time-critical tasks, such as rescue catastrophes, the agent is required to explore as efficiently as…

Robotics · Computer Science 2023-08-01 Xuyang Chen , Ashvin N. Iyer , Zixing Wang , Ahmed H. Qureshi

Predicting future behaviors of road agents is a key task in autonomous driving. While existing models have demonstrated great success in predicting marginal agent future behaviors, it remains a challenge to efficiently predict consistent…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Xin Huang , Xiaoyu Tian , Junru Gu , Qiao Sun , Hang Zhao

While Vision-language models (VLMs) have demonstrated remarkable performance across multi-modal tasks, their choice of vision encoders presents a fundamental weakness: their low-level features lack the robust structural and spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Brandon Huang , Hang Hua , Zhuoran Yu , Trevor Darrell , Rogerio Feris , Roei Herzig

Accurately modeling agent behaviors is an important task in self-driving. It is also a task with many symmetries, such as equivariance to the order of agents and objects in the scene or equivariance to arbitrary roto-translations of the…

Robotics · Computer Science 2026-04-03 Scott Xu , Dian Chen , Kelvin Wong , Chris Zhang , Kion Fallah , Raquel Urtasun

Program code serves as a bridge linking vision and logic, providing a feasible supervisory approach for enhancing the multimodal reasoning capability of large models through geometric operations such as auxiliary line construction and…

Artificial Intelligence · Computer Science 2026-02-10 Zhenyu Wu , Yanxi Long , Jian Li , Hua Huang

Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Mingkun Wang , Xiaoguang Ren , Ruochun Jin , Minglong Li , Xiaochuan Zhang , Changqian Yu , Mingxu Wang , Wenjing Yang

While Vision-Language Models (VLMs) show significant promise for end-to-end autonomous driving by leveraging the common sense embedded in language models, their reliance on 2D image cues for complex scene understanding and decision-making…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Weijie Wei , Zhipeng Luo , Ling Feng , Venice Erin Liong

Long-term visual localization in outdoor environment is a challenging problem, especially faced with the cross-seasonal, bi-directional tasks and changing environment. In this paper we propose a novel visual inertial localization framework…

Robotics · Computer Science 2018-03-06 Xiaqing Ding , Yue Wang , Dongxuan Li , Li Tang , Huan Yin , Rong Xiong

Data driven approaches for decision making applied to automated driving require appropriate generalization strategies, to ensure applicability to the world's variability. Current approaches either do not generalize well beyond the training…

Machine Learning · Computer Science 2022-03-11 Karl Kurzer , Philip Schörner , Alexander Albers , Hauke Thomsen , Karam Daaboul , J. Marius Zöllner

Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing approaches commonly rely on a single pre-trained vision…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Matvey Skripkin , Elizaveta Goncharova , Dmitrii Tarasov , Andrey Kuznetsov

Forecasting the long-term future motion of road actors is a core challenge to the deployment of safe autonomous vehicles (AVs). Viable solutions must account for both the static geometric context, such as road lanes, and dynamic social…

Machine Learning · Computer Science 2020-08-25 Siddhesh Khandelwal , William Qi , Jagjeet Singh , Andrew Hartnett , Deva Ramanan

Urban forecasting has increasingly benefited from high-dimensional spatial data through two primary approaches: graph-based methods that rely on predefined spatial structures, and region-based methods that focus on learning expressive urban…

Artificial Intelligence · Computer Science 2025-06-18 Yuhao Jia , Zile Wu , Shengao Yi , Yifei Sun , Xiao Huang
‹ Prev 1 3 4 5 6 7 10 Next ›