中文
相关论文

相关论文: GoRela: Go Relative for Viewpoint-Invariant Motion…

200 篇论文

As the key component in multimodal large language models (MLLMs), the ability of the visual encoder greatly affects MLLM's understanding on diverse image content. Although some large-scale pretrained vision encoders such as vision encoders…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Zhuofan Zong , Bingqi Ma , Dazhong Shen , Guanglu Song , Hao Shao , Dongzhi Jiang , Hongsheng Li , Yu Liu

This paper studies the problem of multi-agent cooperative localization of a common reference coordinate frame in $\mathbb{R}^3$. Each agent in a system maintains a body-fixed coordinate frame and its actual \textit{frame transformation}…

最优化与控制 · 数学 2019-07-11 Quoc Van Tran , Hyo-Sung Ahn

Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Jingjing Qian , Boyao Han , Chen Shi , Lei Xiao , Long Yang , Shaoshuai Shi , Li Jiang

Estimating vehicles' locations is one of the key components in intelligent traffic management systems (ITMSs) for increasing traffic scene awareness. Traditionally, stationary sensors have been employed in this regard. The development of…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Elnaz Namazi , Rudolf Mester , Chaoru Lu , Jingyue Li

Accurately recognizing a revisited place is crucial for embodied agents to localize and navigate. This requires visual representations to be distinct, despite strong variations in camera viewpoint and scene appearance. Existing visual place…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Kartik Garg , Sai Shubodh Puligilla , Shishir Kolathaya , Madhava Krishna , Sourav Garg

In this paper, we explore the use of vehicle-to-vehicle (V2V) communication to improve the perception and motion forecasting performance of self-driving vehicles. By intelligently aggregating the information received from multiple nearby…

计算机视觉与模式识别 · 计算机科学 2020-08-18 Tsun-Hsuan Wang , Sivabalan Manivasagam , Ming Liang , Bin Yang , Wenyuan Zeng , James Tu , Raquel Urtasun

Visual Teach-and-Repeat Navigation is a direct solution for mobile robot to be deployed in unknown environments. However, robust trajectory repeat navigation still remains challenged due to environmental changing and dynamic objects. In…

机器人学 · 计算机科学 2025-10-13 Jikai Wang , Yunqi Cheng , Kezhi Wang , Zonghai Chen

In autonomous driving, deep learning enabled motion prediction is a popular topic. A critical gap in traditional motion prediction methodologies lies in ensuring equivariance under Euclidean geometric transformations and maintaining…

机器人学 · 计算机科学 2025-08-05 Yuping Wang , Jier Chen

The robot exploration task has been widely studied with applications spanning from novel environment mapping to item delivery. For some time-critical tasks, such as rescue catastrophes, the agent is required to explore as efficiently as…

机器人学 · 计算机科学 2023-08-01 Xuyang Chen , Ashvin N. Iyer , Zixing Wang , Ahmed H. Qureshi

Predicting future behaviors of road agents is a key task in autonomous driving. While existing models have demonstrated great success in predicting marginal agent future behaviors, it remains a challenge to efficiently predict consistent…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Xin Huang , Xiaoyu Tian , Junru Gu , Qiao Sun , Hang Zhao

While Vision-language models (VLMs) have demonstrated remarkable performance across multi-modal tasks, their choice of vision encoders presents a fundamental weakness: their low-level features lack the robust structural and spatial…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Brandon Huang , Hang Hua , Zhuoran Yu , Trevor Darrell , Rogerio Feris , Roei Herzig

Accurately modeling agent behaviors is an important task in self-driving. It is also a task with many symmetries, such as equivariance to the order of agents and objects in the scene or equivariance to arbitrary roto-translations of the…

机器人学 · 计算机科学 2026-04-03 Scott Xu , Dian Chen , Kelvin Wong , Chris Zhang , Kion Fallah , Raquel Urtasun

Program code serves as a bridge linking vision and logic, providing a feasible supervisory approach for enhancing the multimodal reasoning capability of large models through geometric operations such as auxiliary line construction and…

人工智能 · 计算机科学 2026-02-10 Zhenyu Wu , Yanxi Long , Jian Li , Hua Huang

Most prior motion prediction endeavors in autonomous driving have inadequately encoded future scenarios, leading to predictions that may fail to accurately capture the diverse movements of agents (e.g., vehicles or pedestrians). To address…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Mingkun Wang , Xiaoguang Ren , Ruochun Jin , Minglong Li , Xiaochuan Zhang , Changqian Yu , Mingxu Wang , Wenjing Yang

While Vision-Language Models (VLMs) show significant promise for end-to-end autonomous driving by leveraging the common sense embedded in language models, their reliance on 2D image cues for complex scene understanding and decision-making…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Weijie Wei , Zhipeng Luo , Ling Feng , Venice Erin Liong

Long-term visual localization in outdoor environment is a challenging problem, especially faced with the cross-seasonal, bi-directional tasks and changing environment. In this paper we propose a novel visual inertial localization framework…

机器人学 · 计算机科学 2018-03-06 Xiaqing Ding , Yue Wang , Dongxuan Li , Li Tang , Huan Yin , Rong Xiong

Data driven approaches for decision making applied to automated driving require appropriate generalization strategies, to ensure applicability to the world's variability. Current approaches either do not generalize well beyond the training…

机器学习 · 计算机科学 2022-03-11 Karl Kurzer , Philip Schörner , Alexander Albers , Hauke Thomsen , Karam Daaboul , J. Marius Zöllner

Multimodal language models (MLMs) integrate visual and textual information by coupling a vision encoder with a large language model through the specific adapter. While existing approaches commonly rely on a single pre-trained vision…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Matvey Skripkin , Elizaveta Goncharova , Dmitrii Tarasov , Andrey Kuznetsov

Forecasting the long-term future motion of road actors is a core challenge to the deployment of safe autonomous vehicles (AVs). Viable solutions must account for both the static geometric context, such as road lanes, and dynamic social…

机器学习 · 计算机科学 2020-08-25 Siddhesh Khandelwal , William Qi , Jagjeet Singh , Andrew Hartnett , Deva Ramanan

Urban forecasting has increasingly benefited from high-dimensional spatial data through two primary approaches: graph-based methods that rely on predefined spatial structures, and region-based methods that focus on learning expressive urban…

人工智能 · 计算机科学 2025-06-18 Yuhao Jia , Zile Wu , Shengao Yi , Yifei Sun , Xiao Huang