中文
相关论文

相关论文: VisionTrap: Vision-Augmented Trajectory Prediction…

200 篇论文

The advancement of autonomous driving technologies necessitates increasingly sophisticated methods for understanding and predicting real-world scenarios. Vision language models (VLMs) are emerging as revolutionary tools with significant…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Yongjie Fu , Anmol Jain , Xuan Di , Xu Chen , Zhaobin Mo

Latent Action Models (LAMs) have rapidly gained traction as an important component in the pre-training pipelines of leading Vision-Language-Action models. However, they fail when observations contain action-correlated distractors, often…

Large vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models' vulnerability to deliberately placed adversarial texts, such texts are often easily…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Yue Cao , Yun Xing , Jie Zhang , Di Lin , Tianwei Zhang , Ivor Tsang , Yang Liu , Qing Guo

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Nan Song , Bozhou Zhang , Xiatian Zhu , Jiankang Deng , Li Zhang

Recognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Md Zahid Hasan , Jiajing Chen , Jiyang Wang , Mohammed Shaiqur Rahman , Ameya Joshi , Senem Velipasalar , Chinmay Hegde , Anuj Sharma , Soumik Sarkar

In this paper, we propose a novel approach for agent motion prediction in cluttered environments. One of the main challenges in predicting agent motion is accounting for location and context-specific information. Our main contribution is…

机器人学 · 计算机科学 2020-07-08 Igor Gilitschenski , Guy Rosman , Arjun Gupta , Sertac Karaman , Daniela Rus

We present a multi-modal trajectory generation and selection algorithm for real-world mapless outdoor navigation in human-centered environments. Such environments contain rich features like crosswalks, grass, and curbs, which are easily…

机器人学 · 计算机科学 2025-05-19 Daeun Song , Jing Liang , Xuesu Xiao , Dinesh Manocha

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then…

机器人学 · 计算机科学 2025-09-30 Chaoran Zhu , Hengyi Wang , Yik Lung Pang , Changjae Oh

Pedestrian trajectory prediction is an essential task in robotic applications such as autonomous driving and robot navigation. State-of-the-art trajectory predictors use a conditional variational autoencoder (CVAE) with recurrent neural…

计算机视觉与模式识别 · 计算机科学 2020-11-17 Yu Yao , Ella Atkins , Matthew Johnson-Roberson , Ram Vasudevan , Xiaoxiao Du

Procedure planning requires a model to predict a sequence of actions that transform a start visual observation into a goal in instructional videos. While most existing methods rely primarily on visual observations as input, they often…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Lei Shi , Victor Aregbede , Andreas Persson , Martin Längkvist , Amy Loutfi , Stephanie Lowry

In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object detection. We extend…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Linfeng He , Yiming Sun , Sihao Wu , Jiaxu Liu , Xiaowei Huang

Object-aware reasoning in vision-language tasks poses significant challenges for current models, particularly in handling unseen objects, reducing hallucinations, and capturing fine-grained relationships in complex visual scenes. To address…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Antonio Carlos Rivera , Anthony Moore , Steven Robinson

Visual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Yiming Sun , Fan Yu , Shaoxiang Chen , Yu Zhang , Junwei Huang , Chenhui Li , Yang Li , Changbo Wang

Accurately predicting human behaviors is crucial for mobile robots operating in human-populated environments. While prior research primarily focuses on predicting actions in single-human scenarios from an egocentric view, several robotic…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Utsav Panchal , Yuchen Liu , Luigi Palmieri , Ilche Georgievski , Marco Aiello

High-definition (HD) maps have played an integral role in the development of modern autonomous vehicle (AV) stacks, albeit with high associated labeling and maintenance costs. As a result, many recent works have proposed methods for…

机器人学 · 计算机科学 2024-03-26 Xunjiang Gu , Guanyu Song , Igor Gilitschenski , Marco Pavone , Boris Ivanovic

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Fan Yang , Zhiyang Chen , Yousong Zhu , Xin Li , Jinqiao Wang

Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves…

计算机视觉与模式识别 · 计算机科学 2022-05-02 Sibo Song , Jianqiang Wan , Zhibo Yang , Jun Tang , Wenqing Cheng , Xiang Bai , Cong Yao

This paper introduces VLAP, a novel approach that bridges pretrained vision models and large language models (LLMs) to make frozen LLMs understand the visual world. VLAP transforms the embedding space of pretrained vision models into the…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Jungin Park , Jiyoung Lee , Kwanghoon Sohn

This paper proposes a new method, that we call VisualBackProp, for visualizing which sets of pixels of the input image contribute most to the predictions made by the convolutional neural network (CNN). The method heavily hinges on exploring…

计算机视觉与模式识别 · 计算机科学 2017-05-23 Mariusz Bojarski , Anna Choromanska , Krzysztof Choromanski , Bernhard Firner , Larry Jackel , Urs Muller , Karol Zieba

Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent goal 6D pose of a target object in a 3D scene. However, we…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Sangwon Baik , Gunhee Kim , Mingi Choi , Hanbyul Joo