中文
相关论文

相关论文: DVGT-2: Vision-Geometry-Action Model for Autonomou…

200 篇论文

Recent advancements in Visual Language Models (VLMs) have made them crucial for visual question answering (VQA) in autonomous driving, enabling natural human-vehicle interactions. However, existing methods often struggle in dynamic driving…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Siwen Jiao , Yangyi Fang , Baoyun Peng , Wangqun Chen , Bharadwaj Veeravalli

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera.…

机器人学 · 计算机科学 2026-04-24 Songen Gu , Yuhang Zheng , Weize Li , Yupeng Zheng , Yating Feng , Xiang Li , Yilun Chen , Pengfei Li , Wenchao Ding

Effectively capturing intricate interactions among road users is of critical importance to achieving safe navigation for autonomous vehicles. While graph learning (GL) has emerged as a promising approach to tackle this challenge, existing…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Junyao Wang , Arnav Vaibhav Malawade , Junhong Zhou , Shih-Yuan Yu , Mohammad Abdullah Al Faruque

Vision-Language-Action (VLA) models have recently attracted growing attention in end-to-end autonomous driving for their strong reasoning capabilities and rich world knowledge. However, existing VLAs often suffer from limited numerical…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Zhaohui Wang , Tengbo Yu , Hao Tang

Recent vision-language-action (VLA) systems have demonstrated strong capabilities in embodied manipulation. However, most existing VLA policies rely on limited observation windows and end-to-end action prediction, which makes them brittle…

机器人学 · 计算机科学 2026-04-16 Zhen Liu , Xinyu Ning , Zhe Hu , Xinxin Xie , Weize Li , Zhipeng Tang , Chongyu Wang , Zejun Yang , Hanlin Wang , Yitong Liu , Zhongzhu Pu

Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision-Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge. Existing 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Haijier Chen , Bo Xu , Shoujian Zhang , Haoze Liu , Jiaxuan Lin , Jingrong Wang

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

We introduce WAM-Flow, a vision-language-action (VLA) model that casts ego-trajectory planning as discrete flow matching over a structured token space. In contrast to autoregressive decoders, WAM-Flow performs fully parallel, bidirectional…

机器人学 · 计算机科学 2025-12-17 Yifang Xu , Jiahao Cui , Feipeng Cai , Zhihao Zhu , Hanlin Shang , Shan Luan , Mingwang Xu , Neng Zhang , Yaoyi Li , Jia Cai , Siyu Zhu

The UAV-VLA (Visual-Language-Action) system is a tool designed to facilitate communication with aerial robots. By integrating satellite imagery processing with the Visual Language Model (VLM) and the powerful capabilities of GPT, UAV-VLA…

This paper proposes a Video Graph Transformer (VGT) model for Video Quetion Answering (VideoQA). VGT's uniqueness are two-fold: 1) it designs a dynamic graph transformer module which encodes video by explicitly capturing the visual objects,…

计算机视觉与模式识别 · 计算机科学 2022-07-22 Junbin Xiao , Pan Zhou , Tat-Seng Chua , Shuicheng Yan

Current state-of-the-art autonomous vehicles could face safety-critical situations when their local sensors are occluded by large nearby objects on the road. Vehicle-to-vehicle (V2V) cooperative autonomous driving has been proposed as a…

机器人学 · 计算机科学 2026-02-17 Hsu-kuang Chiu , Ryo Hachiuma , Chien-Yi Wang , Yu-Chiang Frank Wang , Min-Hung Chen , Stephen F. Smith

Spatial reasoning poses a particular challenge for intelligent agents and is at the same time a prerequisite for their successful interaction and communication in the physical world. One such reasoning task is to describe the position of a…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Kyra Ahrens , Matthias Kerzel , Jae Hee Lee , Cornelius Weber , Stefan Wermter

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most existing VLAs rely primarily on 2D visual representations,…

机器人学 · 计算机科学 2026-05-21 Shizhe Chen , Paul Pacaud , Cordelia Schmid

Single-vehicle Vision-Language Models (VLMs) are fundamentally constrained by sensor occlusions. While Vehicle-to-Everything (V2X) systems mitigate this, current benchmarks lack the cooperative reasoning required for resolving ambiguities…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Kevin Richard , Alphin Varghese , Colin Pham , David Oh , Srijan Das

Vision-language models (VLMs) have become a promising approach to enhancing perception and decision-making in autonomous driving. The gap remains in applying VLMs to understand complex scenarios interacting with pedestrians and efficient…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Haoxiang Gao , Li Zhang , Yu Zhao , Zhou Yang , Jinghan Cao

Online dense mapping of urban scenes forms a fundamental cornerstone for scene understanding and navigation of autonomous vehicles. Recent advancements in mapping methods are mainly based on NeRF, whose rendering speed is too slow to meet…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Ke Wu , Kaizhao Zhang , Zhiwei Zhang , Shanshuai Yuan , Muer Tie , Julong Wei , Zijun Xu , Jieru Zhao , Zhongxue Gan , Wenchao Ding

We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural…

机器人学 · 计算机科学 2025-11-19 Ishika Singh , Ankit Goyal , Stan Birchfield , Dieter Fox , Animesh Garg , Valts Blukis

Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While the resulting geometric VLAs often show improved…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yurou Yang , Muyuan Lin , Roberto Martin-Martin , Martin Labrie , Shreekant Gayaka , Cheng-Hao Kuo , Luca Carlone

Recent advancements in 3D Gaussian Splatting (3D-GS) enable high-quality 3D scene reconstruction from RGB images. Many studies extend this paradigm for language-driven open-vocabulary scene understanding. However, most of them simply…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Jiazhong Cen , Xudong Zhou , Jiemin Fang , Changsong Wen , Lingxi Xie , Xiaopeng Zhang , Wei Shen , Qi Tian

The autonomous driving community is increasingly focused on addressing the challenges posed by out-of-distribution (OOD) driving scenarios. A dominant research trend seeks to enhance end-to-end (E2E) driving systems by integrating…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Yingzi Ma , Yulong Cao , Wenhao Ding , Shuibai Zhang , Yan Wang , Boris Ivanovic , Ming Jiang , Marco Pavone , Chaowei Xiao