中文
相关论文

相关论文: OmniScene: Attention-Augmented Multimodal 4D Scene…

200 篇论文

Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Saurav Jha , Stefan K. Ehrlich

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are…

计算机视觉与模式识别 · 计算机科学 2023-04-07 Songyou Peng , Kyle Genova , Chiyu "Max" Jiang , Andrea Tagliasacchi , Marc Pollefeys , Thomas Funkhouser

Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However, objects in the BEV representation typically exhibit small sizes, and the associated point cloud…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Junbo Yin , Jianbing Shen , Runnan Chen , Wei Li , Ruigang Yang , Pascal Frossard , Wenguan Wang

Autonomous driving on water surfaces plays an essential role in executing hazardous and time-consuming missions, such as maritime surveillance, survivors rescue, environmental monitoring, hydrography mapping and waste cleaning. This work…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Shanliang Yao , Runwei Guan , Zhaodong Wu , Yi Ni , Zile Huang , Ryan Wen Liu , Yong Yue , Weiping Ding , Eng Gee Lim , Hyungjoon Seo , Ka Lok Man , Jieming Ma , Xiaohui Zhu , Yutao Yue

Accurate classification of autonomous vehicle (AV) driving behaviors is critical for safety validation, performance diagnosis, and traffic integration analysis. However, existing approaches primarily rely on numerical time-series modeling…

Large vision-language models (VLMs) are increasingly used in autonomous-vehicle (AV) stacks, but hallucination limits their reliability in safety-critical pipelines. We present Shapley-credited Context-Aware Dawid-Skene with Agreement, a…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yuxiang Feng , Keyang Zhang , Hassane Ouchouid , Ashwil Kaniamparambil , Ioannis Souflas , Panagiotis Angeloudis

To maximize safety and driving comfort, autonomous driving systems can benefit from implementing foresighted action choices that take different potential scenario developments into account. While artificial scene prediction methods are…

机器人学 · 计算机科学 2022-04-15 Chao Wang , Thomas H. Weisswange , Matti Krueger , Christiane B. Wiebel-Herboth

Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems. In the context of intelligent vehicles, leveraging the power of…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Sheng Luo , Wei Chen , Wanxin Tian , Rui Liu , Luanxuan Hou , Xiubao Zhang , Haifeng Shen , Ruiqi Wu , Shuyi Geng , Yi Zhou , Ling Shao , Yi Yang , Bojun Gao , Qun Li , Guobin Wu

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Thomas Monninger , Shaoyuan Xie , Qi Alfred Chen , Sihao Ding

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Zirui Wang , Tao Zhang

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This…

Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for…

机器人学 · 计算机科学 2026-05-22 Wenxuan Guo , Xiuwei Xu , Yichen Liu , Xiangyu Li , Hang Yin , Huangxing Chen , Wenzhao Zheng , Jianjiang Feng , Jie Zhou , Jiwen Lu

Over the past few years, the advancement of Multimodal Large Language Models (MLLMs) has captured the wide interest of researchers, leading to numerous innovations to enhance MLLMs' comprehension. In this paper, we present AdaptVision, a…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Yonghui Wang , Wengang Zhou , Hao Feng , Houqiang Li

The integration of electric vehicles (EVs) into smart grids presents unique opportunities to enhance both transportation systems and energy networks. However, ensuring safe and interpretable interactions between drivers, vehicles, and the…

Multi-sensor fusion plays a critical role in enhancing perception for autonomous driving, overcoming individual sensor limitations, and enabling comprehensive environmental understanding. This paper first formalizes multi-sensor fusion…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Chuheng Wei , Ziye Qin , Ziyan Zhang , Guoyuan Wu , Matthew J. Barth

Taking over arbitrary tasks like humans do with a mobile service robot in open-world settings requires a holistic scene perception for decision-making and high-level control. This paper presents a human-inspired scene perception model to…

机器人学 · 计算机科学 2024-07-09 Florenz Graf , Jochen Lindermayr , Birgit Graf , Werner Kraus , Marco F. Huber

We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Xinhao Cai , Minghang Zheng , Xin Jin , Yang Liu

Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various…

计算机视觉与模式识别 · 计算机科学 2023-10-24 Zhecan Wang , Haoxuan You , Yicheng He , Wenhao Li , Kai-Wei Chang , Shih-Fu Chang

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Ziang Guo , Feng Yang , Xuefeng Zhang , Jiaqi Guo , Kun Zhao , Yixiao Zhou , Peng Lu , Sifa Zheng , Zufeng Zhang

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Fan-Yun Sun , Weiyu Liu , Siyi Gu , Dylan Lim , Goutam Bhat , Federico Tombari , Manling Li , Nick Haber , Jiajun Wu