中文
相关论文

相关论文: SEER-VAR: Semantic Egocentric Environment Reasoner…

200 篇论文

An environment representation (ER) is a substantial part of every autonomous system. It introduces a common interface between perception and other system components, such as decision making, and allows downstream algorithms to deal with…

计算机视觉与模式识别 · 计算机科学 2019-07-30 Lukas Hoyer , Patrick Kesper , Anna Khoreva , Volker Fischer

Visual perception plays an important role in autonomous driving. One of the primary tasks is object detection and identification. Since the vision sensor is rich in color and texture information, it can quickly and accurately identify…

计算机视觉与模式识别 · 计算机科学 2022-12-23 Fei Liu , Zihao Lu , Xianke Lin

Augmented reality (AR) requires the seamless integration of visual, auditory, and linguistic channels for optimized human-computer interaction. While auditory and visual inputs facilitate real-time and contextual user guidance, the…

计算与语言 · 计算机科学 2023-10-19 Jing Bi , Nguyen Manh Nguyen , Ali Vosoughi , Chenliang Xu

Recently, autoregressive (AR) models have shown strong potential in image generation, offering better scalability and easier integration with unified multi-modal systems compared to diffusion-based methods. However, extending AR models to…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Dongyang Jin , Ryan Xu , Jianhao Zeng , Rui Lan , Yancheng Bai , Lei Sun , Xiangxiang Chu

In Augmented Reality (AR) environment, realistic interactions between the virtual and real objects play a crucial role in user experience. Much of recent advances in AR has been largely focused on developing geometry-aware environment, but…

计算机视觉与模式识别 · 计算机科学 2018-03-19 Long Chen , Karl Francis , Wen Tang

Ego-pose estimation and dynamic object tracking are two key issues in an autonomous driving system. Two assumptions are often made for them, i.e. the static world assumption of simultaneous localization and mapping (SLAM) and the exact…

机器人学 · 计算机科学 2022-02-24 Xuebo Tian , Junqiao Zhao , Chen Ye

CAR-Scenes is a frame-level dataset for autonomous driving that enables training and evaluation of vision-language models (VLMs) for interpretable, scene-level understanding. We annotate 5,192 images drawn from Argoverse 1, Cityscapes,…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Yuankai He , Weisong Shi

Maps have played an indispensable role in enabling safe and automated driving. Although there have been many advances on different fronts ranging from SLAM to semantics, building an actionable hierarchical semantic representation of urban…

机器人学 · 计算机科学 2024-09-12 Elias Greve , Martin Büchner , Niclas Vödisch , Wolfram Burgard , Abhinav Valada

Visual simultaneous localization and mapping (SLAM) plays a critical role in autonomous robotic systems, especially where accurate and reliable measurements are essential for navigation and sensing. In feature-based SLAM, the quantityand…

机器人学 · 计算机科学 2025-09-03 Haolan Zhang , Chenghao Li , Thanh Nguyen Canh , Lijun Wang , Nak Young Chong

Object Simultaneous Localization and Mapping (SLAM) systems struggle to correctly associate semantically similar objects in close proximity, especially in cluttered indoor environments and when scenes change. We present Semantic Enhancement…

机器人学 · 计算机科学 2025-06-18 Jungseok Hong , Ran Choi , John J. Leonard

Current techniques in Visual Simultaneous Localization and Mapping (VSLAM) estimate camera displacement by comparing image features of consecutive scenes. These algorithms depend on scene continuity, hence requires frequent camera inputs.…

机器人学 · 计算机科学 2024-01-25 Mingyang Li , Yue Ma , Qinru Qiu

Decreasing costs of vision sensors and advances in embedded hardware boosted lane related research detection, estimation, and tracking in the past two decades. The interest in this topic has increased even more with the demand for advanced…

计算机视觉与模式识别 · 计算机科学 2018-06-18 Rodrigo F. Berriel , Edilson de Aguiar , Alberto F. de Souza , Thiago Oliveira-Santos

Precise 6-DoF simultaneous localization and mapping (SLAM) from onboard sensors is critical for wearable devices capturing egocentric data, which exhibits specific challenges, such as a wider diversity of motions and viewpoints, prevalent…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Anusha Krishnan , Shaohui Liu , Paul-Edouard Sarlin , Oscar Gentilhomme , David Caruso , Maurizio Monge , Richard Newcombe , Jakob Engel , Marc Pollefeys

Monocular simultaneous localization and mapping (SLAM) is emerging in advanced driver assistance systems and autonomous driving, because a single camera is cheap and easy to install. Conventional monocular SLAM has two major challenges…

计算机视觉与模式识别 · 计算机科学 2022-12-16 Jinkyu Lee , Muhyun Back , Sung Soo Hwang , Il Yong Chun

Free-text promptable 3D medical image segmentation offers an intuitive and clinically flexible interaction paradigm. However, current methods are highly sensitive to linguistic variability: minor changes in phrasing can cause substantial…

图像与视频处理 · 电气工程与系统科学 2026-03-10 Tongrui Zhang , Chenhui Wang , Yongming Li , Zhihao Chen , Xufeng Zhan , Hongming Shan

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Accurate localization in challenging garage environments -- marked by poor lighting, sparse textures, repetitive structures, dynamic scenes, and the absence of GPS -- is crucial for automated valet parking (AVP) tasks. Addressing these…

机器人学 · 计算机科学 2024-07-02 Ye Li , Wenchao Yang , Dekun Lin , Qianlei Wang , Zhe Cui , Xiaolin Qin

Reliable and accurate localization and mapping are key components of most autonomous systems. Besides geometric information about the mapped environment, the semantics plays an important role to enable intelligent navigation behaviors. In…

机器人学 · 计算机科学 2021-05-25 Xieyuanli Chen , Andres Milioto , Emanuele Palazzolo , Philippe Giguère , Jens Behley , Cyrill Stachniss

Embodied scene understanding serves as the cornerstone for autonomous agents to perceive, interpret, and respond to open driving scenarios. Such understanding is typically founded upon Vision-Language Models (VLMs). Nevertheless, existing…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Yunsong Zhou , Linyan Huang , Qingwen Bu , Jia Zeng , Tianyu Li , Hang Qiu , Hongzi Zhu , Minyi Guo , Yu Qiao , Hongyang Li

Ensuring safe decision-making in autonomous vehicles remains a fundamental challenge despite rapid advances in end-to-end learning approaches. Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse…

机器人学 · 计算机科学 2026-03-20 Zilin Huang , Zihao Sheng , Zhengyang Wan , Yansong Qu , Junwei You , Sicong Jiang , Sikai Chen
‹ 上一页 1 2 3 10 下一页 ›