English
Related papers

Related papers: Transformer-based Localization from Embodied Dialo…

200 papers

In this work we propose an alternative formulation to the problem of ground reflectivity grid based localization involving laser scanned data from multiple LIDARs mounted on autonomous vehicles. The driving idea of our localization…

Robotics · Computer Science 2017-10-09 Juan Castorena , Siddharth Agarwal

Self-supervised vision-and-language pretraining (VLP) aims to learn transferable multi-modal representations from large-scale image-text data and to achieve strong performances on a broad scope of vision-language tasks after finetuning.…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Yongfei Liu , Chenfei Wu , Shao-yen Tseng , Vasudev Lal , Xuming He , Nan Duan

Over the last few years, we have witnessed tremendous progress on many subtasks of autonomous driving, including perception, motion forecasting, and motion planning. However, these systems often assume that the car is accurately localized…

Computer Vision and Pattern Recognition · Computer Science 2021-04-13 John Phillips , Julieta Martinez , Ioan Andrei Bârsan , Sergio Casas , Abbas Sadat , Raquel Urtasun

Motivated by the intuitive understanding humans have about the space of possible interactions, and the ease with which they can generalize this understanding to previously unseen scenes, we develop an approach for learning visual…

Robotics · Computer Science 2023-05-30 Homanga Bharadhwaj , Abhinav Gupta , Shubham Tulsiani

Structured scene representations are a core component of embodied agents, helping to consolidate raw sensory streams into readable, modular, and searchable formats. Due to their high computational overhead, many approaches build such…

Artificial Intelligence · Computer Science 2025-06-03 Muhammad Qasim Ali , Saeejith Nair , Alexander Wong , Yuchen Cui , Yuhao Chen

Methods for learning latent user representations from historical behavior logs have gained traction for recommendation tasks in e-commerce, content streaming, and other settings. However, this area still remains relatively underexplored in…

Accurate 6D object pose estimation is an important task for a variety of robotic applications such as grasping or localization. It is a challenging task due to object symmetries, clutter and occlusion, but it becomes more challenging when…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Thomas Jantos , Mohamed Amin Hamdad , Wolfgang Granig , Stephan Weiss , Jan Steinbrener

This paper investigates the performance of transformer-based architectures for person identification in natural, face-to-face conversation scenario. We implement and evaluate a two-stream framework that separately models spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Masoumeh Chapariniya , Teodora Vukovic , Sarah Ebling , Volker Dellwo

Predicting pedestrian motion trajectories is crucial for path planning and motion control of autonomous vehicles. Accurately forecasting crowd trajectories is challenging due to the uncertain nature of human motions in different…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Yu Liu , Yuexin Zhang , Kunming Li , Yongliang Qiao , Stewart Worrall , You-Fu Li , He Kong

Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, unified representation of the scene topology will allow…

Computer Vision and Pattern Recognition · Computer Science 2022-05-18 Zachary Seymour , Niluthpol Chowdhury Mithun , Han-Pang Chiu , Supun Samarasekera , Rakesh Kumar

This paper focuses on the DialFRED task, which is the task of embodied instruction following in a setting where an agent can actively ask questions about the task. To address this task, we propose DialMAT. DialMAT introduces Moment-based…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Kanta Kaneda , Ryosuke Korekata , Yuiga Wada , Shunya Nagashima , Motonari Kambara , Yui Iioka , Haruka Matsuo , Yuto Imai , Takayuki Nishimura , Komei Sugiura

Language-driven action localization in videos is a challenging task that involves not only visual-linguistic matching but also action boundary prediction. Recent progress has been achieved through aligning language query to video segments,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Shuo Yang , Xinxiao Wu

Graph neural networks (GNNs) largely rely on the message-passing paradigm, where nodes iteratively aggregate information from their neighbors. Yet, standard message passing neural networks (MPNNs) face well-documented theoretical and…

Machine Learning · Computer Science 2026-05-15 Juan Amboage , Ernst Röell , Patrick Schnider , Bastian Rieck

We propose to learn tasks directly from visual demonstrations by learning to predict the outcome of human and robot actions on an environment. We enable a robot to physically perform a human demonstrated task without knowledge of the…

Robotics · Computer Science 2017-03-09 Adam Tow , Niko Sünderhauf , Sareh Shirazi , Michael Milford , Jürgen Leitner

Meta-learning has emerged as an efficient approach for constructing target models based on support sets. For example, the meta-learned embeddings enable the construction of target nearest-neighbor classifiers for specific tasks by pulling…

Machine Learning · Computer Science 2023-09-19 Han-Jia Ye , Da-Wei Zhou , Lanqing Hong , Zhenguo Li , Xiu-Shen Wei , De-Chuan Zhan

Robots are traditionally bounded by a fixed embodiment during their operational lifetime, which limits their ability to adapt to their surroundings. Co-optimizing control and morphology of a robot, however, is often inefficient due to the…

Robotics · Computer Science 2022-12-20 Chen Yu , Weinan Zhang , Hang Lai , Zheng Tian , Laurent Kneip , Jun Wang

In a busy city street, a pedestrian surrounded by distractions can pick out a single sign if it is relevant to their route. Artificial agents in outdoor Vision-and-Language Navigation (VLN) are also confronted with detecting supervisory…

Machine Learning · Computer Science 2022-11-21 Jason Armitage , Leonardo Impett , Rico Sennrich

Learning from demonstration for motion planning is an ongoing research topic. In this paper we present a model that is able to learn the complex mapping from raw 2D-laser range findings and a target position to the required steering…

Robotics · Computer Science 2018-11-07 Mark Pfeiffer , Michael Schaeuble , Juan Nieto , Roland Siegwart , Cesar Cadena

An increasing amount of location-based service (LBS) data is being accumulated and helps to study urban dynamics and human mobility. GPS coordinates and other location indicators are normally low dimensional and only representing spatial…

Social and Information Networks · Computer Science 2022-10-11 Chenyu Tian , Yuchun Zhang , Zefeng Weng , Xiusen Gu , Wai Kin Victor Chan

We learn, in an unsupervised way, an embedding from sequences of radar images that is suitable for solving the place recognition problem with complex radar data. Our method is based on invariant instance feature learning but is tailored for…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Matthew Gadd , Daniele De Martini , Paul Newman
‹ Prev 1 8 9 10 Next ›