中文
相关论文

相关论文: Where Are You? Localization from Embodied Dialog

200 篇论文

The use of wireless signals for purposes of localization enables a host of applications relating to the determination and verification of the positions of network participants, ranging from radar to satellite navigation. Consequently, it…

网络与互联网体系结构 · 计算机科学 2020-12-02 Matthias Schäfer , Martin Strohmeier , Mauro Leonardi , Vincent Lenders

Lifelong user behavior sequences are crucial for capturing user interests and predicting user responses in modern recommendation systems. A two-stage paradigm is typically adopted to handle these long sequences: a subset of relevant…

信息检索 · 计算机科学 2025-03-27 Ningya Feng , Junwei Pan , Jialong Wu , Baixu Chen , Ximei Wang , Qian Li , Xian Hu , Jie Jiang , Mingsheng Long

Visual localization is the problem of estimating the position and orientation from which a given image (or a sequence of images) is taken in a known scene. It is an important part of a wide range of computer vision and robotics…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Ara Jafarzadeh , Manuel Lopez Antequera , Pau Gargallo , Yubin Kuang , Carl Toft , Fredrik Kahl , Torsten Sattler

Human-robot collaboration is an essential research topic in artificial intelligence (AI), enabling researchers to devise cognitive AI systems and affords an intuitive means for users to interact with the robot. Of note, communication plays…

人工智能 · 计算机科学 2021-08-09 Qi Wu , Cheng-Ju Wu , Yixin Zhu , Jungseock Joo

While multimodal conversation agents are gaining importance in several domains such as retail, travel etc., deep learning research in this area has been limited primarily due to the lack of availability of large-scale, open chatlogs. To…

计算与语言 · 计算机科学 2018-02-01 Amrita Saha , Mitesh Khapra , Karthik Sankaranarayanan

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, thus grounding each…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Paul Voigtlaender , Soravit Changpinyo , Jordi Pont-Tuset , Radu Soricut , Vittorio Ferrari

In multilingual societies, social conversations often involve code-mixed speech. The current speech technology may not be well equipped to extract information from multi-lingual multi-speaker conversations. The DISPLACE challenge entails a…

The ability of robots to estimate their location is crucial for a wide variety of autonomous operations. In settings where GPS is unavailable, measurements of transmissions from fixed beacons provide an effective means of estimating a…

机器人学 · 计算机科学 2017-09-21 Charles Schaff , David Yunis , Ayan Chakrabarti , Matthew R. Walter

Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to explain success or failure in closed-loop embodied tasks. Following an empirical AI…

人工智能 · 计算机科学 2026-05-20 Oussama Zenkri , Oliver Brock

Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitly represent the observer viewpoint, making spatial relations…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Ayaka Nanri , Klara Reichard , Mert Kiray , Federico Tombari , Benjamin Busam , Asako Kanezaki

Humans learn from life events to form intuitions towards the understanding of visual environments and languages. Envision that you are instructed by a high-level instruction, "Go to the bathroom in the master bedroom and replace the blue…

计算机视觉与模式识别 · 计算机科学 2021-03-25 Xiangru Lin , Guanbin Li , Yizhou Yu

Image geolocalization, in which an AI model traditionally predicts the precise GPS coordinates of an image, is a challenging task with many downstream applications. However, the user cannot utilize the model to further their knowledge…

计算机视觉与模式识别 · 计算机科学 2025-09-04 Ron Campos , Ashmal Vayani , Parth Parag Kulkarni , Rohit Gupta , Aizan Zafar , Aritra Dutta , Mubarak Shah

The embedded sensors in widely used smartphones and other wearable devices make the data of human activities more accessible. However, recognizing different human activities from the wearable sensor data remains a challenging research…

机器学习 · 计算机科学 2023-07-25 Taoran Sheng , Manfred Huber

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Ruohan Gao , Changan Chen , Ziad Al-Halah , Carl Schissler , Kristen Grauman

In this paper we present a system that detects and tracks objects and agents, computes spatial relations, and communicates those relations to the user using speech. Our system is able to detect multiple objects and agents at 30 frames per…

机器人学 · 计算机科学 2019-09-16 E. Akin Sisbot , Jonathan H. Connell

We are witnessing significant progress on perception models, specifically those trained on large-scale internet images. However, efficiently generalizing these perception models to unseen embodied tasks is insufficiently studied, which will…

机器人学 · 计算机科学 2023-03-21 Ya Jing , Tao Kong

Developing autonomous home robots controlled by natural language has long been a pursuit of humanity. While advancements in large language models (LLMs) and embodied intelligence make this goal closer, several challenges persist: the lack…

机器人学 · 计算机科学 2025-05-16 Dongping Li , Tielong Cai , Tianci Tang , Wenhao Chai , Katherine Rose Driggs-Campbell , Gaoang Wang

Accurate prediction of future person location and movement trajectory from an egocentric wearable camera can benefit a wide range of applications, such as assisting visually impaired people in navigation, and the development of mobility…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Jianing Qiu , Frank P. -W. Lo , Xiao Gu , Yingnan Sun , Shuo Jiang , Benny Lo

Background: Verbal deception detection research relies on narratives and commonly assumes statements as truthful or deceptive. A more realistic perspective acknowledges that the veracity of statements exists on a continuum with truthful and…

计算与语言 · 计算机科学 2025-07-24 Riccardo Loconte , Bennett Kleinberg

We introduce GuessWhat?!, a two-player guessing game as a testbed for research on the interplay of computer vision and dialogue systems. The goal of the game is to locate an unknown object in a rich image scene by asking a sequence of…

人工智能 · 计算机科学 2017-02-08 Harm de Vries , Florian Strub , Sarath Chandar , Olivier Pietquin , Hugo Larochelle , Aaron Courville
‹ 上一页 1 8 9 10 下一页 ›