中文
相关论文

相关论文: DiaLoc: An Iterative Approach to Embodied Dialog L…

200 篇论文

Deep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential…

计算机视觉与模式识别 · 计算机科学 2019-10-29 Bing Wang , Changhao Chen , Chris Xiaoxuan Lu , Peijun Zhao , Niki Trigoni , Andrew Markham

Existing visual localization methods are typically either 2D image-based, which are easy to build and maintain but limited in effective geometric reasoning, or 3D structure-based, which achieve high accuracy but require a centralized…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Xudong Jiang , Fangjinhua Wang , Silvano Galliani , Christoph Vogel , Marc Pollefeys

Precise object placement remains underexplored in aerial manipulation, where most systems rely on predefined target coordinates and focus primarily on grasping and control. Specifying exact placement poses, however, is cumbersome in…

机器人学 · 计算机科学 2026-03-10 Sarthak Mishra , Rishabh Dev Yadav , Naveen Nair , Wei Pan , Spandan Roy

Acquiring count annotations generally requires less human effort than point-level and bounding box annotations. Thus, we propose the novel problem setup of localizing objects in dense scenes under this weaker supervision. We propose LOOC, a…

计算机视觉与模式识别 · 计算机科学 2020-07-06 Issam H. Laradji , Rafael Pardinas , Pau Rodriguez , David Vazquez

Multimodal large language models (MLLMs), built on large-scale pre-trained vision towers and language models, have shown great capabilities in multimodal understanding. However, most existing MLLMs are trained on single-turn vision…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Jiazheng Liu , Sipeng Zheng , Börje F. Karlsson , Zongqing Lu

Pragmatic reasoning plays a pivotal role in deciphering implicit meanings that frequently arise in real-life conversations and is essential for the development of communicative social agents. In this paper, we introduce a novel challenge,…

计算与语言 · 计算机科学 2023-06-21 Hengli Li , Song-Chun Zhu , Zilong Zheng

Tactile sensing is a fundamental modality for embodied intelligence, offering unique and direct feedback on contact geometry, material properties, and interaction dynamics that remote sensors cannot replace. However, unimodal tactile…

Integrating visual and linguistic information into a single multimodal representation is an unsolved problem with wide-reaching applications to both natural language processing and computer vision. In this paper, we present a simple method…

机器学习 · 统计学 2017-03-28 Guillem Collell , Teddy Zhang , Marie-Francine Moens

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing…

机器人学 · 计算机科学 2026-05-14 Yiran Ling , Qing Lian , Jinghang Li , Qing Jiang , Tianming Zhang , Xiaoke Jiang , Chuanxiu Liu , Jie Liu , Lei Zhang

We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a real-life visual…

计算机视觉与模式识别 · 计算机科学 2020-05-19 Howard Chen , Alane Suhr , Dipendra Misra , Noah Snavely , Yoav Artzi

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional use of templates or pre-set information from training sets…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Yunhao Li , Xiaoqiong Liu , Luke Liu , Heng Fan , Libo Zhang

Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an…

计算与语言 · 计算机科学 2019-12-19 Feilong Chen , Fandong Meng , Jiaming Xu , Peng Li , Bo Xu , Jie Zhou

In this paper, we present a novel end-to-end learning-based LiDAR relocalization framework, termed PointLoc, which infers 6-DoF poses directly using only a single point cloud as input, without requiring a pre-built map. Compared to RGB…

机器人学 · 计算机科学 2021-11-23 Wei Wang , Bing Wang , Peijun Zhao , Changhao Chen , Ronald Clark , Bo Yang , Andrew Markham , Niki Trigoni

In this paper, we propose a probabilistic framework for solving the task of `Visual Dialog'. Solving this task requires reasoning and understanding of visual modality, language modality, and common sense knowledge to answer. Various…

计算机视觉与模式识别 · 计算机科学 2019-10-18 Badri N. Patro , Anupriy , Vinay P. Namboodiri

Predicting future sensory states is crucial for learning agents such as robots, drones, and autonomous vehicles. In this paper, we couple multiple sensory modalities with exploratory actions and propose a predictive neural network…

机器人学 · 计算机科学 2021-09-17 Xiaohui Chen , Ramtin Hosseini , Karen Panetta , Jivko Sinapov

Global localization is a critical problem in autonomous navigation, enabling precise positioning without reliance on GPS. Modern global localization techniques often depend on dense LiDAR maps, which, while precise, require extensive…

Humans use spatial language to naturally describe object locations and their relations. Interpreting spatial language not only adds a perceptual modality for robots, but also reduces the barrier of interfacing with humans. Previous work…

机器人学 · 计算机科学 2021-08-03 Kaiyu Zheng , Deniz Bayazit , Rebecca Mathew , Ellie Pavlick , Stefanie Tellex

Fingerprinting-based localization often suffers from poor cross-environment generalization, especially when only a few labeled samples are available in the target environment. Existing methods mitigate distribution shifts through domain…

信号处理 · 电气工程与系统科学 2026-05-20 Jun Gao , Zheng Xing , Wenliang Lin , Weibing Zhao , Xuhui Zhang , Junting Chen , Zhongliang Deng , Shuguang Cui

Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the traditional satellite-centric paradigm limits robustness when high-resolution or up-to-date…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Zixuan Song , Jing Zhang , Di Wang , Zidie Zhou , Wenbin Liu , Haonan Guo , En Wang , Bo Du

The field of multimodal robot navigation in indoor environments has garnered significant attention in recent years. However, as tasks and methods become more advanced, the action decision systems tend to become more complex and operate as…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Haru Kondoh , Asako Kanezaki