中文
相关论文

相关论文: Less is More: Generating Grounded Navigation Instr…

200 篇论文

Vision Language Models (VLMs) exhibit persistent hallucinations in counting tasks, with accuracy substantially lower than other visual reasoning tasks (excluding sentiment). This phenomenon persists even in state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Boyuan Chen , Minghao Shao , Siddharth Garg , Ramesh Karri , Muhammad Shafique

Navigation instruction generation for visually impaired (VI) individuals (NIG-VI) is critical yet relatively underexplored. This study focuses on generating precise, in-situ, step-by-step navigation instructions that are practically usable…

计算与语言 · 计算机科学 2025-12-19 Yi Zhao , Siqi Wang , Jing Li

GUI grounding, the task of mapping natural-language instructions to pixel coordinates, is crucial for autonomous agents, yet remains difficult for current VLMs. The core bottleneck is reliable patch-to-pixel mapping, which breaks when…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Suyuchen Wang , Tianyu Zhang , Ahmed Masry , Christopher Pal , Spandana Gella , Bang Liu , Perouz Taslakian

Manually generating catchy descriptions and names is labor intensive and a slow process for retailers. Although generative AI provides an automation solution in form of Vision to Language Models (VLM), the current VLMs are prone to factual…

机器学习 · 计算机科学 2025-10-28 Nayan Kumar Singh

The Reference Remote Sensing Image Segmentation (RRSIS) task generates segmentation masks for specified objects in images based on textual descriptions, which has attracted widespread attention and research interest. Current RRSIS methods…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Shuyang Li , Shuang Wang , Zhuangzhuang Sun , Jing Xiao

We propose a learning system in which language is grounded in visual percepts without specific pre-defined categories of terms. We present a unified generative method to acquire a shared semantic/visual embedding that enables the learning…

计算与语言 · 计算机科学 2021-08-02 Nisha Pillai , Cynthia Matuszek , Francis Ferraro

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Rafi Ibn Sultan , Hui Zhu , Xiangyu Zhou , Chengyin Li , Prashant Khanduri , Marco Brocanelli , Dongxiao Zhu

Vision-and-Language Navigation (VLN) requires an agent to find a path to a remote location on the basis of natural-language instructions and a set of photo-realistic panoramas. Most existing methods take the words in the instructions and…

计算与语言 · 计算机科学 2021-08-26 Yuankai Qi , Zizheng Pan , Yicong Hong , Ming-Hsuan Yang , Anton van den Hengel , Qi Wu

Robots operating in shared human environments must not only navigate, interact, and detect their surroundings, they must also interpret and respond to dynamic, and often unpredictable, human behaviours. Although recent advances have shown…

Multimodal large language models (MLLMs) hold promise for integrating diverse data modalities, but current medical adaptations such as LLaVA-Med often fail to fully exploit the synergy between color fundus photography (CFP) and optical…

This work studies object goal navigation task, which involves navigating to the closest object related to the given semantic category in unseen environments. Recent works have shown significant achievements both in the end-to-end…

人工智能 · 计算机科学 2021-09-21 Aleksey Staroverov , Aleksandr I. Panov

Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) face two critical challenges: the lack of cross-image reasoning…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Lihao Zheng , Jiawei Chen , Xintian Shen , Hao Ma , Tao Wei

Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations of sub-goals implied by the instructions can serve as…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Akhil Perincherry , Jacob Krantz , Stefan Lee

Multi-modality foundation models, as represented by GPT-4V, have brought a new paradigm for low-level visual perception and understanding tasks, that can respond to a broad range of natural human instructions in a model. While existing…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Haoning Wu , Zicheng Zhang , Erli Zhang , Chaofeng Chen , Liang Liao , Annan Wang , Kaixin Xu , Chunyi Li , Jingwen Hou , Guangtao Zhai , Geng Xue , Wenxiu Sun , Qiong Yan , Weisi Lin

Event cameras capture changes in brightness with microsecond precision and remain reliable under motion blur and challenging illumination, offering clear advantages for modeling highly dynamic scenes. Yet, their integration with natural…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Lingdong Kong , Dongyue Lu , Ao Liang , Rong Li , Yuhao Dong , Tianshuai Hu , Lai Xing Ng , Wei Tsang Ooi , Benoit R. Cottereau

The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with interesting findings:…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Wayner Barrios , Mattia Soldan , Alberto Mario Ceballos-Arroyo , Fabian Caba Heilbron , Bernard Ghanem

Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate on-screen coordinates for actions such as clicks and keystrokes. However, recent Vision Language…

Large language models (LLMs) have enabled the automatic generation of step-by-step augmented reality (AR) instructions for a wide range of physical tasks. However, existing LLM-based AR guidance often lacks rich visual augmentations to…

人机交互 · 计算机科学 2025-09-25 Ada Yi Zhao , Aditya Gunturu , Ellen Yi-Luen Do , Ryo Suzuki

This paper presents a new approach for integrating semantic information for vision-based vehicle navigation. Although vision-based vehicle navigation systems using pre-mapped visual landmarks are capable of achieving submeter level accuracy…

计算机视觉与模式识别 · 计算机科学 2018-01-04 Varun Murali , Han-Pang Chiu , Supun Samarasekera , Rakesh , Kumar

To mitigate potential risks associated with language models, recent AI detection research proposes incorporating watermarks into machine-generated text through random vocabulary restrictions and utilizing this information for detection.…

计算与语言 · 计算机科学 2024-02-14 Yu Fu , Deyi Xiong , Yue Dong