English
Related papers

Related papers: Anticipating the Unseen Discrepancy for Vision and…

200 papers

It is expensive to collect training data for every possible domain that a vision model may encounter when deployed. We instead consider how simply verbalizing the training domain (e.g. "photos of birds") as well as domains we want to extend…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Lisa Dunlap , Clara Mohri , Devin Guillory , Han Zhang , Trevor Darrell , Joseph E. Gonzalez , Aditi Raghunathan , Anja Rohrbach

Contrastive self-supervised learning has emerged as a promising approach to unsupervised visual representation learning. In general, these methods learn global (image-level) representations that are invariant to different views (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Pedro O. Pinheiro , Amjad Almahairi , Ryan Y. Benmalek , Florian Golemo , Aaron Courville

We explore the use of language as a perceptual representation for vision-and-language navigation (VLN), with a focus on low-data settings. Our approach uses off-the-shelf vision systems for image captioning and object detection to convert…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Bowen Pan , Rameswar Panda , SouYoung Jin , Rogerio Feris , Aude Oliva , Phillip Isola , Yoon Kim

The existing methods for Vision and Language Navigation in the Continuous Environment (VLN-CE) commonly incorporate a waypoint predictor to discretize the environment. This simplifies the navigation actions into a view selection task and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Yue Zhang , Parisa Kordjamshidi

Vision-and-Language Navigation (VLN) tasks such as Room-to-Room (R2R) require machine agents to interpret natural language instructions and learn to act in visually realistic environments to achieve navigation goals. The overall task…

Computer Vision and Pattern Recognition · Computer Science 2019-08-14 Haoshuo Huang , Vihan Jain , Harsh Mehta , Alexander Ku , Gabriel Magalhaes , Jason Baldridge , Eugene Ie

Learned communication makes multi-agent systems more effective by aggregating distributed information. However, it also exposes individual agents to the threat of erroneous messages they might receive. In this paper, we study the setting…

Computer Vision and Pattern Recognition · Computer Science 2020-11-11 Nicholas Vadivelu , Mengye Ren , James Tu , Jingkang Wang , Raquel Urtasun

Efficient ObjectGoal navigation (ObjectNav) in novel environments requires an understanding of the spatial and semantic regularities in environment layouts. In this work, we present a straightforward method for learning these regularities…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Albert J. Zhai , Shenlong Wang

The challenge of navigation in environments with dynamic objects continues to be a central issue in the study of autonomous agents. While predictive methods hold promise, their reliance on precise state information makes them less practical…

Robotics · Computer Science 2024-10-28 Hsuan-Kung Yang , Tsung-Chih Chiang , Ting-Ru Liu , Chun-Wei Huang , Jou-Min Liu , Chun-Yi Lee

Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for diverse applications. However, the scarcity of training data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Sen Wang , Dongliang Zhou , Liang Xie , Chao Xu , Ye Yan , Erwei Yin

Visual perception is an important component for autonomous navigation of unmanned surface vessels (USV), particularly for the tasks related to autonomous inspection and tracking. These tasks involve vision-based navigation techniques to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Muhayyuddin Ahmed , Ahsan Baidar Bakht , Taimur Hassan , Waseem Akram , Ahmed Humais , Lakmal Seneviratne , Shaoming He , Defu Lin , Irfan Hussain

We present the Frontier Aware Search with backTracking (FAST) Navigator, a general framework for action decoding, that achieves state-of-the-art results on the Room-to-Room (R2R) Vision-and-Language navigation challenge of Anderson et. al.…

Computation and Language · Computer Science 2019-04-03 Liyiming Ke , Xiujun Li , Yonatan Bisk , Ari Holtzman , Zhe Gan , Jingjing Liu , Jianfeng Gao , Yejin Choi , Siddhartha Srinivasa

Despite the recent progress in deep learning based computer vision, domain shifts are still one of the major challenges. Semantic segmentation for autonomous driving faces a wide range of domain shifts, e.g. caused by changing weather…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Manuel Schwonberg , Claus Werner , Hanno Gottschalk , Carsten Meyer

Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying on visual evidence. While prior work attempts to mitigate this issue through decoding…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Seulbi Lee , Sangheum Hwang

The academic field of learning instruction-guided visual navigation can be generally categorized into high-level category-specific search and low-level language-guided navigation, depending on the granularity of language instruction, in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Gengze Zhou , Yicong Hong , Zun Wang , Chongyang Zhao , Mohit Bansal , Qi Wu

Humans have a natural ability to perform semantic associations with the surrounding objects in the environment. This allows them to create a mental map of the environment, allowing them to navigate on-demand when given linguistic…

Vision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN agents are trained typically under disturbance-free…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Bingqian Lin , Yanxin Long , Yi Zhu , Fengda Zhu , Xiaodan Liang , Qixiang Ye , Liang Lin

Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict high-dimensional visual states can distribute model capacity…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yi Yang , Xueqi Li , Yiyang Chen , Jin Song , Yihan Wang , Zipeng Xiao , Jiadi Su , You Qiaoben , Pengfei Liu , Zhijie Deng

Vision-language navigation (VLN) is a challenging task due to its large searching space in the environment. To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2022-03-09 Xiwen Liang , Fengda Zhu , Lingling Li , Hang Xu , Xiaodan Liang

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ipsita Praharaj , Yukta Butala , Badrikanath Praharaj , Yash Butala

The ability of gaze estimation models to generalize is often significantly hindered by various factors unrelated to gaze, especially when the training dataset is limited. Current strategies aim to address this challenge through different…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Pengwei Yin , Jingjing Wang , Guanzhong Zeng , Di Xie , Jiang Zhu
‹ Prev 1 4 5 6 7 8 10 Next ›