English
Related papers

Related papers: Situational Fusion of Visual Representation for Vi…

200 papers

Human texture perception is a weighted average of multi-sensory inputs: visual and tactile. While the visual sensing mechanism extracts global features, the tactile mechanism complements it by extracting local features. The lack of coupled…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Prasanna Kumar Routray , Aditya Sanjiv Kanade , Jay Bhanushali , Manivannan Muniyandi

Real world visual navigation requires robots to operate in unfamiliar, human-occupied dynamic environments. Navigation around humans is especially difficult because it requires anticipating their future motion, which can be quite…

Robotics · Computer Science 2021-02-16 Varun Tolani , Somil Bansal , Aleksandra Faust , Claire Tomlin

Image fusion aims to combine complementary information from multiple source images to generate more comprehensive scene representations. Existing methods primarily rely on the stacking and design of network architectures to enhance the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Linli Ma , Suzhen Lin , Jianchao Zeng , Zanxia Jin , Yanbo Wang , Fengyuan Li , Yubing Luo

The robustness of visual navigation policies trained through imitation often hinges on the augmentation of the training image-action pairs. Traditionally, this has been done by collecting data from multiple cameras, by using standard data…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Dhruv Sharma , Alihusein Kuwajerwala , Florian Shkurti

Objects, in the real world, rarely occur in isolation and exhibit typical arrangements governed by their independent utility, and their expected interaction with humans and other objects in the context. For example, a chair is expected near…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Sharat Agarwal

Appearance-based gaze estimation has been actively studied in recent years. However, its generalization performance for unseen head poses is still a significant limitation for existing methods. This work proposes a generalizable multi-view…

Computer Vision and Pattern Recognition · Computer Science 2023-11-16 Yoichiro Hisadome , Tianyi Wu , Jiawei Qin , Yusuke Sugano

Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Shaohong Wang , Bin Lu , Xinyu Xiao , Hanzhi Zhong , Bowen Pang , Tong Wang , Zhiyu Xiang , Hangguan Shan , Eryun Liu

To complete a complex task where a robot navigates to a goal object and fetches it, the robot needs to have a good understanding of the instructions and the surrounding environment. Large pre-trained models have shown capabilities to…

Robotics · Computer Science 2024-08-21 Yu Li , Dayou Li , Chenkun Zhao , Ruifeng Wang , Ran Song , Wei Zhang

Visual perception is an effective way to obtain the spatial characteristics of wireless channels and to reduce the overhead for communications system. A critical problem for the visual assistance is that the communications system needs to…

Signal Processing · Electrical Eng. & Systems 2024-12-17 Weihua Xu , Feifei Gao , Yong Zhang , Chengkang Pan , Guangyi Liu

Visual localization determines an agent's precise position and orientation within an environment using visual data. It has become a critical task in the field of robotics, particularly in applications such as autonomous navigation. This is…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Nanda Febri Istighfarin , HyungGi Jo

We propose a framework for aligning and fusing multiple images into a single view using neural image representations (NIRs), also known as implicit or coordinate-based neural representations. Our framework targets burst images that exhibit…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Seonghyeon Nam , Marcus A. Brubaker , Michael S. Brown

The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that attention mechanism is beneficial to the fusion process. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Bin Duan , Hao Tang , Wei Wang , Ziliang Zong , Guowei Yang , Yan Yan

The emerging vision-and-language navigation (VLN) problem aims at learning to navigate an agent to the target location in unseen photo-realistic environments according to the given language instruction. The main challenges of VLN arise…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Weixia Zhang , Chao Ma , Qi Wu , Xiaokang Yang

Vision-language navigation (VLN) requires an agent to navigate through an 3D environment based on visual observations and natural language instructions. It is clear that the pivotal factor for successful navigation lies in the comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Rui Liu , Wenguan Wang , Yi Yang

Vision-based Interfaces (VIs) are pivotal in advancing Human-Computer Interaction (HCI), particularly in enhancing context awareness. However, there are significant opportunities for these interfaces due to rapid advancements in multimodal…

Human-Computer Interaction · Computer Science 2024-08-15 Yongquan Hu , Wen Hu , Aaron Quigley

Visual sensation and perception refers to the process of sensing, organizing, identifying, and interpreting visual information in environmental awareness and understanding. Computational models inspired by visual perception have the…

Artificial Intelligence · Computer Science 2021-09-09 Bing Wei , Yudi Zhao , Kuangrong Hao , Lei Gao

Typical attempts to improve the capability of visual place recognition techniques include the use of multi-sensor fusion and integration of information over time from image sequences. These approaches can improve performance but have…

Robotics · Computer Science 2019-03-11 Stephen Hausler , Adam Jacobson , Michael Milford

An emerging paradigm in vision-and-language navigation (VLN) is the use of history-aware multi-modal transformer models. Given a language instruction, these models process observation and navigation history to predict the most appropriate…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Dongwoo Kang , Akhil Perincherry , Zachary Coalson , Aiden Gabriel , Stefan Lee , Sanghyun Hong

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Haolong Yan , Yeqing Shen , Xin Huang , Jia Wang , Kaijun Tan , Zhixuan Liang , Hongxin Li , Zheng Ge , Osamu Yoshie , Si Li , Xiangyu Zhang , Daxin Jiang

Humans can collaborate and complete tasks based on visual signals and instruction from the environment. Training such a robot is difficult especially due to the understanding of the instruction and the complicated environment. Previous…

Artificial Intelligence · Computer Science 2023-05-12 Kairui Zhou