English
Related papers

Related papers: SUGAMAN: Describing Floor Plans for Visually Impai…

200 papers

Multi-task indoor scene understanding is widely considered as an intriguing formulation, as the affinity of different tasks may lead to improved performance. In this paper, we tackle the new problem of joint semantic, affordance and…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Xiaoxue Chen , Tianyu Liu , Hao Zhao , Guyue Zhou , Ya-Qin Zhang

This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Models (LLMs) to…

Robotics · Computer Science 2025-03-20 Yao Cheng , Zhe Han , Fengyang Jiang , Huaizhen Wang , Fengyu Zhou , Qingshan Yin , Lei Wei

Recently, there have been efforts to improve the performance in sign language recognition by designing self-supervised learning methods. However, these methods capture limited information from sign pose data in a frame-wise learning manner,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Weichao Zhao , Wengang Zhou , Hezhen Hu , Min Wang , Houqiang Li

Simultaneous mapping and localization (SLAM) in an real indoor environment is still a challenging task. Traditional SLAM approaches rely heavily on low-level geometric constraints like corners or lines, which may lead to tracking failure in…

Robotics · Computer Science 2019-10-01 Xueyang Kang , Shunying Yuan

This paper describes a system whereby a robot detects and track human-meaningful navigational cues as it navigates in an indoor environment. It is intended as the sensor front-end for a mobile robot system that can communicate its…

Robotics · Computer Science 2019-03-12 Payam Nikdel , Richard Vaughan

Vision language models (VLMs) can simultaneously reason about images and texts to tackle many tasks, from visual question answering to image captioning. This paper focuses on map parsing, a novel task that is unexplored within the VLM…

Robotics · Computer Science 2025-11-26 David DeFazio , Hrudayangam Mehta , Meng Wang , Ping Yang , Jeremy Blackburn , Shiqi Zhang

Vision-Language Navigation (VLN) requires an embodied agent to navigate complex environments by following natural language instructions, which typically demands tight fusion of visual and language modalities. Existing VLN methods often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Daojie Peng , Fulong Ma , Jun Ma

Visual affordance learning is crucial for robots to understand and interact effectively with the physical world. Recent advances in this field attempt to leverage pre-trained knowledge of vision-language foundation models to learn…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Qian Zhang , Lin Zhang , Xing Fang , Mingxin Zhang , Zhiyuan Wei , Ran Song , Wei Zhang

This paper presents GoodSAM++, a novel framework utilizing the powerful zero-shot instance segmentation capability of SAM (i.e., teacher) to learn a compact panoramic semantic segmentation model, i.e., student, without requiring any labeled…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Weiming Zhang , Yexin Liu , Xu Zheng , Lin Wang

We consider an autonomous navigation robot that can accept human commands through natural language to provide services in an indoor environment. These natural language commands may include time, position, object, and action components.…

Robotics · Computer Science 2024-10-18 Kuan-Lin Chen , Tzu-Ti Wei , Li-Tzu Yeh , Elaine Kao , Yu-Chee Tseng , Jen-Jee Chen

Person search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are…

Computer Vision and Pattern Recognition · Computer Science 2021-09-28 Yichao Yan , Jinpeng Li , Shengcai Liao , Jie Qin , Bingbing Ni , Xiaokang Yang , Ling Shao

Recent models achieve promising results in visually grounded dialogues. However, existing datasets often contain undesirable biases and lack sophisticated linguistic analyses, which make it difficult to understand how well current models…

Computation and Language · Computer Science 2020-10-08 Takuma Udagawa , Takato Yamazaki , Akiko Aizawa

Exploring an unfamiliar indoor environment and avoiding obstacles is challenging for visually impaired people. Currently, several approaches achieve the avoidance of static obstacles based on the mapping of indoor scenes. To solve the issue…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Wenyan Ou , Jiaming Zhang , Kunyu Peng , Kailun Yang , Gerhard Jaworek , Karin Müller , Rainer Stiefelhagen

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However, existing VPG approaches are…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Chaolei Tan , Jianhuang Lai , Wei-Shi Zheng , Jian-Fang Hu

Pre-training has proven effective for learning transferable features in sign language understanding (SLU) tasks. Recently, skeleton-based methods have gained increasing attention because they can robustly handle variations in subjects and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Muxin Pu , Mei Kuan Lim , Chun Yong Chong , Chen Change Loy

Indoor built environments like homes and offices often present complex and cluttered layouts that pose significant challenges for individuals who are blind or visually impaired, especially when performing tasks that involve locating and…

Robotics · Computer Science 2025-09-03 Yifan Xu , Qianwei Wang , Vineet Kamat , Carol Menassa

Vision-and-Language Navigation (VLN) tasks require an agent to navigate through the environment based on language instructions. In this paper, we aim to solve two key challenges in this task: utilizing multilingual instructions for improved…

Computer Vision and Pattern Recognition · Computer Science 2022-07-06 Jialu Li , Hao Tan , Mohit Bansal

Image-text matching has received growing interest since it bridges vision and language. The key challenge lies in how to learn correspondence between image and text. Existing works learn coarse correspondence based on object co-occurrence…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Chunxiao Liu , Zhendong Mao , Tianzhu Zhang , Hongtao Xie , Bin Wang , Yongdong Zhang

In this paper, we present a novel siamese motion-aware network (SiamMan) for visual tracking, which consists of the siamese feature extraction subnetwork, followed by the classification, regression, and localization branches in parallel.…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Wenzhang Zhou , Longyin Wen , Libo Zhang , Dawei Du , Tiejian Luo , Yanjun Wu

We present a new method, PARsing And visual GrOuNding (ParaGon), for grounding natural language in object placement tasks. Natural language generally describes objects and spatial relations with compositionality and ambiguity, two major…

Robotics · Computer Science 2023-03-14 Zirui Zhao , Wee Sun Lee , David Hsu