English
Related papers

Related papers: Position-aware Location Regression Network for Tem…

200 papers

Temporally language grounding in untrimmed videos is a newly-raised task in video understanding. Most of the existing methods suffer from inferior efficiency, lacking interpretability, and deviating from the human perception mechanism.…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Jie Wu , Guanbin Li , Si Liu , Liang Lin

Image-text matching tasks have recently attracted a lot of attention in the computer vision field. The key point of this cross-domain problem is how to accurately measure the similarity between the visual and the textual contents, which…

Computation and Language · Computer Science 2019-07-24 Yaxiong Wang , Hao Yang , Xueming Qian , Lin Ma , Jing Lu , Biao Li , Xin Fan

A robot's ability to understand or ground natural language instructions is fundamentally tied to its knowledge about the surrounding world. We present an approach to grounding natural language utterances in the context of factual…

Robotics · Computer Science 2018-11-19 Rohan Paul , Andrei Barbu , Sue Felshin , Boris Katz , Nicholas Roy

Given a video, video grounding aims to retrieve a temporal moment that semantically corresponds to a language query. In this work, we propose a Parallel Attention Network with Sequence matching (SeqPAN) to address the challenges in this…

Computation and Language · Computer Science 2023-04-26 Hao Zhang , Aixin Sun , Wei Jing , Liangli Zhen , Joey Tianyi Zhou , Rick Siow Mong Goh

Succinct representation of complex signals using coordinate-based neural representations (CNRs) has seen great progress, and several recent efforts focus on extending them for handling videos. Here, the main challenge is how to (a)…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Subin Kim , Sihyun Yu , Jaeho Lee , Jinwoo Shin

Most existing work that grounds natural language phrases in images starts with the assumption that the phrase in question is relevant to the image. In this paper we address a more realistic version of the natural language grounding task…

Computer Vision and Pattern Recognition · Computer Science 2020-10-14 Bryan A. Plummer , Kevin J. Shih , Yichen Li , Ke Xu , Svetlana Lazebnik , Stan Sclaroff , Kate Saenko

Exploring fine-grained relationship between entities(e.g. objects in image or words in sentence) has great contribution to understand multimedia content precisely. Previous attention mechanism employed in image-text matching either takes…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Yaxian Xia , Lun Huang , Wenmin Wang , Xiaoyong Wei , Wenmin Wang

In visual place recognition (VPR), filtering and sequence-based matching approaches can improve performance by integrating temporal information across image sequences, especially in challenging conditions. While these methods are commonly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Somayeh Hussaini , Tobias Fischer , Michael Milford

Place recognition is a cornerstone of vehicle navigation and mapping, which is pivotal in enabling systems to determine whether a location has been previously visited. This capability is critical for tasks such as loop closure in…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Zhenyu Li , Tianyi Shang , Pengjie Xu , Zhaojun Deng

This paper presents a novel approach for the Vision-and-Language Navigation (VLN) task in continuous 3D environments, which requires an autonomous agent to follow natural language instructions in unseen environments. Existing end-to-end…

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yilun Chen , Shuai Yang , Haifeng Huang , Tai Wang , Runsen Xu , Ruiyuan Lyu , Dahua Lin , Jiangmiao Pang

To solve video-and-language grounding tasks, the key is for the network to understand the connection between the two modalities. For a pair of video and language description, their semantic relation is reflected by their encodings'…

Computer Vision and Pattern Recognition · Computer Science 2023-02-28 Yubo Zhang , Feiyang Niu , Qing Ping , Govind Thattai

Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlooking the inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Cai Chen , Runzhong Zhang , Jianjun Gao , Kejun Wu , Kim-Hui Yap , Yi Wang

Since the superiority of Transformer in learning long-term dependency, the sign language Transformer model achieves remarkable progress in Sign Language Recognition (SLR) and Translation (SLT). However, there are several issues with the…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Pan Xie , Mengyi Zhao , Xiaohui Hu

Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework…

Computation and Language · Computer Science 2025-10-21 Deyi Ji , Yuekui Yang , Haiyang Wu , Shaoping Ma , Tianrun Chen , Lanyun Zhu

Recurrent neural networks (RNNs) have shown the ability to improve scene parsing through capturing long-range dependencies among image units. In this paper, we propose dense RNNs for scene labeling by exploring various long-range semantic…

Computer Vision and Pattern Recognition · Computer Science 2018-11-13 Heng Fan , Peng Chu , Longin Jan Latecki , Haibin Ling

Recent advances in pixel-level tasks (e.g. segmentation) illustrate the benefit of of long-range interactions between aggregated region-based representations that can enhance local features. However, such aggregated representations, often…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Mir Rayat Imtiaz Hossain , Leonid Sigal , James J. Little

Mobile robots and autonomous vehicles are often required to function in environments where critical position estimates from sensors such as GPS become uncertain or unreliable. Single image visual place recognition (VPR) provides an…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Connor Malone , Ankit Vora , Thierry Peynot , Michael Milford

Recent years pretrained language models (PLMs) hit a success on several downstream tasks, showing their power on modeling language. To better understand and leverage what PLMs have learned, several techniques have emerged to explore…

Computation and Language · Computer Science 2021-09-23 Qian Liu , Dejian Yang , Jiahui Zhang , Jiaqi Guo , Bin Zhou , Jian-Guang Lou

Large language models (LLMs) exhibit a variety of promising capabilities in robotics, including long-horizon planning and commonsense reasoning. However, their performance in place recognition is still underexplored. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Zonglin Lyu , Juexiao Zhang , Mingxuan Lu , Yiming Li , Chen Feng