English
Related papers

Related papers: Real-Time Referring Expression Comprehension by Si…

200 papers

For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper,…

Robotics · Computer Science 2021-11-08 Junha Roh , Karthik Desingh , Ali Farhadi , Dieter Fox

Video Referring Expression Comprehension (REC) aims to localize a target object in video frames referred by the natural language expression. Recently, the Transformerbased methods have greatly boosted the performance limit. However, we…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Ji Jiang , Meng Cao , Tengtao Song , Yuexian Zou

Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering.…

Computation and Language · Computer Science 2025-11-11 Akshar Tumu , Varad Shinde , Parisa Kordjamshidi

Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target…

Computer Vision and Pattern Recognition · Computer Science 2021-04-15 Zhengyao Lv , Xiaoming Li , Xin Li , Fu Li , Tianwei Lin , Dongliang He , Wangmeng Zuo

In this paper, we introduce the task of visual grounding for remote sensing data (RSVG). RSVG aims to localize the referred objects in remote sensing (RS) images with the guidance of natural language. To retrieve rich information from RS…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Yang Zhan , Zhitong Xiong , Yuan Yuan

Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as visual question answering. However, it has not been widely…

Computer Vision and Pattern Recognition · Computer Science 2021-05-06 Wei Suo , Mengyang Sun , Peng Wang , Qi Wu

Referring image segmentation aims to segment the target object referred by a natural language expression. However, previous methods rely on the strong assumption that one sentence must describe one target in the image, which is often not…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Yutao Hu , Qixiong Wang , Wenqi Shao , Enze Xie , Zhenguo Li , Jungong Han , Ping Luo

This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment in an untrimmed video according to a given sentence query. Traditional TSG methods mainly follow…

Multimedia · Computer Science 2026-05-26 Xiang Fang , Daizong Liu , Pan Zhou , Zichuan Xu , Ruixuan Li

Scene Graph Generation (SGG) has achieved significant progress recently. However, most previous works rely heavily on fixed-size entity representations based on bounding box proposals, anchors, or learnable queries. As each representation's…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Hengyue Liu , Bir Bhanu

Referring expression grounding is an important and challenging task in computer vision. To avoid the laborious annotation in conventional referring grounding, unpaired referring grounding is introduced, where the training data only contains…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Hengcan Shi , Munawar Hayat , Jianfei Cai

Even without auxiliary images, single hyperspectral image super-resolution (SHSR) methods can be designed to improve the spatial resolution of hyperspectral images. However, failing to explore coherence thoroughly along bands and…

Image and Video Processing · Electrical Eng. & Systems 2025-06-25 Xufei Wang , Mingjian Zhang , Fei Ge , Jinchen Zhu , Wen Sha , Jifen Ren , Zhimeng Hou , Shouguo Zheng , ling Zheng , Shizhuang Weng

Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on high-level visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xiang Fang , Daizong Liu , Pan Zhou , Guoshun Nan

Recently, one-stage visual grounders attract high attention due to their comparable accuracy but significantly higher efficiency than two-stage grounders. However, inter-object relation modeling has not been well studied for one-stage…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yang Jiao , Zequn Jie , Jingjing Chen , Lin Ma , Yu-Gang Jiang

Vision-and-language models trained to match images with text can be combined with visual explanation methods to point to the locations of specific objects in an image. Our work shows that the localization --"grounding"-- abilities of these…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Ruozhen He , Paola Cascante-Bonilla , Ziyan Yang , Alexander C. Berg , Vicente Ordonez

Referring image segmentation aims to produce a pixel-level mask for the image region described by a natural-language expression. Although pretrained vision-language models have improved semantic grounding, many existing methods still rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Alaa Dalaq , Muzammil Behzad

Deeply learned representations have achieved superior image retrieval performance in a retrieve-then-rerank manner. Recent state-of-the-art single stage model, which heuristically fuses local and global features, achieves promising…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Yuxin Song , Ruolin Zhu , Min Yang , Dongliang He

Recognition of rodent behavior is important for understanding neural and behavioral mechanisms. Traditional manual scoring is time-consuming and prone to human error. We propose MSGL-Transformer, a Multi-Scale Global-Local Transformer for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Muhammad Imran Sharif , Doina Caragea

Referring segmentation aims to generate a segmentation mask for the target instance indicated by a natural language expression. There are typically two kinds of existing methods: one-stage methods that directly perform segmentation on the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-27 Chang Liu , Xudong Jiang , Henghui Ding

Referring Remote Sensing Image Segmentation provides a flexible and fine-grained framework for remote sensing scene analysis via vision-language collaborative interpretation. Current approaches predominantly utilize a three-stage pipeline…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Keyan Chen , Chenyang Liu , Bowen Chen , Jiafan Zhang , Zhengxia Zou , Zhenwei Shi

While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: constrained by the pre-trained fixed visual encoder and failed…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Haotian Zhang , Haoxuan You , Philipp Dufter , Bowen Zhang , Chen Chen , Hong-You Chen , Tsu-Jui Fu , William Yang Wang , Shih-Fu Chang , Zhe Gan , Yinfei Yang
‹ Prev 1 4 5 6 7 8 10 Next ›