English
Related papers

Related papers: Real-Time Referring Expression Comprehension by Si…

200 papers

Most of the existing work in one-stage referring expression comprehension (REC) mainly focuses on multi-modal fusion and reasoning, while the influence of other factors in this task lacks in-depth exploration. To fill this gap, we conduct…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Gen Luo , Yiyi Zhou , Jiamu Sun , Xiaoshuai Sun , Rongrong Ji

Intelligent robots designed to interact with humans in real scenarios need to be able to refer to entities actively by natural language. In spatial referring expression generation, the ambiguity is unavoidable due to the diversity of…

Robotics · Computer Science 2022-04-05 Mingjiang Liu , Chengli Xiao , Chunlin Chen

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks of VG using the cycle training strategy, with abundant…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Xiaoyu Yang , Lijian Xu , Hao Sun , Hongsheng Li , Shaoting Zhang

The task in referring expression comprehension is to localise the object instance in an image described by a referring expression phrased in natural language. As a language-to-vision matching task, the key to this problem is to learn a…

Computer Vision and Pattern Recognition · Computer Science 2018-12-13 Peng Wang , Qi Wu , Jiewei Cao , Chunhua Shen , Lianli Gao , Anton van den Hengel

Referring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time…

Computer Vision and Pattern Recognition · Computer Science 2020-04-28 Yue Liao , Si Liu , Guanbin Li , Fei Wang , Yanjie Chen , Chen Qian , Bo Li

In this study, we establish a baseline for a new task named multimodal multi-round referring and grounding (MRG), opening up a promising direction for instance-level multimodal dialogues. We present a new benchmark and an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Yunjie Tian , Tianren Ma , Lingxi Xie , Jihao Qiu , Xi Tang , Yuan Zhang , Jianbin Jiao , Qi Tian , Qixiang Ye

Most existing approaches to referring segmentation achieve strong performance only through fine-tuning or by composing multiple pre-trained models, often at the cost of additional training and architectural modifications. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Anna Kukleva , Enis Simsar , Alessio Tonioni , Muhammad Ferjad Naeem , Federico Tombari , Jan Eric Lenssen , Bernt Schiele

Single image super resolution is a very important computer vision task, with a wide range of applications. In recent years, the depth of the super-resolution model has been constantly increasing, but with a small increase in performance, it…

Computer Vision and Pattern Recognition · Computer Science 2018-02-01 Xi Cheng , Xiang Li , Ying Tai , Jian Yang

We propose a unified framework that integrates object detection (OD) and visual grounding (VG) for remote sensing (RS) imagery. To support conventional OD and establish an intuitive prior for VG task, we fine-tune an open-set object…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Karim Radouane , Hanane Azzag , Mustapha lebbah

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC extends the scope to a…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Yaxian Wang , Henghui Ding , Shuting He , Xudong Jiang , Bifan Wei , Jun Liu

Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods attempt to develop various region of interest (RoI)…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Liang Qiao , Ying Chen , Zhanzhan Cheng , Yunlu Xu , Yi Niu , Shiliang Pu , Fei Wu

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding…

Computer Vision and Pattern Recognition · Computer Science 2021-07-30 Zhihao Yuan , Xu Yan , Yinghong Liao , Ruimao Zhang , Sheng Wang , Zhen Li , Shuguang Cui

Scene Graph Generation (SGG) is a challenging task of detecting objects and predicting relationships between objects. After DETR was developed, one-stage SGG models based on a one-stage object detector have been actively studied. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Jinbae Im , JeongYeon Nam , Nokyung Park , Hyungmin Lee , Seunghyun Park

A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence description. Existing methods mainly leverage vanilla soft…

Computer Vision and Pattern Recognition · Computer Science 2021-09-15 Daizong Liu , Xiaoye Qu , Pan Zhou

The rich textual information of large vision-language models (VLMs) combined with the powerful generative prior of pre-trained text-to-image (T2I) diffusion models has achieved impressive performance in single-image super-resolution (SISR).…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Haodong He , Yancheng Bai , Rui Lan , Xu Duan , Lei Sun , Xiangxiang Chu , Gui-Song Xia

Referring image segmentation aims to predict the foreground mask of the object referred by a natural language sentence. Multimodal context of the sentence is crucial to distinguish the referent from the background. Existing methods either…

Computer Vision and Pattern Recognition · Computer Science 2020-10-06 Tianrui Hui , Si Liu , Shaofei Huang , Guanbin Li , Sansi Yu , Faxi Zhang , Jizhong Han

Recent advances in deep learning have brought significant progress in visual grounding tasks such as language-guided video object segmentation. However, collecting large datasets for these tasks is expensive in terms of annotation time,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-10 Ioannis Kazakos , Carles Ventura , Miriam Bellver , Carina Silberer , Xavier Giro-i-Nieto

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Ming Dai , Wenxuan Cheng , Jiang-Jiang Liu , Lingfeng Yang , Zhenhua Feng , Wankou Yang , Jingdong Wang

We propose Score-based Relaxation-guided Generation (SRG), a generative framework based on an approximate formulation of relaxation-guided stochastic differential equations (SDEs) for mixed-integer linear programming. SRG employs a…

Machine Learning · Computer Science 2026-05-13 Ruobing Wang , Xin Li , Yujie Fang , Mingzhong Wang