English
Related papers

Related papers: Panoptic Perception: A Novel Task and Fine-grained…

200 papers

The application of Vision-language foundation models (VLFMs) to remote sensing (RS) imagery has garnered significant attention due to their superior capability in various downstream tasks. A key challenge lies in the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yiguo He , Junjie Zhu , Yiying Li , Xiaoyu Zhang , Chunping Qiu , Jun Wang , Qiangjuan Huang , Ke Yang

This work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step towards panoptic captioning by formulating it as a task of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Kun-Yu Lin , Hongjun Wang , Weining Ren , Kai Han

The multimodal fusion of images and scene captions has been extensively explored and applied in various fields. However, when dealing with complex remote sensing (RS) scenes, existing studies have predominantly concentrated on architectural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Ziyi Wang , Xianping Ma , Ziyao Wang , Hongyang Zhang , Man On Pun

Transparent objects are a common part of everyday life, yet they possess unique visual properties that make them incredibly difficult for standard 3D sensors to produce accurate depth estimates for. In many cases, they often appear as noisy…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 Shreeyak S. Sajjan , Matthew Moore , Mike Pan , Ganesh Nagaraja , Johnny Lee , Andy Zeng , Shuran Song

Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. It serves as an essential testing ground for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Xuzheng Yang , Junzhuo Liu , Peng Wang , Guoqing Wang , Yang Yang , Heng Tao Shen

Language-guided robotic grasping is a rapidly advancing field where robots are instructed using human language to grasp specific objects. However, existing methods often depend on dense camera views and struggle to quickly update scenes,…

Robotics · Computer Science 2024-12-04 Junqiu Yu , Xinlin Ren , Yongchong Gu , Haitao Lin , Tianyu Wang , Yi Zhu , Hang Xu , Yu-Gang Jiang , Xiangyang Xue , Yanwei Fu

Semantic image segmentation is an essential component of modern autonomous driving systems, as an accurate understanding of the surrounding scene is crucial to navigation and action planning. Current state-of-the-art approaches in semantic…

Computer Vision and Pattern Recognition · Computer Science 2016-12-07 Tobias Pohlen , Alexander Hermans , Markus Mathias , Bastian Leibe

Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Yuan Liu , Saihui Hou , Saijie Hou , Jiabao Du , Shibei Meng , Yongzhen Huang

Remote sensing image restoration (RSIR) is essential for recovering high-fidelity imagery from degraded observations, enabling accurate downstream analysis. However, most existing methods focus on single degradation types within homogeneous…

Image and Video Processing · Electrical Eng. & Systems 2026-04-06 Wenli Huang , Yang Wu , Xiaomeng Xin , Zhihong Liu , Jinjun Wang , Ye Deng

The paradigm of self-supervision focuses on representation learning from raw data without the need of labor-consuming annotations, which is the main bottleneck of current data-driven methods. Self-supervision tasks are often used to…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Shuai Chen , Subhradeep Kayal , Marleen de Bruijne

In this age of information, images are a critical medium for storing and transmitting information. With the rapid growth of image data amount, visual compression and visual data perception are two important research topics attracting a lot…

Image and Video Processing · Electrical Eng. & Systems 2024-07-02 Yuefeng Zhang , Chuanmin Jia , Jiannhui Chang , Siwei Ma

Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A critical challenge is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Zilun Zhang , Tiancheng Zhao , Yulong Guo , Jianwei Yin

Remote sensing image segmentation is a specific task of remote sensing image interpretation. A good remote sensing image segmentation algorithm can provide guidance for environmental protection, agricultural production, and urban…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Jiacheng Li

Super Resolution is the problem of recovering a high-resolution image from a single or multiple low-resolution images of the same scene. It is an ill-posed problem since high frequency visual details of the scene are completely lost in…

Image and Video Processing · Electrical Eng. & Systems 2020-04-22 Hamid Reza Vaezi Joze , Ilya Zharkov , Karlton Powell , Carl Ringler , Luming Liang , Andy Roulston , Moshe Lutz , Vivek Pradeep

Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights. Since natural…

Computer Vision and Pattern Recognition · Computer Science 2023-07-19 Di Wang , Jing Zhang , Bo Du , Gui-Song Xia , Dacheng Tao

Multi-task dense scene understanding is a thriving research domain that requires simultaneous perception and reasoning on a series of correlated tasks with pixel-wise prediction. Most existing works encounter a severe limitation of modeling…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Hanrong Ye , Dan Xu

Generalized few-shot semantic segmentation (GFSS) is fundamentally limited by the coverage of novel-class appearances under scarce annotations. While diffusion models can synthesize novel-class images at scale, practical gains are often…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Guohuan Xie , Xin He , Dingying Fan , Le Zhang , Ming-Ming Cheng , Yun Liu

Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visual question answering, and visual grounding. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Ruizhe Ou , Yuan Hu , Fan Zhang , Jiaxin Chen , Yu Liu

We investigate the problem of understanding the message (gist) conveyed by images and their captions as found, for instance, on websites or news articles. To this end, we propose a methodology to capture the meaning of image-caption pairs…

Information Retrieval · Computer Science 2019-04-19 Lydia Weiland , Ioana Hulpus , Simone Paolo Ponzetto , Wolfgang Effelsberg , Laura Dietz

Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic perception datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Tim Brödermann , David Bruggemann , Christos Sakaridis , Kevin Ta , Odysseas Liagouris , Jason Corkill , Luc Van Gool
‹ Prev 1 8 9 10 Next ›