中文
相关论文

相关论文: Copy-Move Forgery Detection and Question Answering…

200 篇论文

Surveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration…

图像与视频处理 · 电气工程与系统科学 2026-02-10 Yanwei Jiang , Wei Sun , Yingjie Zhou , Xiangyang Zhu , Yuqin Cao , Jun Jia , Yunhao Li , Sijing Wu , Dandan Zhu , Xingkuo Min , Guangtao Zhai

Moving target detection is a challenging computer vision task aimed at generating accurate segmentation maps in diverse in-the-wild color videos captured by static cameras. If backgrounds and targets can be simultaneously extracted and…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Liyang Wang , Shiqian Wu , Shun Fang , Qile Zhu , Jiaxin Wu , Sos Again

While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "temporal blindness". Existing architectures lack intrinsic…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Xiaohe Li , Jiahao Li , Kaixin Zhang , Yuqiang Fang , Leilei Lin , Hong Wang , Haohua Wu , Zide Fan

As a pivotal task that bridges remote visual and linguistic understanding, Remote Sensing Image-Text Retrieval (RSITR) has attracted considerable research interest in recent years. However, almost all RSITR methods implicitly assume that…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Qiya Song , Yiqiang Xie , Yuan Sun , Renwei Dian , Xudong Kang

With abundant, unlabeled real faces, how can we learn robust and transferable facial representations to boost generalization across various face security tasks? We make the first attempt and propose FS-VFM, a scalable self-supervised…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Gaojian Wang , Feng Lin , Tong Wu , Zhisheng Yan , Kui Ren

Feature matching is a crucial task in the field of computer vision, which involves finding correspondences between images. Previous studies achieve remarkable performance using learning-based feature comparison. However, the pervasive…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yesheng Zhang , Xu Zhao

Using sensor data from multiple modalities presents an opportunity to encode redundant and complementary features that can be useful when one modality is corrupted or noisy. Humans do this everyday, relying on touch and proprioceptive…

机器人学 · 计算机科学 2020-12-02 Michelle A. Lee , Matthew Tan , Yuke Zhu , Jeannette Bohg

Video-Question-Answering (VideoQA) comprises the capturing of complex visual relation changes over time, remaining a challenge even for advanced Video Language Models (VLM), i.a., because of the need to represent the visual content to a…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Sofian Chaybouti , Walid Bousselham , Moritz Wolter , Hilde Kuehne

Remote sensing image dehazing (RSID) aims to remove nonuniform and physically irregular haze factors for high-quality image restoration. The emergence of CNNs and Transformers has taken extraordinary strides in the RSID arena. However,…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Huiling Zhou , Xianhao Wu , Hongming Chen , Xiang Chen , Xin He

Manipulation and re-use of images in scientific publications is a concerning problem that currently lacks a scalable solution. Current tools for detecting image duplication are mostly manual or semi-automated, despite the availability of an…

计算机视觉与模式识别 · 计算机科学 2020-03-18 M. Cicconet , H. Elliott , D. L. Richmond , D. Wainstock , M. Walsh

Unsupervised remote sensing change detection aims to monitor and analyze changes from multi-temporal remote sensing images in the same geometric region at different times, without the need for labeled training data. Previous unsupervised…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yating Liu , Yan Lu

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, which struggles to exploit…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Ke Li , Ting Wang , Di Wang , Yongshan Zhu , Yiming Zhang , Tao Lei , Quan Wang

Query-based 3D object detection methods using multi-view images often struggle to efficiently leverage dynamic multi-scale information, e.g., the relationship between the object features and the geometric of the queries are not sufficiently…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mingxi Pang , Dingheng Wang , Zekun Li , Zhenping Sun , Bo Wang , Zhihang Wang , Zhao-Xu Yang

Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods struggle to separate camouflaged object edges due to model freezing, and (2) MLLM-based…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Hua Zhang , Changjiang Luo , Ruoyu Chen

Semantic segmentation of Very High Resolution (VHR) remote sensing images is a fundamental task for many applications. However, large variations in the scales of objects in those VHR images pose a challenge for performing accurate semantic…

计算机视觉与模式识别 · 计算机科学 2023-05-02 Yuanzhi Cai , Lei Fan , Yuan Fang

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Chao Pang , Xingxing Weng , Jiang Wu , Jiayu Li , Yi Liu , Jiaxing Sun , Weijia Li , Shuai Wang , Litong Feng , Gui-Song Xia , Conghui He

Synthetic datasets, recognized for their cost effectiveness, play a pivotal role in advancing computer vision tasks and techniques. However, when it comes to remote sensing image processing, the creation of synthetic datasets becomes…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Jian Song , Hongruixuan Chen , Naoto Yokoya

We present the Surveillance Forgery Image Test Range (SurFITR), a dataset for surveillance-style image forgery detection and localisation, in response to recent advances in open-access image generation models that raise concerns about…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Qizhou Wang , Guansong Pang , Christopher Leckie

Recently, there has been increasing interest in multimodal applications that integrate text with other modalities, such as images, audio and video, to facilitate natural language interactions with multimodal AI systems. While applications…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Roger Ferrod , Luigi Di Caro , Dino Ienco

Visual Question Answering (VQA) is a challenging task of predicting the answer to a question about the content of an image. Prior works directly evaluate the answering models by simply calculating the accuracy of predicted answers. However,…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Kun Li , George Vosselman , Michael Ying Yang