English
Related papers

Related papers: EraseLoRA: MLLM-Driven Foreground Exclusion and Ba…

200 papers

Object removal requires eliminating not only the target object but also its associated visual effects such as shadows and reflections. However, diffusion-based inpainting and removal methods often introduce artifacts, hallucinate contents,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jixin Zhao , Zhouxia Wang , Peiqing Yang , Shangchen Zhou

Object removal aims to eliminate specified objects from images while plausibly inpainting the affected regions with background content. Current training-free methods typically block attention to object regions within self-attention layers…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Dingming Liu

Recently, diffusion models have emerged as promising newcomers in the field of generative models, shining brightly in image generation. However, when employed for object removal tasks, they still encounter issues such as generating random…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Wenhao Sun , Benlei Cui , Xue-Mei Dong , Jingqun Tang

Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often rely on fine-tuning, prompt engineering, or inference-time…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Dinh-Khoi Vo , Van-Loc Nguyen , Tam V. Nguyen , Minh-Triet Tran , Trung-Nghia Le

Recent advances in object-centric representation learning have shown that slot attention-based methods can effectively decompose visual scenes into object slot representations without supervision. However, existing approaches typically…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Huankun Sheng , Ming Li , Yixiang Wei , Yeying Fan , Yu-Hui Wen , Tieliang Gong , Yong-Jin Liu

In Omnimatte, one aims to decompose a given video into semantically meaningful layers, including the background and individual objects along with their associated effects, such as shadows and reflections. Existing methods often require…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Dvir Samuel , Matan Levy , Nir Darshan , Gal Chechik , Rami Ben-Ari

Attention mechanism, being frequently used to train networks for better feature representations, can effectively disentangle the target object from irrelevant objects in the background. Given an arbitrary image, we find that the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Zhemin Zhang , Xun Gong , Jinyi Wu

Video object removal aims to eliminate target objects from videos while plausibly completing missing regions and preserving spatio-temporal consistency. Although diffusion models have recently advanced this task, it remains challenging to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Dingming Liu , Wenjing Wang , Chen Li , Jing Lyu

The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (LVLMs). Current approaches to address this issue - ranging from expensive data-driven…

Artificial Intelligence · Computer Science 2026-05-26 Yuanzhi Xu , Qian Gao , Jun Fan , Guohui Ding , Zhenyu Yang , Sixue Lin , Yuteng Xiao

Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some of the current methods tend to cheat from the background,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Jinpeng Wang , Yuting Gao , Ke Li , Yiqi Lin , Andy J. Ma , Hao Cheng , Pai Peng , Feiyue Huang , Rongrong Ji , Xing Sun

Multimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious…

Cryptography and Security · Computer Science 2025-05-23 Xuankun Rong , Wenke Huang , Jian Liang , Jinhe Bi , Xun Xiao , Yiming Li , Bo Du , Mang Ye

Incremental or continual learning has been extensively studied for image classification tasks to alleviate catastrophic forgetting, a phenomenon that earlier learned knowledge is forgotten when learning new concepts. For class incremental…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Zekang Zhang , Guangyu Gao , Zhiyuan Fang , Jianbo Jiao , Yunchao Wei

Attention is a core operation in large language models (LLMs) and vision-language models (VLMs). We present BD Attention (BDA), the first lossless algorithmic reformulation of attention. BDA is enabled by a simple matrix identity from Basis…

Machine Learning · Computer Science 2025-10-03 Jialin Zhao

Background subtraction (BGS) aims to extract all moving objects in the video frames to obtain binary foreground segmentation masks. Deep learning has been widely used in this field. Compared with supervised-based BGS methods, unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Yongqi An , Xu Zhao , Tao Yu , Haiyun Guo , Chaoyang Zhao , Ming Tang , Jinqiao Wang

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded samples into the training data to implant behaviors that can be…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Zhifang Zhang , Bojun Yang , Shuo He , Weitong Chen , Wei Emma Zhang , Olaf Maennel , Lei Feng , Miao Xu

Creative processes such as painting often involve creating different components of an image one by one. Can we build a computational model to perform this task? Prior works often fail by making global changes to the image, inserting objects…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Alper Canberk , Maksym Bondarenko , Ege Ozguroglu , Ruoshi Liu , Carl Vondrick

Vision-language models (VLMs) have recently shown remarkable capabilities in visual understanding and generation, but remain vulnerable to adversarial manipulations of visual content. Prior object-hiding attacks primarily rely on…

Cryptography and Security · Computer Science 2026-03-18 Amira Guesmi , Muhammad Shafique

Downstream fine-tuning of vision-language-action (VLA) models enhances robotics, yet exposes the pipeline to backdoor risks. Attackers can pretrain VLAs on poisoned data to implant backdoors that remain stealthy but can trigger harmful…

State-of-the-art diffusion models often rely on parameter-efficient fine-tuning to perform specialized image editing tasks. However, real-world applications require continual adaptation to new tasks while preserving previously learned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yuehao Liu , Weijia Zhang , Xuanming Shang , Zhizhou Chen , Yanhao Ge , Shanyan Guan , Chao Ma

Extracting accurate foreground objects from a scene is an essential step for many video applications. Traditional background subtraction algorithms can generate coarse estimates, but generating high quality masks requires professional…

Computer Vision and Pattern Recognition · Computer Science 2020-09-29 Xiran Wang , Jason Juang , Stanley H. Chan
‹ Prev 1 2 3 10 Next ›