中文
相关论文

相关论文: Object-WIPER : Training-Free Object and Associated…

200 篇论文

Video Instance Removal (VIR) requires removing target objects while maintaining background integrity and physical consistency, such as specular reflections and illumination interactions. Despite advancements in text-guided editing, current…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zirui Li , Xinghao Chen , Lingyu Jiang , Dengzhe Hou , Fangzhou Lin , Kazunori Yamada , Xiangbo Gao , Zhengzhong Tu

Deep-learning and large scale language-image training have produced image object detectors that generalise well to diverse environments and semantic classes. However, single-image object detectors trained on internet data are not optimally…

机器人学 · 计算机科学 2024-02-07 Nicolas Harvey Chapman , Feras Dayoub , Will Browne , Chris Lehnert

We propose a novel framework for the task of object-centric video prediction, i.e., extracting the compositional structure of a video sequence, as well as modeling objects dynamics and interactions from visual observations in order to…

计算机视觉与模式识别 · 计算机科学 2023-08-01 Angel Villar-Corrales , Ismail Wahdan , Sven Behnke

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a multimodal query (comprising a reference image and a modification text), without training on annotated triplets. Existing methods typically convert the…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Tianyue Wang , Leigang Qu , Tianyu Yang , Xiangzhao Hao , Yifan Xu , Haiyun Guo , Jinqiao Wang

Video object removal and inpainting are critical tasks in the fields of computer vision and multimedia processing, aimed at restoring missing or corrupted regions in video sequences. Traditional methods predominantly rely on flow-based…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Jie Liu , Zheng Hui

Video editing and synthesis often introduce object inconsistencies, such as frame flicker and identity drift that degrade perceptual quality. To address these issues, we introduce ObjectAlign, a novel framework that seamlessly blends…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Mustafa Munir , Harsh Goel , Xiwen Wei , Minkyu Choi , Sahil Shah , Kartikeya Bhardwaj , Paul Whatmough , Sandeep Chinchali , Radu Marculescu

The traditional image inpainting task aims to restore corrupted regions by referencing surrounding background and foreground. However, the object erasure task, which is in increasing demand, aims to erase objects and generate harmonious…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Fan Li , Zixiao Zhang , Yi Huang , Jianzhuang Liu , Renjing Pei , Bin Shao , Songcen Xu

Object removal refers to the process of erasing designated objects from an image while preserving the overall appearance. Existing works on object removal erase removal targets using image inpainting networks. However, image inpainting…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Changsuk Oh , H. Jin Kim

With the revolution of generative AI, video-related tasks have been widely studied. However, current state-of-the-art video models still lag behind image models in visual quality and user control over generated content. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Haiming Zhu , Yangyang Xu , Jun Yu , Shengfeng He

Recognizing objects from sparse and noisy events becomes extremely difficult when paired images and category labels do not exist. In this paper, we study label-free event-based object recognition where category labels and paired images are…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Hoonhee Cho , Hyeonseong Kim , Yujeong Chae , Kuk-Jin Yoon

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

计算机视觉与模式识别 · 计算机科学 2019-08-22 Ruohan Gao , Kristen Grauman

Motivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm \emph{AutoRemover}, designed specifically for generating street-view videos without any moving objects. In our…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Rong Zhang , Wei Li , Peng Wang , Chenye Guan , Jin Fang , Yuhang Song , Jinhui Yu , Baoquan Chen , Weiwei Xu , Ruigang Yang

Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dynamic object interactions, often resulting in unrealistic movements and frequent violations of…

机器学习 · 计算机科学 2026-04-21 Hiroki Furuta , Heiga Zen , Dale Schuurmans , Aleksandra Faust , Yutaka Matsuo , Percy Liang , Sherry Yang

Diffusion Transformers (DiTs) excel at generation, but their global self-attention makes controllable, reference-image-based editing a distinct challenge. Unlike U-Nets, naively injecting local appearance into a DiT can disrupt its holistic…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Shengrong Gu , Ye Wang , Song Wu , Rui Ma , Qian Wang , Lanjun Wang , Zili Yi

In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visual context-aware editing methods incur data collection effort…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Rui Song , Guo-Hua Wang , Qing-Guo Chen , Weihua Luo , Tongda Xu , Zhening Liu , Yan Wang , Zehong Lin , Jun Zhang

Self-supervised pre-training for images without labels has recently achieved promising performance in image classification. The success of transformer-based methods, ViT and MAE, draws the community's attention to the design of backbone…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Jiantao Wu , Shentong Mo

Object tracking can be formulated as "finding the right object in a video". We observe that recent approaches for class-agnostic tracking tend to focus on the "finding" part, but largely overlook the "object" part of the task, essentially…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Achal Dave , Pavel Tokmakov , Cordelia Schmid , Deva Ramanan

We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification. Our model extracts spatio-temporal tokens from the input video, which are then encoded by a series of…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Anurag Arnab , Mostafa Dehghani , Georg Heigold , Chen Sun , Mario Lučić , Cordelia Schmid

Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Xinyu Chen , Yuyi Qian , Jiang Lin , Shenyi Wang , Gao Wang , Zhiqiu Zhang , Jizhi Zhang , Mingjie Wang , Qiang Tang , Qian Wang , Song Wu , Zili Yi

We present a novel approach to weakly supervised object detection. Instead of annotated images, our method only requires two short videos to learn to detect a new object: 1) a video of a moving object and 2) one or more "negative" videos of…

计算机视觉与模式识别 · 计算机科学 2019-10-01 Rico Jonschkowski , Austin Stone