English
Related papers

Related papers: SUGAR: Subject-Driven Video Customization in a Zer…

200 papers

Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a fundamental requirement…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Mingxiao Li , Mang Ning , Marie-Francine Moens

Compositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains including product…

Current diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However, these approaches are often ineffective when the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yuchao Gu , Yipin Zhou , Bichen Wu , Licheng Yu , Jia-Wei Liu , Rui Zhao , Jay Zhangjie Wu , David Junhao Zhang , Mike Zheng Shou , Kevin Tang

Given an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and video simultaneously. One of the most challenging issues is an…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Dahye Kim , Jungin Park , Jiyoung Lee , Seongheon Park , Kwanghoon Sohn

Zero-shot paraphrase generation has drawn much attention as the large-scale high-quality paraphrase corpus is limited. Back-translation, also known as the pivot-based method, is typical to this end. Several works leverage different…

Computation and Language · Computer Science 2022-09-23 Zhe Lin , Xiaojun Wan

The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Han Yang , Chuanguang Yang , Qiuli Wang , Zhulin An , Weilun Feng , Libo Huang , Yongjun Xu

Zero-shot learning (ZSL) can be defined by correctly solving a task where no training data is available, based on previous acquired knowledge from different, but related tasks. So far, this area has mostly drawn the attention from computer…

Computer Vision and Pattern Recognition · Computer Science 2018-10-25 Joao Reis , Gil Gonçalves

Audio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Jie Hong , Zeeshan Hayder , Junlin Han , Pengfei Fang , Mehrtash Harandi , Lars Petersson

Existing literature typically treats style-driven and subject-driven generation as two disjoint tasks: the former prioritizes stylistic similarity, whereas the latter insists on subject consistency, resulting in an apparent antagonism. We…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Shaojin Wu , Mengqi Huang , Yufeng Cheng , Wenxu Wu , Jiahe Tian , Yiming Luo , Fei Ding , Qian He

A major obstacle to the wide-spread adoption of neural retrieval models is that they require large supervised training sets to surpass traditional term-based techniques, which are constructed from raw corpora. In this paper, we propose an…

Information Retrieval · Computer Science 2021-01-28 Ji Ma , Ivan Korotkov , Yinfei Yang , Keith Hall , Ryan McDonald

We study the problem of recognizing visual entities from the textual descriptions of their classes. Specifically, given birds' images with free-text descriptions of their species, we learn to classify images of previously-unseen species…

Computation and Language · Computer Science 2020-10-08 Tzuf Paz-Argaman , Yuval Atzmon , Gal Chechik , Reut Tsarfaty

Clustering tabular data remains a significant open challenge in data analysis and machine learning. Unlike for image data, similarity between tabular records often varies across datasets, making the definition of clusters highly…

Machine Learning · Computer Science 2025-10-27 Patryk Marszałek , Tomasz Kuśmierczyk , Witold Wydmański , Jacek Tabor , Marek Śmieja

Feedforward monocular face capture methods seek to reconstruct posed faces from a single image of a person. Current state of the art approaches have the ability to regress parametric 3D face models in real-time across a wide range of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Kelian Baert , Shrisha Bharadwaj , Fabien Castan , Benoit Maujean , Marc Christie , Victoria Abrevaya , Adnane Boukhayma

Video personalization aims to generate videos that faithfully reflect a user-provided subject while following a text prompt. However, existing approaches often rely on heavy video-based finetuning or large-scale video datasets, which impose…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Hyunkoo Lee , Wooseok Jang , Jini Yang , Taehwan Kim , Sangoh Kim , Sangwon Jung , Seungryong Kim

We present Magic Insert, a method for dragging-and-dropping subjects from a user-provided image into a target image of a different style in a physically plausible manner while matching the style of the target image. This work formalizes the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Nataniel Ruiz , Yuanzhen Li , Neal Wadhwa , Yael Pritch , Michael Rubinstein , David E. Jacobs , Shlomi Fruchter

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

Multimedia · Computer Science 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

Recent advancements in large-scale pre-training of visual-language models on paired image-text data have demonstrated impressive generalization capabilities for zero-shot tasks. Building on this success, efforts have been made to adapt…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Shahzad Ahmad , Sukalpa Chanda , Yogesh S Rawat

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Inkyu Shin , Qihang Yu , Xiaohui Shen , In So Kweon , Kuk-Jin Yoon , Liang-Chieh Chen

Existing datasets for manually labelled query-based video summarization are costly and thus small, limiting the performance of supervised deep video summarization models. Self-supervision can address the data sparsity challenge by using a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Jia-Hong Huang , Luka Murn , Marta Mrak , Marcel Worring

Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondences. However, these methods often fail to accurately edit the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Xiangpeng Yang , Linchao Zhu , Hehe Fan , Yi Yang