English
Related papers

Related papers: Multi-subject Open-set Personalization in Video Ge…

200 papers

Personalizing generative text-to-image models has seen remarkable progress, but extending this personalization to text-to-video models presents unique challenges. Unlike static concepts, personalizing text-to-video models has the potential…

Computer vision is increasingly effective at segmenting objects in images and videos; however, scene effects related to the objects -- shadows, reflections, generated smoke, etc -- are typically overlooked. Identifying such scene effects…

Computer Vision and Pattern Recognition · Computer Science 2021-10-04 Erika Lu , Forrester Cole , Tali Dekel , Andrew Zisserman , William T. Freeman , Michael Rubinstein

Generating high-quality and person-generic visual dubbing remains a challenge. Recent innovation has seen the advent of a two-stage paradigm, decoupling the rendering and lip synchronization process facilitated by intermediate…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Tao Liu , Chenpeng Du , Shuai Fan , Feilong Chen , Kai Yu

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Pengxiang Cai , Mengyang Li

Describing video content according to users' needs is a long-held goal. Although existing video captioning methods have made significant progress, the generated captions may not focus on the entity that users are particularly interested in.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Chang Teng , Yunchuan Ma , Guorong Li , Yuankai Qi , Laiyu Qing , Qingming Huang

This paper investigates a solution for enabling in-context capabilities of video diffusion transformers, with minimal tuning required for activation. Specifically, we propose a simple pipeline to leverage in-context generation:…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Zhengcong Fei , Di Qiu , Debang Li , Changqian Yu , Mingyuan Fan

Although the problem of automatic video summarization has recently received a lot of attention, the problem of creating a video summary that also highlights elements relevant to a search query has been less studied. We address this problem…

Computer Vision and Pattern Recognition · Computer Science 2017-09-29 Arun Balajee Vasudevan , Michael Gygli , Anna Volokitin , Luc Van Gool

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or human-annotated captions to generate preference…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yogesh Kulkarni , Pooyan Fazli

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu

Generating video descriptions in natural language (a.k.a. video captioning) is a more challenging task than image captioning as the videos are intrinsically more complicated than images in two aspects. First, videos cover a broader range of…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Shizhe Chen , Jia Chen , Qin Jin

The performance of supervised deep learning algorithms depends significantly on the scale, quality and diversity of the data used for their training. Collecting and manually annotating large amount of data can be both time-consuming and…

Computer Vision and Pattern Recognition · Computer Science 2021-07-02 C. Symeonidis , P. Nousi , P. Tosidis , K. Tsampazis , N. Passalis , A. Tefas , N. Nikolaidis

Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the state-of-the-art of vid2vid has advanced significantly,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ting-Chun Wang , Ming-Yu Liu , Andrew Tao , Guilin Liu , Jan Kautz , Bryan Catanzaro

Traditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video summarization is one…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Jia-Hong Huang , Luka Murn , Marta Mrak , Marcel Worring

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limitations of previous…

Sound · Computer Science 2021-10-15 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Diffusion-based video editing have reached impressive quality and can transform either the global style, local structure, and attributes of given video inputs, following textual edit prompts. However, such solutions typically incur heavy…

Computer Vision and Pattern Recognition · Computer Science 2024-09-02 Kumara Kahatapitiya , Adil Karjauv , Davide Abati , Fatih Porikli , Yuki M. Asano , Amirhossein Habibian

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

While 2D diffusion models have achieved remarkable success in identity-preserving personalization, extending this capability to 3D assets remains a significant challenge due to the complexities of multi-view consistency and spatial control.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jinxin Ai , Matthias Nießner , Ziya Erkoç

This article aims to use graphic engines to simulate a large number of training data that have free annotations and possibly strongly resemble to real-world data. Between synthetic and real, a two-level domain gap exists, involving content…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Yue Yao , Liang Zheng , Xiaodong Yang , Milind Napthade , Tom Gedeon

The goal of this paper is to bypass the need for labelled examples in few-shot video understanding at run time. While proven effective, in many practical video settings even labelling a few examples appears unrealistic. This is especially…

Computer Vision and Pattern Recognition · Computer Science 2022-04-20 Pengwan Yang , Yuki M. Asano , Pascal Mettes , Cees G. M. Snoek

Large text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and furniture, will not be captured by the original model. This…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Xingzhe He , Zhiwen Cao , Nicholas Kolkin , Lantao Yu , Kun Wan , Helge Rhodin , Ratheesh Kalarot