English
Related papers

Related papers: Learning Zero-Shot Subject-Driven Video Generation…

200 papers

Few-shot video action recognition is an effective approach to recognizing new categories with only a few labeled examples, thereby reducing the challenges associated with collecting and annotating large-scale video datasets. Existing…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Sarinda Samarasinghe , Mamshad Nayeem Rizve , Navid Kardan , Mubarak Shah

Generative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target dataset distribution. As a result, they capture only the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Zerun Wang , Jiafeng Mao , Xueting Wang , Toshihiko Yamasaki

Deep learning thrives with large neural networks and large datasets. However, larger networks and larger datasets result in longer training times that impede research and development progress. Distributed synchronous SGD offers a potential…

Computer Vision and Pattern Recognition · Computer Science 2018-05-02 Priya Goyal , Piotr Dollár , Ross Girshick , Pieter Noordhuis , Lukasz Wesolowski , Aapo Kyrola , Andrew Tulloch , Yangqing Jia , Kaiming He

Controllable video generation has attracted significant attention, largely due to advances in video diffusion models. In domains such as autonomous driving, it is essential to develop highly accurate predictions for object motions. This…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Ge Ya Luo , Zhi Hao Luo , Anthony Gosselin , Alexia Jolicoeur-Martineau , Christopher Pal

Visual Speech Recognition (VSR) is the process of recognizing or interpreting speech by watching the lip movements of the speaker. Recent machine learning based approaches model VSR as a classification problem; however, the scarcity of…

A major obstacle to the wide-spread adoption of neural retrieval models is that they require large supervised training sets to surpass traditional term-based techniques, which are constructed from raw corpora. In this paper, we propose an…

Information Retrieval · Computer Science 2021-01-28 Ji Ma , Ivan Korotkov , Yinfei Yang , Keith Hall , Ryan McDonald

Recent progress in personalized image generation using diffusion models has been significant. However, development in the area of open-domain and non-fine-tuning personalized image generation is proceeding rather slowly. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Jian Ma , Junhao Liang , Chen Chen , Haonan Lu

Current unified multimodal models for image generation and editing typically rely on massive parameter scales (e.g., >10B), entailing prohibitive training costs and deployment footprints. In this work, we present DeepGen 1.0, a lightweight…

Incorporating a temporal dimension into pretrained image diffusion models for video generation is a prevalent approach. However, this method is computationally demanding and necessitates large-scale video datasets. More critically, the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Dengsheng Chen , Jie Hu , Xiaoming Wei , Enhua Wu

We propose a novel, zero-shot image generation technique called "Visual Concept Blending" that provides fine-grained control over which features from multiple reference images are transferred to a source image. If only a single reference…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Hiroya Makino , Takahiro Yamaguchi , Hiroyuki Sakai

Large-scale demonstration data has powered key breakthroughs in robot manipulation, but collecting that data remains costly and time-consuming. We present Constraint-Preserving Data Generation (CP-Gen), a method that uses a single expert…

Robotics · Computer Science 2025-08-07 Kevin Lin , Varun Ragunath , Andrew McAlinden , Aaditya Prasad , Jimmy Wu , Yuke Zhu , Jeannette Bohg

We study the problem of single-image zero-shot 3D shape reconstruction. Recent works learn zero-shot shape reconstruction through generative modeling of 3D assets, but these models are computationally expensive at train and inference time.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Zixuan Huang , Stefan Stojanov , Anh Thai , Varun Jampani , James M. Rehg

Notable breakthroughs in diffusion modeling have propelled rapid improvements in video generation, yet current foundational model still face critical challenges in simultaneously balancing prompt following, motion plausibility, and visual…

Efficient video-language modeling should consider the computational cost because of a large, sometimes intractable, number of video frames. Parametric approaches such as the attention mechanism may not be ideal since its computational cost…

Computer Vision and Pattern Recognition · Computer Science 2023-01-30 Sungdong Kim , Jin-Hwa Kim , Jiyoung Lee , Minjoon Seo

Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is highly uneven. We present Goal-Driven Data Optimization (GDO), a framework that computes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Rujie Wu , Haozhe Zhao , Hai Ci , Yizhou Wang

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance. However, existing methods struggle with a critical trade-off between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zebin Yao , Lei Ren , Huixing Jiang , Wei Chen , Xiaojie Wang , Ruifan Li , Fangxiang Feng

We propose a simple yet effective zero-shot framework for subject-driven image generation using a vanilla Flux model. By framing the task as grid-based image completion and simply replicating the subject image(s) in a mosaic layout, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Hao Kang , Stathi Fotiadis , Liming Jiang , Qing Yan , Yumin Jia , Zichuan Liu , Min Jin Chong , Xin Lu

Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled attention mechanisms…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Hannan Lu , Xiaohe Wu , Shudong Wang , Xiameng Qin , Xinyu Zhang , Junyu Han , Wangmeng Zuo , Ji Tao

3D-aware image generation necessitates extensive training data to ensure stable training and mitigate the risk of overfitting. This paper first considers a novel task known as One-shot 3D Generative Domain Adaptation (GDA), aimed at…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Ziqiang Li , Yi Wu , Chaoyue Wang , Xue Rui , Bin Li

Subject-driven generation has garnered significant interest recently due to its ability to personalize text-to-image generation. Typical works focus on learning the new subject's private attributes. However, an important fact has not been…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Pengchong Qiao , Lei Shang , Chang Liu , Baigui Sun , Xiangyang Ji , Jie Chen