中文
相关论文

相关论文: Few-shot multi-token DreamBooth with LoRa for styl…

200 篇论文

When reading a story, humans can quickly understand new fictional characters with a few observations, mainly by drawing analogies to fictional and real people they already know. This reflects the few-shot and meta-learning essence of…

计算与语言 · 计算机科学 2024-02-06 Mo Yu , Qiujing Wang , Shunchi Zhang , Yisi Sang , Kangsheng Pu , Zekai Wei , Han Wang , Liyan Xu , Jing Li , Yue Yu , Jie Zhou

Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes with multiple…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Zhengguang Zhou , Jing Li , Huaxia Li , Nemo Chen , Xu Tang

Image classification systems often inherit biases from uneven group representation in training data. For example, in face datasets for hair color classification, blond hair may be disproportionately associated with females, reinforcing…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Abhipsa Basu , Aviral Gupta , Abhijnya Bhat , R. Venkatesh Babu

Pre-trained large text-to-image (T2I) models with an appropriate text prompt has attracted growing interests in customized images generation field. However, catastrophic forgetting issue make it hard to continually synthesize new…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Chenxi Liu , Gan Sun , Wenqi Liang , Jiahua Dong , Can Qin , Yang Cong

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as iterative shot…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Kaiwen Zhang , Liming Jiang , Angtian Wang , Jacob Zhiyuan Fang , Tiancheng Zhi , Qing Yan , Hao Kang , Xin Lu , Xingang Pan

Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for contemporary diffusion-based models, which often suffer from identity drift and are limited to…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Jingxuan He , Busheng Su , Finn Wong

We present DreamBooth3D, an approach to personalize text-to-3D generative models from as few as 3-6 casually captured images of a subject. Our approach combines recent advances in personalizing text-to-image models (DreamBooth) with…

Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods struggle with character consistency, leading to artifact generation and inaccurate dialogue…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Ayan Banerjee , Josep Llados , Umapada Pal , Anjan Dutta

Subject-driven image generation aims to synthesize novel depictions of a specific subject across diverse contexts while preserving its core identity features. Achieving both strong identity consistency and high prompt diversity presents a…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Aditi Singhania , Arushi Jain , Krutik Malani , Riddhi Dhawan , Souymodip Chakraborty , Vineet Batra , Ankit Phogat

In the realm of subject-driven text-to-image (T2I) generative models, recent developments like DreamBooth and BLIP-Diffusion have led to impressive results yet encounter limitations due to their intensive fine-tuning demands and substantial…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Shyam Marjit , Harshit Singh , Nityanand Mathur , Sayak Paul , Chia-Mu Yu , Pin-Yu Chen

In this paper, we propose \textbf{CharacterShot}, a controllable and consistent 4D character animation framework that enables any individual designer to create dynamic 3D characters (i.e., 4D character animation) from a single reference…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Junyao Gao , Jiaxing Li , Wenran Liu , Yanhong Zeng , Fei Shen , Kai Chen , Yanan Sun , Cairong Zhao

This work focuses on generating high-quality images with specific style of reference images and content of provided textual descriptions. Current leading algorithms, i.e., DreamBooth and LoRA, require fine-tuning for each style, leading to…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Zhouxia Wang , Xintao Wang , Liangbin Xie , Zhongang Qi , Ying Shan , Wenping Wang , Ping Luo

In this work, we present MotionBooth, an innovative framework designed for animating customized subjects with precise control over both object and camera movements. By leveraging a few images of a specific object, we efficiently fine-tune a…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Jianzong Wu , Xiangtai Li , Yanhong Zeng , Jiangning Zhang , Qianyu Zhou , Yining Li , Yunhai Tong , Kai Chen

Recent advances in diffusion models and parameter-efficient fine-tuning (PEFT) have made text-to-image generation and customization widely accessible, with Low Rank Adaptation (LoRA) able to replicate an artist's style or subject using…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Chenxi Liu , Towaki Takikawa , Alec Jacobson

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Kaihang Pan , Wang Lin , Zhongqi Yue , Tenglong Ao , Liyu Jia , Wei Zhao , Juncheng Li , Siliang Tang , Hanwang Zhang

Lifelong few-shot customization for text-to-image diffusion aims to continually generalize existing models for new tasks with minimal data while preserving old knowledge. Current customization diffusion models excel in few-shot tasks but…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Nan Song , Xiaofeng Yang , Ze Yang , Guosheng Lin

Diffusion-based models have demonstrated impressive capabilities for text-to-image generation and are expected for personalized applications of subject-driven generation, which require the generation of customized concepts with one or a few…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Miao Hua , Jiawei Liu , Fei Ding , Wei Liu , Jie Wu , Qian He

Beyond general recognition tasks, specialized domains and fine-grained settings often encounter data scarcity, especially for tail classes. To obtain less biased and more reliable models under such scarcity, practitioners leverage diffusion…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hoyoung Kim , Minwoo Jang , Jabin Koo , Sangdoo Yun , Jungseul Ok

Recent diffusion models achieve personalization by learning specific subjects, allowing learned attributes to be integrated into generated images. However, personalized human image generation remains challenging due to the need for precise…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Jeongho Kim , Sunghyun Park , Hyoungwoo Park , Sungrack Yun , Jaegul Choo , Seokeon Choi

The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Mohan Zhou , Yalong Bai , Qing Yang , Tiejun Zhao