中文
相关论文

相关论文: ID-Crafter: VLM-Grounded Online RL for Composition…

200 篇论文

Text-to-video generation has made remarkable advancements through diffusion models. However, Multi-Concept Video Customization (MCVC) remains a significant challenge. We identify two key challenges for this task: 1) the identity decoupling…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Yuzhou Huang , Ziyang Yuan , Quande Liu , Qiulin Wang , Xintao Wang , Ruimao Zhang , Pengfei Wan , Di Zhang , Kun Gai

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

We present Concat-ID, a unified framework for identity-preserving video generation. Concat-ID employs variational autoencoders to extract image features, which are then concatenated with video latents along the sequence dimension. It relies…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Yong Zhong , Zhuoyi Yang , Jiayan Teng , Xiaotao Gu , Chongxuan Li

Despite the promising progress in subject-driven image generation, current models often deviate from the reference identities and struggle in complex scenes with multiple subjects. To address this challenge, we introduce OpenSubject, a…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Yexin Liu , Manyuan Zhang , Yueze Wang , Hongyu Li , Dian Zheng , Weiming Zhang , Changsheng Lu , Xunliang Cai , Yan Feng , Peng Pei , Harry Yang

Current video generation models excel at creating short, realistic clips, but struggle with longer, multi-scene videos. We introduce \texttt{DreamFactory}, an LLM-based framework that tackles this challenge. \texttt{DreamFactory} leverages…

人工智能 · 计算机科学 2024-08-22 Zhifei Xie , Daniel Tang , Dingwei Tan , Jacques Klein , Tegawend F. Bissyand , Saad Ezzini

Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory degradation. Most existing solutions either compress historical frames using predefined…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jinzhuo Liu , Jiangning Zhang , Wencan Jiang , Yabiao Wang , Dingkang Liang , Zhucun Xue , Ran Yi , Yong Liu

Recent advances in generative models have achieved high-fidelity in 3D human reconstruction, yet their utility for specific tasks (e.g., human 3D segmentation) remains constrained. We propose HumanCrafter, a unified framework that enables…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Panwang Pan , Tingting Shen , Chenxin Li , Yunlong Lin , Kairun Wen , Jingjing Zhao , Yixuan Yuan

Existing sports video captioning methods often focus on the action yet overlook player identities, limiting their applicability. Although some methods integrate extra information to generate identity-aware descriptions, the player…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zeyu Xi , Haoying Sun , Yaofei Wu , Junchi Yan , Haoran Zhang , Lifang Wu , Liang Wang , Changwen Chen

Recent approaches to controllable 4D video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This dominant paradigm is computationally expensive, requiring large-scale datasets and architectural modifications,…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Yeobin Hong , Suhyeon Lee , Hyungjin Chung , Jong Chul Ye

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

Multimodal learning plays a critical role in e-commerce recommendation platforms today, enabling accurate recommendations and product understanding. However, existing vision-language models, such as CLIP, face key challenges in e-commerce…

Visible-Infrared Person Re-identification (VI-ReID) is a challenging cross-modal pedestrian retrieval task, due to significant intra-class variations and cross-modal discrepancies among different cameras. Existing works mainly focus on…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Kaijie Ren , Lei Zhang

Generating VectorArt from text prompts is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. However, existing research has been mostly limited to the generation of single objects,…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Ayan Banerjee , Nityanand Mathur , Josep Llados , Umapada Pal , Anjan Dutta

Text-to-video (T2V) models have shown remarkable capabilities in generating diverse videos. However, they struggle to produce user-desired stylized videos due to (i) text's inherent clumsiness in expressing specific styles and (ii) the…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Gongye Liu , Menghan Xia , Yong Zhang , Haoxin Chen , Jinbo Xing , Yibo Wang , Xintao Wang , Yujiu Yang , Ying Shan

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven creation has achieved great progress with the identity…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Ze Ma , Daquan Zhou , Chun-Hsiao Yeh , Xue-She Wang , Xiuyu Li , Huanrui Yang , Zhen Dong , Kurt Keutzer , Jiashi Feng

Despite impressive advancements in recent multimodal reasoning approaches, they are still limited in flexibility and efficiency, as these models typically process only a few fixed modality inputs and require updates to numerous parameters.…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Shoubin Yu , Jaehong Yoon , Mohit Bansal

Generating high-fidelity human video with specified identities has attracted significant attention in the content generation community. However, existing techniques struggle to strike a balance between training efficiency and identity…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Xuanhua He , Quande Liu , Shengju Qian , Xin Wang , Tao Hu , Ke Cao , Keyu Yan , Jie Zhang

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

While recent works have achieved great success on image-to-3D object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Jinkun Hao , Junshu Tang , Jiangning Zhang , Ran Yi , Yijia Hong , Moran Li , Weijian Cao , Yating Wang , Chengjie Wang , Lizhuang Ma