中文
相关论文

相关论文: EchoShot: Multi-Shot Portrait Video Generation

200 篇论文

We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-context learning for image generation, where a query image is…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Ivona Najdenkoska , Animesh Sinha , Abhimanyu Dubey , Dhruv Mahajan , Vignesh Ramanathan , Filip Radenovic

Video generation models have progressed tremendously through large latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Orest Kupyn , Fabian Manhardt , Federico Tombari , Christian Rupprecht

Producing prompt-faithful videos that preserve a user-specified identity remains challenging: models need to extrapolate facial dynamics from sparse reference while balancing the tension between identity preservation and motion naturalness.…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Yixuan Lai , He Wang , Kun Zhou , Tianjia Shao

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

This work presents AnyDoor, a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations in a harmonious way. Instead of tuning parameters for each object, our model is trained only…

计算机视觉与模式识别 · 计算机科学 2024-05-09 Xi Chen , Lianghua Huang , Yu Liu , Yujun Shen , Deli Zhao , Hengshuang Zhao

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents, controlling…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Yuxin Zhang , Fan Tang , Nisha Huang , Haibin Huang , Chongyang Ma , Weiming Dong , Changsheng Xu

Traditional computer vision models often necessitate extensive data acquisition, annotation, and validation. These models frequently struggle in real-world applications, resulting in high false positive and negative rates, and exhibit poor…

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

We introduce MyStyle, a personalized deep generative prior trained with a few shots of an individual. MyStyle allows to reconstruct, enhance and edit images of a specific person, such that the output is faithful to the person's key facial…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Yotam Nitzan , Kfir Aberman , Qiurui He , Orly Liba , Michal Yarom , Yossi Gandelsman , Inbar Mosseri , Yael Pritch , Daniel Cohen-or

High-quality, large-scale data is essential for robust deep learning models in medical applications, particularly ultrasound image analysis. Diffusion models facilitate high-fidelity medical image generation, reducing the costs associated…

图像与视频处理 · 电气工程与系统科学 2024-04-01 Pooria Ashrafian , Milad Yazdani , Moein Heidari , Dena Shahriari , Ilker Hacihaliloglu

We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Weijie Lyu , Ming-Hsuan Yang , Zhixin Shu

Customized generation using diffusion models has made impressive progress in image generation, but remains unsatisfactory in the challenging video generation task, as it requires the controllability of both subjects and motions. To that…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Yujie Wei , Shiwei Zhang , Zhiwu Qing , Hangjie Yuan , Zhiheng Liu , Yu Liu , Yingya Zhang , Jingren Zhou , Hongming Shan

Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or localized attention often struggle with the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Honghao Cai , Xiangyuan Wang , Jing Li , Yunhao Bai , Tianze Zhou , Haohua Chen , Chao Hui , Changhao Qiao , Runqi Wang , Sijie Xu , Yuyang Hao , Zezhou Cui , Yuyuan Yang , Wei Zhu , Yibo Chen , Xu Tang , Yao Hu , Zhen Li

Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While recent advances in image generation have enabled high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Kazuaki Mishima , Antoni Bigata Casademunt , Stavros Petridis , Maja Pantic , Kenji Suzuki

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face-attribute…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Jiazheng Xing , Fei Du , Hangjie Yuan , Pengwei Liu , Hongbin Xu , Hai Ci , Ruigang Niu , Weihua Chen , Fan Wang , Yong Liu

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Zehuan Huang , Hao Wen , Junting Dong , Yaohui Wang , Yangguang Li , Xinyuan Chen , Yan-Pei Cao , Ding Liang , Yu Qiao , Bo Dai , Lu Sheng

This paper presents UniPortrait, an innovative human image personalization framework that unifies single- and multi-ID customization with high face fidelity, extensive facial editability, free-form input description, and diverse layout…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Junjie He , Yifeng Geng , Liefeng Bo

This paper introduces MultiBooth, a novel and efficient technique for multi-concept customization in image generation from text. Despite the significant advancements in customized generation methods, particularly with the success of…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Chenyang Zhu , Kai Li , Yue Ma , Chunming He , Xiu Li

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking…

计算机视觉与模式识别 · 计算机科学 2025-08-12 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su , Xiangqian Wu

Generating human videos from a single image while ensuring high visual quality and precise control is a challenging task, especially in complex scenarios involving multiple individuals and interactions with objects. Existing methods, while…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Zhenzhi Wang , Yixuan Li , Yanhong Zeng , Yuwei Guo , Dahua Lin , Tianfan Xue , Bo Dai