中文
相关论文

相关论文: EchoShot: Multi-Shot Portrait Video Generation

200 篇论文

In this paper, we propose \textbf{CharacterShot}, a controllable and consistent 4D character animation framework that enables any individual designer to create dynamic 3D characters (i.e., 4D character animation) from a single reference…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Junyao Gao , Jiaxing Li , Wenran Liu , Yanhong Zeng , Fei Shen , Kai Chen , Yanan Sun , Cairong Zhao

The recent innovations and breakthroughs in diffusion models have significantly expanded the possibilities of generating high-quality videos for the given prompts. Most existing works tackle the single-scene scenario with only one video…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Fuchen Long , Zhaofan Qiu , Ting Yao , Tao Mei

Image synthesis is expected to provide value for the translation of machine learning methods into clinical practice. Fundamental problems like model robustness, domain transfer, causal modelling, and operator training become approachable…

计算机视觉与模式识别 · 计算机科学 2024-02-22 Hadrien Reynaud , Mengyun Qiao , Mischa Dombrowski , Thomas Day , Reza Razavi , Alberto Gomez , Paul Leeson , Bernhard Kainz

In controllable generation tasks, flexibly manipulating the generated images to attain a desired appearance or structure based on a single input image cue remains a critical and longstanding challenge. Achieving this requires the effective…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xi Wang , Yichen Peng , Heng Fang , Yilin Wang , Haoran Xie , Xi Yang , Chuntao Li

Foundation models have recently gained significant attention because of their generalizability and adaptability across multiple tasks and data distributions. Although medical foundation models have emerged, solutions for cardiac imaging,…

计算机视觉与模式识别 · 计算机科学 2025-01-30 Sekeun Kim , Pengfei Jin , Sifan Song , Cheng Chen , Yiwei Li , Hui Ren , Xiang Li , Tianming Liu , Quanzheng Li

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view synthesis as a…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Ruishu Zhu , Zhihao Huang , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

Recent advancements in diffusion-based technologies have made significant strides, particularly in identity-preserved portrait generation (IPG). However, when using multiple reference images from the same ID, existing methods typically…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Yifang Xu , Benxiang Zhai , Yunzhuo Sun , Ming Li , Yang Li , Sidan Du

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yuxiang Shi , Zhe Li , Yanwen Wang , Hao Zhu , Xun Cao , Ligang Liu

Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Brian Chao , Lior Yariv , Howard Xiao , Gordon Wetzstein

Diffusion-based models have gained wide adoption in the virtual human generation due to their outstanding expressiveness. However, their substantial computational requirements have constrained their deployment in real-time interactive…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Haojie Yu , Zhaonian Wang , Yihan Pan , Meng Cheng , Hao Yang , Chao Wang , Tao Xie , Xiaoming Xu , Xiaoming Wei , Xunliang Cai

In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Bingyuan Zhang , Xulong Zhang , Ning Cheng , Jun Yu , Jing Xiao , Jianzong Wang

Over-the-shoulder dialogue videos are essential in films, short dramas, and advertisements, providing visual variety and enhancing viewers' emotional connection. Despite their importance, such dialogue scenes remain largely underexplored in…

计算机视觉与模式识别 · 计算机科学 2025-08-18 Yuang Zhang , Junqi Cheng , Haoyu Zhao , Jiaxi Gu , Fangyuan Zou , Zenghui Lu , Peng Shu

In this paper, we introduce PoseCrafter, a one-shot method for personalized video generation following the control of flexible poses. Built upon Stable Diffusion and ControlNet, we carefully design an inference process to produce…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Yong Zhong , Min Zhao , Zebin You , Xiaofeng Yu , Changwang Zhang , Chongxuan Li

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

Subject-driven generation is a critical task in creative AI; yet current state-of-the-art methods present a stark trade-off. They either rely on computationally expensive, per-subject fine-tuning, sacrificing efficiency and zero-shot…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Ruixiao Dong , Zhendong Wang , Keli Liu , Li Li , Ying Chen , Kai Li , Daowen Li , Houqiang Li

We present BootComp, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments. Here, the main bottleneck is data acquisition for training: collecting a large-scale…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Yisol Choi , Sangkyung Kwak , Sihyun Yu , Hyungwon Choi , Jinwoo Shin

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. Recent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Tianqi Zhang , Ziyi Wang , Wenzhao Zheng , Weiliang Chen , Yuanhui Huang , Zhengyang Huang , Jie Zhou , Jiwen Lu

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on low-level visual…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Xiaoxue Wu , Xinyuan Chen , Yaohui Wang , Yu Qiao

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Yuyang Sha , Zijie Lou , Youyun Tang , Xiaochao Qu , Zheng Qu , Ben Xia , Haoxiang Li , Ting Liu , Luoqi Liu