中文
相关论文

相关论文: HunyuanCustom: A Multimodal-Driven Architecture fo…

200 篇论文

Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Huanghao Yin , Shenkun Xu , Kanle Shi , Junhai Yong , Bin Wang

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

声音 · 计算机科学 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consistency. To mitigate these drawbacks, we propose UniAVGen, a…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Guozhen Zhang , Zixiang Zhou , Teng Hu , Ziqiao Peng , Youliang Zhang , Yi Chen , Yuan Zhou , Qinglin Lu , Limin Wang

Recent diffusion-based human image animation techniques have demonstrated impressive success in synthesizing videos that faithfully follow a given reference identity and a sequence of desired movement poses. Despite this, there are still…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Xiang Wang , Shiwei Zhang , Changxin Gao , Jiayu Wang , Xiaoqiang Zhou , Yingya Zhang , Luxin Yan , Nong Sang

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

声音 · 计算机科学 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability…

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFoley, a model…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Ziyang Chen , Prem Seetharaman , Bryan Russell , Oriol Nieto , David Bourgin , Andrew Owens , Justin Salamon

Text-to-3D generation, which synthesizes 3D assets according to an overall text description, has significantly progressed. However, a challenge arises when the specific appearances need customizing at designated viewpoints but referring…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Junkai Yan , Yipeng Gao , Qize Yang , Xihan Wei , Xuansong Xie , Ancong Wu , Wei-Shi Zheng

A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yu Yuan , Xijun Wang , Tharindu Wickremasinghe , Zeeshan Nadir , Bole Ma , Stanley H. Chan

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

图形学 · 计算机科学 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Recent advances in subject-driven video generation with large diffusion models have enabled personalized content synthesis conditioned on user-provided subjects. However, existing methods lack fine-grained temporal control over subject…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Sharath Girish , Viacheslav Ivanov , Tsai-Shien Chen , Hao Chen , Aliaksandr Siarohin , Sergey Tulyakov

Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Luozheng Qin , Jia Gong , Qian Qiao , Tianjiao Li , Li Xu , Haoyu Pan , Chao Qu , Zhiyu Tan , Hao Li

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge this gap often suffer…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Yujie Wei , Xinyu Liu , Shiwei Zhang , Hangjie Yuan , Jinbo Xing , Zhekai Chen , Xiang Wang , Haonan Qiu , Rui Zhao , Yutong Feng , Ruihang Chu , Yingya Zhang , Yike Guo , Xihui Liu , Hongming Shan

Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and next-generation e-commerce. However, despite rapid progress…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Hyun-kyu Ko , Jihyeon Park , Younghyun Kim , Dongheok Park , Eunbyung Park

Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple modalities. In this paper, we present Mogao, a unified…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Chao Liao , Liyang Liu , Xun Wang , Zhengxiong Luo , Xinyu Zhang , Wenliang Zhao , Jie Wu , Liang Li , Zhi Tian , Weilin Huang