中文
相关论文

相关论文: MDT-A2G: Exploring Masked Diffusion Transformers f…

200 篇论文

Diffusion transformers have demonstrated remarkable generation quality, albeit requiring longer training iterations and numerous inference steps. In each denoising step, diffusion transformers encode the noisy inputs to extract the…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Shuai Wang , Zhi Tian , Weilin Huang , Limin Wang

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

多媒体 · 计算机科学 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

Diffusion models that are based on iterative denoising have been recently proposed and leveraged in various generation tasks like image generation. Whereas, as a way inherently built for continuous data, existing diffusion models still have…

计算与语言 · 计算机科学 2023-04-11 Jiaao Chen , Aston Zhang , Mu Li , Alex Smola , Diyi Yang

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

Speech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Chen Qian , Zhaoxiang Zhang , Zhen Lei

Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Jiazhi Guan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Shengyi He , Zhiliang Xu , Haocheng Feng , Errui Ding , Jingdong Wang , Hongtao Xie , Youjian Zhao , Ziwei Liu

Although previous co-speech gesture generation methods are able to synthesize motions in line with speech content, it is still not enough to handle diverse and complicated motion distribution. The key challenges are: 1) the one-to-many…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Lianying Yin , Yijun Wang , Tianyu He , Jinming Liu , Wei Zhao , Bohan Li , Xin Jin , Jianxin Lin

We introduce MDSGen, a novel framework for vision-guided open-domain sound generation optimized for model parameter size, memory consumption, and inference speed. This framework incorporates two key innovations: (1) a redundant video…

声音 · 计算机科学 2025-02-14 Trung X. Pham , Tri Ton , Chang D. Yoo

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Siyuan Wang , Jiawei Liu , Wei Wang , Yeying Jin , Jinsong Du , Zhi Han

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Yichen Peng , Jyun-Ting Song , Siyeol Jung , Ruofan Liu , Haiyang Liu , Xuangeng Chu , Ruicong Liu , Erwin Wu , Hideki Koike , Kris Kitani

Masked diffusion language models (MDMs) have recently gained traction as a viable generative framework for natural language. This can be attributed to its scalability and ease of training compared to other diffusion model paradigms for…

计算与语言 · 计算机科学 2025-08-19 Tejomay Kishor Padole , Suyash P Awate , Pushpak Bhattacharyya

Diffusion Transformers have recently shown remarkable effectiveness in generating high-quality 3D point clouds. However, training voxel-based diffusion models for high-resolution 3D voxels remains prohibitively expensive due to the cubic…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Shentong Mo , Enze Xie , Yue Wu , Junsong Chen , Matthias Nießner , Zhenguo Li

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Jiahui Chen , Yang Huan , Runhua Shi , Chanfan Ding , Xiaoqi Mo , Siyu Xiong , Yinong He

Recent Speech Large Language Models~(LLMs) have achieved impressive capabilities in end-to-end speech interaction. However, the prevailing autoregressive paradigm imposes strict serial constraints, limiting generation efficiency and…

计算与语言 · 计算机科学 2026-02-10 Ziyang Cheng , Yuhao Wang , Heyang Liu , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking the vital role of…

图形学 · 计算机科学 2025-05-09 Jinhe Huang , Yongkang Cheng , Yuming Hang , Gaoge Han , Jinewei Li , Jing Zhang , Xingjian Gu

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the…

人机交互 · 计算机科学 2023-05-09 Sicheng Yang , Zhiyong Wu , Minglei Li , Zhensong Zhang , Lei Hao , Weihong Bao , Ming Cheng , Long Xiao

Existing gesture generation methods primarily focus on upper body gestures based on audio features, neglecting speech content, emotion, and locomotion. These limitations result in stiff, mechanical gestures that fail to convey the true…

声音 · 计算机科学 2026-03-10 Yongkang Cheng , Mingjiang Liang , Shaoli Huang , Gaoge Han , Jifeng Ning , Wei Liu

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yongkang Cheng , Shaoli Huang