English
Related papers

Related papers: DiscoForcing: A Unified Framework for Real-Time Au…

200 papers

Conditional diffusion models have exhibited superior performance in high-fidelity text-guided visual generation and editing. Nevertheless, prevailing text-guided visual diffusion models primarily focus on incorporating text-visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ling Yang , Zhilong Zhang , Zhaochen Yu , Jingwei Liu , Minkai Xu , Stefano Ermon , Bin Cui

With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Haonan Qiu , Menghan Xia , Yong Zhang , Yingqing He , Xintao Wang , Ying Shan , Ziwei Liu

This paper introduces DanceFusion, a novel framework for reconstructing and generating dance movements synchronized to music, utilizing a Spatio-Temporal Skeleton Diffusion Transformer. The framework adeptly handles incomplete and noisy…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Li Zhao , Zhengmin Lu

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Ruineng Li , Daitao Xing , Huiming Sun , Yuanzhou Ha , Jinglin Shen , Chiuman Ho

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

Sound · Computer Science 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

We propose Polyffusion, a diffusion model that generates polyphonic music scores by regarding music as image-like piano roll representations. The model is capable of controllable music generation with two paradigms: internal control and…

Sound · Computer Science 2023-07-21 Lejun Min , Junyan Jiang , Gus Xia , Jingwei Zhao

We introduce Self Forcing, a novel training paradigm for autoregressive video diffusion models. It addresses the longstanding issue of exposure bias, where models trained on ground-truth context must generate sequences conditioned on their…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Xun Huang , Zhengqi Li , Guande He , Mingyuan Zhou , Eli Shechtman

The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Zhiyao Sun , Tian Lv , Sheng Ye , Matthieu Lin , Jenny Sheng , Yu-Hui Wen , Minjing Yu , Yong-Jin Liu

Suppression of boundary-driven Rayleigh streaming has recently been demonstrated for fluids of spatial inhomogeneity in density and compressibility owing to the competition between the boundary-layer-induced streaming stress and the…

Fluid Dynamics · Physics 2025-10-30 Wei Qiu , Jonas T. Karlsen , Henrik Bruus , Per Augustsson

Self-supervised learning has garnered increasing attention in time series analysis for benefiting various downstream tasks and reducing reliance on labeled data. Despite its effectiveness, existing methods often struggle to comprehensively…

Machine Learning · Computer Science 2025-06-12 Daoyu Wang , Mingyue Cheng , Zhiding Liu , Qi Liu

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber, the first video…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Ngoc-Son Nguyen , Thanh V. T. Tran , Jeongsoo Choi , Hieu-Nghia Huynh-Nguyen , Truong-Son Hy , Van Nguyen

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Chenxu Zhang , Zenan Li , Hongyi Xu , You Xie , Xiaochen Zhao , Tianpei Gu , Guoxian Song , Xin Chen , Chao Liang , Jianwen Jiang , Linjie Luo

We present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with diffusion models…

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over…

Sound · Computer Science 2026-02-18 Tali Dror , Iftach Shoham , Moshe Buchris , Oren Gal , Haim Permuter , Gilad Katz , Eliya Nachmani

We present Follow-Your-Emoji-Faster, an efficient diffusion-based framework for freestyle portrait animation driven by facial landmarks. The main challenges in this task are preserving the identity of the reference portrait, accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yue Ma , Zexuan Yan , Hongyu Liu , Hongfa Wang , Heng Pan , Yingqing He , Junkun Yuan , Ailing Zeng , Chengfei Cai , Heung-Yeung Shum , Zhifeng Li , Wei Liu , Linfeng Zhang , Qifeng Chen

Digital audio effects are widely used by audio engineers to alter the acoustic and temporal qualities of audio data. However, these effects can have a large number of parameters which can make them difficult to learn for beginners and…

Machine Learning · Computer Science 2023-10-02 Kieran Grant

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Michal Geyer , Omer Bar-Tal , Shai Bagon , Tali Dekel

With the great success of diffusion models (DMs) in generating realistic synthetic vision data, many researchers have investigated their potential in decision-making and control. Most of these works utilized DMs to sample directly from the…

Machine Learning · Computer Science 2026-05-19 Hanye Zhao , Xiaoshen Han , Zhengbang Zhu , Minghuan Liu , Yong Yu , De-Chuan Zhan , Weinan Zhang

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Guinan Su , Yanwu Yang , Zhifeng Li

Controllable music generation methods are critical for human-centered AI-based music creation, but are currently limited by speed, quality, and control design trade-offs. Diffusion Inference-Time T-optimization (DITTO), in particular,…

Sound · Computer Science 2024-05-31 Zachary Novack , Julian McAuley , Taylor Berg-Kirkpatrick , Nicholas Bryan