中文
相关论文

相关论文: DiscoForcing: A Unified Framework for Real-Time Au…

200 篇论文

Autoregressive (AR) diffusion models offer a promising framework for sequential generation tasks such as video synthesis by combining diffusion modeling with causal inference. Although they support streaming generation, existing AR…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Dingcheng Zhen , Xu Zheng , Ruixin Zhang , Zhiqi Jiang , Yichao Yan , Ming Tao , Shunshun Yin

Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor…

图形学 · 计算机科学 2025-04-24 Lingzhou Mu , Baiji Liu , Ruonan Zhang , Guiming Mo , Jiawei Jin , Kai Zhang , Haozhi Huang

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address…

音频与语音处理 · 电气工程与系统科学 2025-01-17 Siyuan Hou , Shansong Liu , Ruibin Yuan , Wei Xue , Ying Shan , Mangsuo Zhao , Chao Zhang

While current research predominantly focuses on image-based colorization, the domain of video-based colorization remains relatively unexplored. Most existing video colorization techniques operate on a frame-by-frame basis, often overlooking…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Rory Ward , Dan Bigioi , Shubhajit Basak , John G. Breslin , Peter Corcoran

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

图形学 · 计算机科学 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable…

计算与语言 · 计算机科学 2025-06-03 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Ruibo Fu , Wei Liang , Dong Yu

Recent approaches to real-time long video generation typically employ streaming tuning strategies, attempting to train a long-context student using a short-context (memoryless) teacher. In these frameworks, the student performs long…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Shuo Chen , Cong Wei , Sun Sun , Ping Nie , Kai Zhou , Ge Zhang , Ming-Hsuan Yang , Wenhu Chen

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce a novel framework for streaming gesture generation that extends Rolling Diffusion models with structured progressive noise…

机器学习 · 计算机科学 2025-11-20 Evgeniia Vu , Andrei Boiarov , Dmitry Vetrov

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Jiahui Chen , Weida Wang , Runhua Shi , Huan Yang , Chaofan Ding , Zihao Chen

Binaural rendering aims to synthesize binaural audio that mimics natural hearing based on a mono audio and the locations of the speaker and listener. Although many methods have been proposed to solve this problem, they struggle with…

Achieving precise, versatile whole-body character control in physics-based animation remains challenging. Recent diffusion-based policies generate rich and expressive motions but typically rely on gradient-based test-time guidance to…

图形学 · 计算机科学 2026-05-21 Chia-Wen Chen , Yan Wu , Korrawe Karunratanakul , Siyu Tang

Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this…

音频与语音处理 · 电气工程与系统科学 2026-03-24 Jianyi Chen , Rongxiu Zhong , Shilei Zhang , Kun Qian , Jinglei Liu , Yike Guo , Wei Xue

In autonomous driving tasks, trajectory prediction in complex traffic environments requires adherence to real-world context conditions and behavior multimodalities. Existing methods predominantly rely on prior assumptions or generative…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Yiming Xu , Hao Cheng , Monika Sester

Music has the power to evoke intense emotional experiences and regulate the mood of an individual. With the advent of online streaming services, research in music recommendation services has seen tremendous progress. Modern methods…

多媒体 · 计算机科学 2021-10-05 Kunal Vaswani , Yudhik Agrawal , Vinoo Alluri

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking…

计算机视觉与模式识别 · 计算机科学 2025-08-12 He Feng , Yongjia Ma , Donglin Di , Lei Fan , Tonghua Su , Xiangqian Wu

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Diffusion models have seen rapid adoption in robotic imitation learning, enabling autonomous execution of complex dexterous tasks. However, action synthesis is often slow, requiring many steps of iterative denoising, limiting the extent to…

机器人学 · 计算机科学 2024-10-14 Sigmund H. Høeg , Yilun Du , Olav Egeland

Inspired by the impressive performance of recent face image editing methods, several studies have been naturally proposed to extend these methods to the face video editing task. One of the main challenges here is temporal consistency among…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Gyeongman Kim , Hajin Shim , Hyunsu Kim , Yunjey Choi , Junho Kim , Eunho Yang

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey…