中文
相关论文

相关论文: Bidirectional Autoregressive Diffusion Model for D…

200 篇论文

In this paper, we present DreaMoving, a diffusion-based controllable video generation framework to produce high-quality customized human videos. Specifically, given target identity and posture sequences, DreaMoving can generate a video of…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Mengyang Feng , Jinlin Liu , Kai Yu , Yuan Yao , Zheng Hui , Xiefan Guo , Xianhui Lin , Haolan Xue , Chen Shi , Xiaowen Li , Aojie Li , Xiaoyang Kang , Biwen Lei , Miaomiao Cui , Peiran Ren , Xuansong Xie

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially under complex,…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Yan Zhang , Han Zou , Lincong Feng , Cong Xie , Ruiqi Yu , Zhenpeng Zhan

The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the…

声音 · 计算机科学 2024-02-05 Marco Pasini , Maarten Grachten , Stefan Lattner

Music-to-dance generation aims to translate auditory signals into expressive human motion, with broad applications in virtual reality, choreography, and digital entertainment. Despite promising progress, the limited generation efficiency of…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Kaixing Yang , Xulong Tang , Ziqiao Peng , Xiangyue Zhang , Puwei Wang , Jun He , Hongyan Liu

The generation of natural human motion interactions is a hot topic in computer vision and computer animation. It is a challenging task due to the diversity of possible human motion interactions. Diffusion models, which have already shown…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Baptiste Chopin , Hao Tang , Mohamed Daoudi

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2)…

声音 · 计算机科学 2025-11-13 Shulei Ji , Zihao Wang , Jiaxing Yu , Xiangyuan Yang , Shuyu Li , Songruoyao Wu , Kejun Zhang

Diffusion models show promise for dynamic scene deblurring; however, existing studies often fail to leverage the intrinsic nature of the blurring process within diffusion models, limiting their full potential. To address it, we present a…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Jin-Ting He , Fu-Jen Tsai , Yan-Tsung Peng , Min-Hung Chen , Chia-Wen Lin , Yen-Yu Lin

We study a challenging task, conditional human motion generation, which produces plausible human motion sequences according to various conditional inputs, such as action classes or textual descriptors. Since human motions are highly diverse…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Xin Chen , Biao Jiang , Wen Liu , Zilong Huang , Bin Fu , Tao Chen , Jingyi Yu , Gang Yu

Generating 3D human motion from text descriptions remains challenging due to the diverse and complex nature of human motion. While existing methods excel within the training distribution, they often struggle with out-of-distribution…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Zongye Zhang , Bohan Kong , Qingjie Liu , Yunhong Wang

Generating realistic human-human interactions is a challenging task that requires not only high-quality individual body and hand motions, but also coherent coordination among all interactants. Due to limitations in available data and…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Pablo Ruiz-Ponce , Sergio Escalera , José García-Rodríguez , Jiankang Deng , Rolandos Alexandros Potamias

Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Shengqi Liu , Yuhao Cheng , Zhuo Chen , Xingyu Ren , Wenhan Zhu , Lincheng Li , Mengxiao Bi , Xiaokang Yang , Yichao Yan

In medical imaging, the diffusion models have shown great potential for synthetic image generation tasks. However, these approaches often lack the interpretable connections between the generated and real images and can create anatomically…

图像与视频处理 · 电气工程与系统科学 2026-02-12 Jian-Qing Zheng , Yuanhan Mo , Yang Sun , Jiahua Li , Fuping Wu , Ziyang Wang , Tonia Vincent , Bartłomiej W. Papież

Diffusion models have shown promising results in cross-modal generation tasks involving audio and music, such as text-to-sound and text-to-music generation. These text-controlled music generation models typically focus on generating music…

声音 · 计算机科学 2024-10-24 Tornike Karchkhadze , Mohammad Rasool Izadi , Ke Chen , Gerard Assayag , Shlomo Dubnov

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

We introduce the Cross Human Motion Diffusion Model (CrossDiff), a novel approach for generating high-quality human motion based on textual descriptions. Our method integrates 3D and 2D information using a shared transformer network within…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Zeping Ren , Shaoli Huang , Xiu Li

In this work, we propose a novel framework to enable diffusion models to adapt their generation quality based on real-time network bandwidth constraints. Traditional diffusion models produce high-fidelity images by performing a fixed number…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Xi Zhang , Hanwei Zhu , Yan Zhong , Jiamang Wang , Weisi Lin

Fashionable image generation aims to synthesize images of diverse fashion prevalent around the globe, helping fashion designers in real-time visualization by giving them a basic customized structure of how a specific design preference would…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Krishna Sri Ipsit Mantri , Nevasini Sasikumar

Autoregressive models (ARMs) and diffusion models (DMs) represent two leading paradigms in generative modeling, each excelling in distinct areas: ARMs in global context modeling and long-sequence generation, and DMs in generating…

机器学习 · 计算机科学 2024-10-08 Hyungjin Chung , Dohun Lee , Jong Chul Ye

We introduce a method for generating realistic pedestrian trajectories and full-body animations that can be controlled to meet user-defined goals. We draw on recent advances in guided diffusion modeling to achieve test-time controllability…

计算机视觉与模式识别 · 计算机科学 2023-04-05 Davis Rempe , Zhengyi Luo , Xue Bin Peng , Ye Yuan , Kris Kitani , Karsten Kreis , Sanja Fidler , Or Litany

We present DiverseMotion, a new approach for synthesizing high-quality human motions conditioned on textual descriptions while preserving motion diversity.Despite the recent significant process in text-based human motion generation,existing…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Yunhong Lou , Linchao Zhu , Yaxiong Wang , Xiaohan Wang , Yi Yang