中文
相关论文

相关论文: MATHDance: Mamba-Transformer Architecture with Uni…

200 篇论文

Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface,…

Music-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating dance motion for a…

多媒体 · 计算机科学 2023-03-28 Nhat Le , Thang Pham , Tuong Do , Erman Tjiputra , Quang D. Tran , Anh Nguyen

Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of various modalities. To unify multi-modal processing, we…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Chunhao Lu , Qiang Lu , Meichen Dong , Jake Luo

Group dance generation from music has broad applications in film, gaming, and animation production. However, it requires synchronizing multiple dancers while maintaining spatial coordination. As the number of dancers and sequence length…

人工智能 · 计算机科学 2025-07-31 Jing Xu , Weiqiang Wang , Cunjian Chen , Jun Liu , Qiuhong Ke

Generating realistic reactive motions, in which one person reacts to the fixed motions of others, is challenging due to strict interaction constraints and a limited feasible solution space. This paper focuses on a typical scenario: duet…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Xuhai Chen , Zhi Cen , Huaijin Pi , Sida Peng , Xiaowei Zhou , Yong Liu

Recent Mamba-based architectures for video understanding demonstrate promising computational efficiency and competitive performance, yet struggle with overfitting issues that hinder their scalability. To overcome this challenge, we…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yunze Liu , Peiran Wu , Cheng Liang , Junxiao Shen , Limin Wang , Li Yi

Choreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have showcased the feasibility of dance synthesis. However, camera movement synthesis…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Zixuan Wang , Jia Jia , Shikun Sun , Haozhe Wu , Rong Han , Zhenyu Li , Di Tang , Jiaqing Zhou , Jiebo Luo

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Zeyu Zhang , Yiran Wang , Wei Mao , Danning Li , Rui Zhao , Biao Wu , Zirui Song , Bohan Zhuang , Ian Reid , Richard Hartley

Dance and music typically go hand in hand. The complexities in dance, music, and their synchronisation make them fascinating to study from a computational creativity perspective. While several works have looked at generating dance for a…

声音 · 计算机科学 2021-07-21 Gunjan Aggarwal , Devi Parikh

Dance requires skillful composition of complex movements that follow rhythmic, tonal and timbral features of music. Formally, generating dance conditioned on a piece of music can be expressed as a problem of modelling a high-dimensional…

With the popularity of video-based user-generated content (UGC) on social media, harmony, as dictated by human perceptual principles, is critical in assessing the rhythmic consistency of audio-visual UGCs for better user engagement. In this…

多媒体 · 计算机科学 2025-06-10 Xinyi Wu , Haohong Wang , Aggelos K. Katsaggelos

In the domain of 3D biomedical image segmentation, Mamba exhibits the superior performance for it addresses the limitations in modeling long-range dependencies inherent to CNNs and mitigates the abundant computational overhead associated…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Weitong Wu , Zhaohu Xing , Jing Gong , Qin Peng , Lei Zhu

Generative models for audio-conditioned dance motion synthesis map music features to dance movements. Models are trained to associate motion patterns to audio patterns, usually without an explicit knowledge of the human body. This approach…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Davide Moltisanti , Jinyi Wu , Bo Dai , Chen Change Loy

Music-driven 3D dance generation offers significant creative potential, yet practical applications demand versatile and multimodal control. As the highly dynamic and complex human motion covering various styles and genres, dance generation…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Jinlu Zhang , Zixi Kang , Libin Liu , Jianlong Chang , Qi Tian , Feng Gao , Yizhou Wang

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired…

声音 · 计算机科学 2024-10-08 Han Yang , Kun Su , Yutong Zhang , Jiaben Chen , Kaizhi Qian , Gaowen Liu , Chuang Gan

We present a generative framework for inverse design of five-layer transmissive Huygens' metasurfaces (HMSs), addressing a longstanding challenge in achieving full-phase, high-efficiency unit cell designs with minimal full-wave simulations.…

应用物理 · 物理学 2026-03-05 Natanel Nissan , Sherman W. Marcus , Dan Raviv , Raja Giryes , Ariel Epstein

The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they still face challenges in generating coherent and natural…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Yaosi Hu , Zhenzhong Chen , Chong Luo

Synthesizing camera movements from music and dance is highly challenging due to the contradicting requirements and complexities of dance cinematography. Unlike human movements, which are always continuous, dance camera movements involve…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Zixuan Wang , Jiayi Li , Xiaoyu Qin , Shikun Sun , Songtao Zhou , Jia Jia , Jiebo Luo

Recent advances in Text-to-3D (T23D) generative models have enabled the synthesis of diverse, high-fidelity 3D assets from textual prompts. However, existing challenges restrict the development of reliable T23D quality assessment (T23DQA).…

计算机视觉与模式识别 · 计算机科学 2025-11-06 Bingyang Cui , Yujie Zhang , Qi Yang , Zhu Li , Yiling Xu