中文
相关论文

相关论文: TMD-Bench: A Multi-Level Evaluation Paradigm for M…

200 篇论文

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training,…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Rui Tian , Mingfei Gao , Mingze Xu , Jiaming Hu , Jiasen Lu , Zuxuan Wu , Yinfei Yang , Afshin Dehghan

Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motion and music is a reflection of advanced human intelligence,…

声音 · 计算机科学 2024-11-05 Fuming You , Minghui Fang , Li Tang , Rongjie Huang , Yongqi Wang , Zhou Zhao

In this paper, we introduce a MusIc conditioned 3D Dance GEneraTion model, named MIDGET based on Dance motion Vector Quantised Variational AutoEncoder (VQ-VAE) model and Motion Generative Pre-Training (GPT) model to generate vibrant and…

声音 · 计算机科学 2024-04-19 Jinwu Wang , Wei Mao , Miaomiao Liu

Temporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos. Despite that they focus on different events,…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Yingsen Zeng , Yujie Zhong , Chengjian Feng , Lin Ma

Dance choreography for a piece of music is a challenging task, having to be creative in presenting distinctive stylistic dance elements while taking into account the musical theme and rhythm. It has been tackled by different approaches such…

声音 · 计算机科学 2022-05-05 Shuang Wu , Shijian Lu , Li Cheng

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven dance video…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Zhikang Dong , Weituo Hao , Ju-Chiang Wang , Peng Zhang , Pawel Polak

As Vision-Language Models (VLMs) increasingly gain traction in medical applications, clinicians are progressively expecting AI systems not only to generate textual diagnoses but also to produce corresponding medical images that integrate…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Junjie Yang , Yuhao Yan , Gang Wu , Yuxuan Wang , Ruoyu Liang , Xinjie Jiang , Xiang Wan , Fenglei Fan , Yongquan Zhang , Feiwei Qin , Changmiao Wang

In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grained motion comprehension - remains under-explored in current benchmarks. To address this gap,…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Wenyi Hong , Yean Cheng , Zhuoyi Yang , Weihan Wang , Lefan Wang , Xiaotao Gu , Shiyu Huang , Yuxiao Dong , Jie Tang

While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error propagation. To…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Ye Tian , Ling Yang , Jiongfan Yang , Anran Wang , Yu Tian , Jiani Zheng , Haochen Wang , Zhiyang Teng , Zhuochen Wang , Yinjie Wang , Yunhai Tong , Mengdi Wang , Xiangtai Li

We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated…

声音 · 计算机科学 2025-05-20 Abhinaba Roy , Geeta Puri , Dorien Herremans

Machine-generated music (MGM) has emerged as a powerful tool with applications in music therapy, personalised editing, and creative inspiration for the music community. However, its unregulated use threatens the entertainment, education,…

声音 · 计算机科学 2026-02-16 Yupei Li , Hanqian Li , Lucia Specia , Björn W. Schuller

Group dance generation from music requires synchronizing multiple dancers while maintaining spatial coordination, making it highly relevant to applications such as film production, gaming, and animation. Recent group dance generation models…

机器学习 · 计算机科学 2026-03-25 Jing Xu , Weiqiang Wang , Cunjian Chen , Jun Liu , Qiuhong Ke

While text-to-image models like DALLE-3 and Stable Diffusion are rapidly proliferating, they often encounter challenges such as hallucination, bias, and the production of unsafe, low-quality output. To effectively address these issues, it…

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Jing Lin , Ruisi Wang , Junzhe Lu , Ziqi Huang , Guorui Song , Ailing Zeng , Xian Liu , Chen Wei , Wanqi Yin , Qingping Sun , Zhongang Cai , Lei Yang , Ziwei Liu

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Ye Zhu , Yu Wu , Nicu Sebe , Yan Yan

Recommending long-form video content demands joint modeling of visual, audio, and textual modalities, yet most benchmarks address only raw features or narrow fusion. We present ViLLA-MMBench, a reproducible, extensible benchmark for…

信息检索 · 计算机科学 2025-08-07 Fatemeh Nazary , Ali Tourani , Yashar Deldjoo , Tommaso Di Noia

We propose MDSC(Music-Dance-Style Consistency), the first evaluation metric that assesses to what degree the dance moves and music match. Existing metrics can only evaluate the motion fidelity and diversity and the degree of rhythmic…

声音 · 计算机科学 2023-12-04 Zixiang Zhou , Weiyuan Li , Baoyuan Wang

We propose MoodNet - A Deep Convolutional Neural Network based architecture to effectively predict the emotion associated with a piece of music given its audio and lyrical content.We evaluate different architectures consisting of varying…

音频与语音处理 · 电气工程与系统科学 2018-11-15 Aniruddha Bhattacharya , K. V. Kadambari

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Jiaben Chen , Xin Yan , Yihang Chen , Siyuan Cen , Zixin Wang , Qinwei Ma , Haoyu Zhen , Kaizhi Qian , Lie Lu , Chuang Gan