中文
相关论文

相关论文: Mozart's Touch: A Lightweight Multi-modal Music Ge…

200 篇论文

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

AI illustrator aims to automatically design visually appealing images for books to provoke rich thoughts and emotions. To achieve this goal, we propose a framework for translating raw descriptions with complex semantics into semantically…

计算机视觉与模式识别 · 计算机科学 2022-09-09 Yiyang Ma , Huan Yang , Bei Liu , Jianlong Fu , Jiaying Liu

Knowledge retrieval with multi-modal queries plays a crucial role in supporting knowledge-intensive multi-modal applications. However, existing methods face challenges in terms of their effectiveness and training efficiency, especially when…

信息检索 · 计算机科学 2024-01-17 Xinwei Long , Jiali Zeng , Fandong Meng , Zhiyuan Ma , Kaiyan Zhang , Bowen Zhou , Jie Zhou

Accompaniment arrangement is a difficult music generation task involving intertwined constraints of melody, harmony, texture, and music structure. Existing models are not yet able to capture all these constraints effectively, especially for…

声音 · 计算机科学 2021-08-26 Jingwei Zhao , Gus Xia

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

Recently, multi-instrument music generation has become a hot topic. Different from single-instrument generation, multi-instrument generation needs to consider inter-track harmony besides intra-track coherence. This is usually achieved by…

声音 · 计算机科学 2023-05-29 Xipin Wei , Junhui Chen , Zirui Zheng , Li Guo , Lantian Li , Dong Wang

Generating music has a few notable differences from generating images and videos. First, music is an art of time, necessitating a temporal model. Second, music is usually composed of multiple instruments/tracks with their own temporal…

音频与语音处理 · 电气工程与系统科学 2020-08-06 Hao-Wen Dong , Wen-Yi Hsiao , Li-Chia Yang , Yi-Hsuan Yang

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Mingshuang Luo , Ruibing Hou , Zhuo Li , Hong Chang , Zimo Liu , Yaowei Wang , Shiguang Shan

This paper presents a study on the use of a real-time music-to-image system as a mechanism to support and inspire musicians during their creative process. The system takes MIDI messages from a keyboard as input which are then interpreted…

人机交互 · 计算机科学 2024-07-09 Meng Yang , Maria Teresa Llano , Jon McCormack

Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reasoning over symbolic scores and expressive performance audio,…

多媒体 · 计算机科学 2026-01-21 Qihao Zhao , Yunqi Cao , Yangyu Huang , Hui Yi Leong , Fan Zhang , Kim-Hui Yap , Wei Hu

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

This document presents some early explorations of applying Softly Masked Language Modelling (SMLM) to symbolic music generation. SMLM can be seen as a generalisation of masked language modelling (MLM), where instead of each element of the…

声音 · 计算机科学 2023-05-12 Nicolas Jonason , Bob L. T. Sturm

Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-generated videos remains challenging due to limited…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Xuanyu Zhang , Weiqi Li , Shijie Zhao , Junlin Li , Li Zhang , Jian Zhang

Generating human portraits is a hot topic in the image generation area, e.g. mask-to-face generation and text-to-face generation. However, these unimodal generation methods lack controllability in image generation. Controllability can be…

计算机视觉与模式识别 · 计算机科学 2024-09-18 Debin Meng , Christos Tzelepis , Ioannis Patras , Georgios Tzimiropoulos

Music is a powerful medium for altering the emotional state of the listener. In recent years, with significant advancement in computing capabilities, artificial intelligence-based (AI-based) approaches have become popular for creating…

人机交互 · 计算机科学 2023-01-18 Adyasha Dash , Kat R. Agres

The ultimate purpose of generative music AI is music production. The studio-lab, a social form within the art-science branch of cross-disciplinarity, is a way to advance music production with AI music models. During a studio-lab experiment…

声音 · 计算机科学 2026-05-18 Emmanuel Deruty , Maarten Grachten

Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Junming Huang , Chi Wang , Letian Li , Guangkai Xu , Donglin Huang , Hao Chen , Qiang Dai , Weiwei Xu

Long-tail recognition is challenging because it requires the model to learn good representations from tail categories and address imbalances across all categories. In this paper, we propose a novel generative and fine-tuning framework,…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Qihao Zhao , Yalun Dai , Hao Li , Wei Hu , Fan Zhang , Jun Liu

By generating plausible and smooth transitions between two image frames, video inbetweening is an essential tool for video editing and long video synthesis. Traditional works lack the capability to generate complex large motions. While…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Maham Tanveer , Yang Zhou , Simon Niklaus , Ali Mahdavi Amiri , Hao Zhang , Krishna Kumar Singh , Nanxuan Zhao