中文
相关论文

相关论文: MuPT: A Generative Symbolic Music Pretrained Trans…

200 篇论文

Music often shares notable parallels with language, motivating the use of pretrained large language models (LLMs) for symbolic music understanding and generation. Despite growing interest, the practical effectiveness of adapting…

声音 · 计算机科学 2026-02-02 Deepak Kumar , Emmanouil Karystinaios , Gerhard Widmer , Markus Schedl

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

声音 · 计算机科学 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model…

计算与语言 · 计算机科学 2025-07-18 Andrew Shin , Kunitake Kaneko

Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this…

多媒体 · 计算机科学 2026-01-30 Meng Yang , Jon McCormack , Maria Teresa Llano , Wanchao Su , Chao Lei

As large language models continue to develop, the feasibility and significance of text-based symbolic music tasks have become increasingly prominent. While symbolic music has been widely used in generation tasks, LLM capabilities in…

声音 · 计算机科学 2025-09-30 Jiahao Zhao , Yunjia Li , Wei Li , Kazuyoshi Yoshii

Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a…

声音 · 计算机科学 2026-04-21 Hao Meng , Siyuan Zheng , Shuran Zhou , Qiangqiang Wang , Yang Song

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or…

声音 · 计算机科学 2025-03-04 Daniel Chin , Gus Xia

We explore the use of large language models (LLMs) for music generation using a retrieval system to select relevant examples. We find promising initial results for music generation in a dialogue with the user, especially considering the…

声音 · 计算机科学 2023-12-29 Nicolas Jonason , Luca Casini , Carl Thomé , Bob L. T. Sturm

Large language models perform strongly on general tasks but remain constrained in specialized settings such as music, particularly in the music-entertainment domain, where corpus scale, purity, and the match between data and training…

计算与语言 · 计算机科学 2025-11-19 Kai Tian , Yirong Mao , Wendong Bi , Hanjie Wang , Que Wenhui

Symbolic music understanding, which refers to the understanding of music from the symbolic data (e.g., MIDI format, but not audio), covers many music applications such as genre classification, emotion classification, and music pieces…

声音 · 计算机科学 2021-06-11 Mingliang Zeng , Xu Tan , Rui Wang , Zeqian Ju , Tao Qin , Tie-Yan Liu

Large Language Models (LLMs) are becoming very popular and are used for many different purposes, including creative tasks in the arts. However, these models sometimes have trouble with specific reasoning tasks, especially those that involve…

计算与语言 · 计算机科学 2024-09-10 Anna Kruspe

Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the symbolic music domain including understanding and generation.…

声音 · 计算机科学 2024-08-01 Ziya Zhou , Yuhang Wu , Zhiyue Wu , Xinyue Zhang , Ruibin Yuan , Yinghao Ma , Lu Wang , Emmanouil Benetos , Wei Xue , Yike Guo

Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training. However, symbolic-domain controllable music generation has lagged behind partly due to the lack of a…

声音 · 计算机科学 2025-06-17 Weihan Xu , Julian McAuley , Taylor Berg-Kirkpatrick , Shlomo Dubnov , Hao-Wen Dong

Large Language Models (LLMs) have ushered in a new wave of artificial intelligence advancements impacting every scientific field and discipline. We live in a world where most of the data around us, e.g., text, audio, and music, has a…

信号处理 · 电气工程与系统科学 2025-02-11 Prateek Verma

Recent progress in text-based Large Language Models (LLMs) and their extended ability to process multi-modal sensory data have led us to explore their applicability in addressing music information retrieval (MIR) challenges. In this paper,…

信息检索 · 计算机科学 2025-01-24 Kun Fang , Ziyu Wang , Gus Xia , Ichiro Fujinaga

The ''pretraining-and-finetuning'' paradigm has become a norm for training domain-specific models in natural language processing and computer vision. In this work, we aim to examine this paradigm for symbolic music generation through…

声音 · 计算机科学 2023-11-22 Weihan Xu , Julian McAuley , Shlomo Dubnov , Hao-Wen Dong

Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers only abstract cues,…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Qing Wang , Xiaohang Yang , Yilan Dong , Naveen Raj Govindaraj , Gregory Slabaugh , Shanxin Yuan

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its…

Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the reported success,…

机器学习 · 计算机科学 2024-09-19 Yannis Vasilakis , Rachel Bittner , Johan Pauwels

Multimodal Large Language Models (LLMs) claim "musical understanding" via evaluations that conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music…

声音 · 计算机科学 2025-10-28 Brandon James Carone , Iran R. Roman , Pablo Ripollés
‹ 上一页 1 2 3 10 下一页 ›