中文
相关论文

相关论文: M6: Multi-generator, Multi-domain, Multi-lingual a…

200 篇论文

Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative…

声音 · 计算机科学 2025-04-21 Heda Zuo , Weitao You , Junxian Wu , Shihong Ren , Pei Chen , Mingxu Zhou , Yujia Lu , Lingyun Sun

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zhaokai Wang , Chenxi Bao , Le Zhuo , Jingrui Han , Yang Yue , Yihong Tang , Victor Shea-Jay Huang , Yue Liao

Music classification, a cornerstone of music information retrieval, supports a wide array of applications. To address the lack of comprehensive datasets and effective methods for sub-genre classification in mainstage dance music, we…

声音 · 计算机科学 2025-08-05 Hongzhi Shu , Xinglin Li , Hongyu Jiang , Minghao Fu , Xinyu Li

Music source separation has been intensively studied in the last decade and tremendous progress with the advent of deep learning could be observed. Evaluation campaigns such as MIREX or SiSEC connected state-of-the-art models and…

音频与语音处理 · 电气工程与系统科学 2022-05-24 Yuki Mitsufuji , Giorgio Fabbro , Stefan Uhlich , Fabian-Robert Stöter , Alexandre Défossez , Minseok Kim , Woosung Choi , Chin-Yun Yu , Kin-Wai Cheuk

The development of artificial intelligent composition has resulted in the increasing popularity of machine-generated pieces, with frequent copyright disputes consequently emerging. There is an insufficient amount of research on the…

人工智能 · 计算机科学 2020-10-16 Yang Deng , Ziyao Xu , Li Zhou , Huanping Liu , Anqi Huang

Music plagiarism detection is gaining more and more attention due to the popularity of music production and society's emphasis on intellectual property. We aim to find fine-grained plagiarism in music pairs since conventional methods are…

声音 · 计算机科学 2023-07-04 Wenxuan Liu , Tianyao He , Chen Gong , Ning Zhang , Hua Yang , Junchi Yan

We propose a novel symbolic music representation and Generative Adversarial Network (GAN) framework specially designed for symbolic multitrack music generation. The main theme of symbolic music generation primarily encompasses the…

声音 · 计算机科学 2024-09-04 Jinlong Zhu , Keigo Sakurai , Ren Togo , Takahiro Ogawa , Miki Haseyama

Generating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically informed control in…

声音 · 计算机科学 2023-09-22 Ben Maman , Johannes Zeitler , Meinard Müller , Amit H. Bermano

We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align…

声音 · 计算机科学 2024-12-31 Shaopeng Wei , Manzhen Wei , Haoyu Wang , Yu Zhao , Gang Kou

Digital advances have transformed the face of automatic music generation since its beginnings at the dawn of computing. Despite the many breakthroughs, issues such as the musical tasks targeted by different machines and the degree to which…

声音 · 计算机科学 2018-12-12 Dorien Herremans , Ching-Hua Chuan , Elaine Chew

While recent years have seen remarkable progress in music generation models, research on their biases across countries, languages, cultures, and musical genres remains underexplored. This gap is compounded by the lack of datasets and…

声音 · 计算机科学 2025-10-03 Ahmet Solak , Florian Grötschla , Luca A. Lanzendörfer , Roger Wattenhofer

We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song. To the best of our knowledge, there are no open-source high-quality dataset representing popular and well-known songs for generative music modeling tasks…

声音 · 计算机科学 2025-06-26 Tawsif Ahmed , Andrej Radonjic , Gollam Rabby

Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental…

音频与语音处理 · 电气工程与系统科学 2025-06-18 Zhongweiyang Xu , Debottam Dutta , Yu-Lin Wei , Romit Roy Choudhury

Most current music source separation (MSS) methods rely on supervised learning, limited by training data quantity and quality. Though web-crawling can bring abundant data, platform-level track labeling often causes metadata mismatches,…

声音 · 计算机科学 2025-10-13 Ji Yu , Yang shuo , Xu Yuetonghui , Liu Mengmei , Ji Qiang , Han Zerui

Procedural Music Generation (PMG) is an emerging field that algorithmically creates music content for video games. By leveraging techniques from simple rule-based approaches to advanced machine learning algorithms, PMG has the potential to…

声音 · 计算机科学 2025-12-16 Shangxuan Luo , Joshua Reiss

Musical instrument classification, a key area in Music Information Retrieval, has gained considerable interest due to its applications in education, digital music production, and consumer media. Recent advances in machine learning,…

声音 · 计算机科学 2024-11-04 Joanikij Chulev

Parallel to rapid advancements in foundation model research, the past few years have witnessed a surge in music AI applications. As AI-generated and AI-augmented music become increasingly mainstream, many researchers in the music AI…

声音 · 计算机科学 2026-02-13 Megan Wei , Mateusz Modrzejewski , Aswin Sivaraman , Dorien Herremans

The use of deep learning to solve problems in literary arts has been a recent trend that has gained a lot of attention and automated generation of music has been an active area. This project deals with the generation of music using raw…

声音 · 计算机科学 2016-12-16 Vasanth Kalingeri , Srikanth Grandhe

While most music generation models generate a mixture of stems (in mono or stereo), we propose to train a multi-stem generative model with 3 stems (bass, drums and other) that learn the musical dependencies between them. To do so, we train…

声音 · 计算机科学 2025-01-08 Simon Rouard , Robin San Roman , Yossi Adi , Axel Roebel

Multimodal Generative Models (MGMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with other sensory modalities…

多媒体 · 计算机科学 2025-11-25 Longzhen Han , Awes Mubarak , Almas Baimagambetov , Nikolaos Polatidis , Thar Baker