中文
相关论文

相关论文: Exploring how a Generative AI interprets music

200 篇论文

Melody is one of the most important components in music. Unlike other components in music theory, such as harmony and counterpoint, computable features for melody is urgently in need. These features are highly demanded as data-driven…

声音 · 计算机科学 2020-03-23 Zehao Wang , Shicheng Zhang , Xiaoou Chen

This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about the creative…

声音 · 计算机科学 2024-08-15 Jesse Allison , Drew Farrar , Treya Nash , Carlos Román , Morgan Weeks , Fiona Xue Ju

Music is a universal feature of human culture, linked to embodied cognitive functions that drive learning, action, and the emergence of creativity and individuality. Evidence highlights the critical role of statistical learning an implicit…

神经元与认知 · 定量生物学 2025-11-21 Tatsuya Daikoku

The ability of deep neural networks to learn complex data relations and representations is established nowadays, but it generally relies on large sets of training data. This work explores a "piece-specific" autoencoding scheme, in which a…

声音 · 计算机科学 2022-03-09 Axel Marmoret , Jérémy E. Cohen , Frédéric Bimbot

This paper introduces a new large-scale music dataset, MusicNet, to serve as a source of supervision and evaluation of machine learning methods for music research. MusicNet consists of hundreds of freely-licensed classical music recordings…

机器学习 · 统计学 2017-04-07 John Thickstun , Zaid Harchaoui , Sham Kakade

Based on a review of anecdotal beliefs, we explored patterns of track-sequencing within professional music albums. We found that songs with high levels of valence, energy and loudness are more likely to be positioned at the beginning of…

多媒体 · 计算机科学 2024-08-09 Pedro Neto , Martin Hartmann , Geoff Luck , Petri Toiviainen

Generative models have been successfully applied to image style transfer and domain translation. However, there is still a wide gap in the quality of results when learning such tasks on musical audio. Furthermore, most translation models…

声音 · 计算机科学 2018-10-02 Adrien Bitton , Philippe Esling , Axel Chemla-Romeu-Santos

Emotional aspects play an important part in our interaction with music. However, modelling these aspects in MIR systems have been notoriously challenging since emotion is an inherently abstract and subjective experience, thus making it…

声音 · 计算机科学 2019-07-09 Shreyan Chowdhury , Andreu Vall , Verena Haunschmid , Gerhard Widmer

We present a general computational approach that enables a machine to generate a dance for any input music. We encode intuitive, flexible heuristics for what a 'good' dance is: the structure of the dance should align with the structure of…

人工智能 · 计算机科学 2020-06-25 Purva Tendulkar , Abhishek Das , Aniruddha Kembhavi , Devi Parikh

The system presented here shows the feasibility of modeling the knowledge involved in a complex musical activity by integrating sub-symbolic and symbolic processes. This research focuses on the question of whether there is any advantage in…

人工智能 · 计算机科学 2007-05-23 Claudia V. Goldman , Dan Gang , Jeffrey S. Rosenschein , Daniel Lehmann

Symbolic music understanding, which refers to the understanding of music from the symbolic data (e.g., MIDI format, but not audio), covers many music applications such as genre classification, emotion classification, and music pieces…

声音 · 计算机科学 2021-06-11 Mingliang Zeng , Xu Tan , Rui Wang , Zeqian Ju , Tao Qin , Tie-Yan Liu

Current models for audio--sheet music retrieval via multimodal embedding space learning use convolutional neural networks with a fixed-size window for the input audio. Depending on the tempo of a query performance, this window captures more…

声音 · 计算机科学 2018-09-18 Matthias Dorfer , Jan Hajič , Gerhard Widmer

Neurons in the brain communicate information via punctual events called spikes. The timing of spikes is thought to carry rich information, but it is not clear how to leverage this in digital systems. We demonstrate that event-based encoding…

声音 · 计算机科学 2024-02-05 Martim Lisboa , Guillaume Bellec

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Chuang Gan , Deng Huang , Peihao Chen , Joshua B. Tenenbaum , Antonio Torralba

Predictive models for music are studied by researchers of algorithmic composition, the cognitive sciences and machine learning. They serve as base models for composition, can simulate human prediction and provide a multidisciplinary…

机器学习 · 计算机科学 2017-10-04 Jonas Langhabel

Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a…

声音 · 计算机科学 2024-02-29 Manvi Agarwal , Changhong Wang , Gaël Richard

Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges…

声音 · 计算机科学 2025-05-12 Christos Plachouras , Emmanouil Benetos , Johan Pauwels

Music generation with the aid of computers has been recently grabbed the attention of many scientists in the area of artificial intelligence. Deep learning techniques have evolved sequence production methods for this purpose. Yet, a…

神经与进化计算 · 计算机科学 2020-04-09 Majid Farzaneh , Rahil Mahdian Toroghi

Music is inherently made up of complex structures, and representing them as graphs helps to capture multiple levels of relationships. While music generation has been explored using various deep generation techniques, research on…

音频与语音处理 · 电气工程与系统科学 2024-09-13 Wen Qing Lim , Jinhua Liang , Huan Zhang

Deep generative models have been praised for their ability to learn smooth latent representation of images, text, and audio, which can then be used to generate new, plausible data. However, current generative models are unable to work with…