English
Related papers

Related papers: Structure-Aware Audio-to-Score Alignment using Pro…

200 papers

High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce.…

Machine Learning · Statistics 2025-03-04 Antonio Sclocchi , Alessandro Favero , Noam Itzhak Levi , Matthieu Wyart

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Ruohan Gao , Kristen Grauman

With the recent growth of remote work, online meetings often encounter challenging audio contexts such as background noise, music, and echo. Accurate real-time detection of music events can help to improve the user experience. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-18 Chandan K. A. Reddy , Vishak Gopa , Harishchandra Dubey , Sergiy Matusevych , Ross Cutler , Robert Aichner

This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired musical relationship between a live musical input and a real-time…

Sound · Computer Science 2025-10-01 Balthazar Bujard , Jérôme Nika , Fédéric Bevilacqua , Nicolas Obin

With the rapid development of neural network applications in NLP, model robustness problem is gaining more attention. Different from computer vision, the discrete nature of texts makes it more challenging to explore robustness in NLP.…

Computation and Language · Computer Science 2023-10-16 Linyang Li , Ke Ren , Yunfan Shao , Pengyu Wang , Xipeng Qiu

Distances on symbolic musical sequences are needed for a variety of applications, from music retrieval to automatic music generation. These musical sequences belong to a given corpus (or style) and it is obvious that a good distance on…

Information Retrieval · Computer Science 2017-09-05 Gaëtan Hadjeres , Frank Nielsen

Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonemic analysis and…

This paper proposes a deep convolutional neural network for performing note-level instrument assignment. Given a polyphonic multi-instrumental music signal along with its ground truth or predicted notes, the objective is to assign an…

Sound · Computer Science 2021-07-30 Carlos Lordelo , Emmanouil Benetos , Simon Dixon , Sven Ahlbäck

While score based generative models, or diffusion models, have found success in image synthesis, they are often coupled with text data or image label to be able to manipulate and conditionally generate images. Even though manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Sandesh Ghimire , Armand Comas , Davin Hill , Aria Masoomi , Octavia Camps , Jennifer Dy

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this problem, we develop a two-stage audiovisual learning framework…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Rui Qian , Di Hu , Heinrich Dinkel , Mengyue Wu , Ning Xu , Weiyao Lin

Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that…

Sound · Computer Science 2024-09-13 Tanisha Hisariya , Huan Zhang , Jinhua Liang

Chorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Qiqi He , Xiaoheng Sun , Yi Yu , Wei Li

In this paper, we compare the performance of using binaural audio features in place of single-channel features for sound event detection. Three different binaural features are studied and evaluated on the publicly available TUT Sound Events…

Sound · Computer Science 2017-10-10 Sharath Adavanne , Tuomas Virtanen

This paper offers a precise, formal definition of an audio-to-score alignment. While the concept of an alignment is intuitively grasped, this precision affords us new insight into the evaluation of audio-to-score alignment algorithms.…

Sound · Computer Science 2020-10-01 John Thickstun , Jennifer Brennan , Harsh Verma

Music genre recognition based on visual representation has been successfully explored over the last years. Recently, there has been increasing interest in attempting convolutional neural networks (CNNs) to achieve the task. However, most of…

Sound · Computer Science 2019-01-28 Caifeng Liu , Lin Feng , Guochao Liu , Huibing Wang , Shenglan Liu

The emergence of a variety of graph-based meaning representations (MRs) has sparked an important conversation about how to adequately represent semantic structure. These MRs exhibit structural differences that reflect different theoretical…

Computation and Language · Computer Science 2020-05-01 Lucia Donatelli , Jonas Groschwitz , Alexander Koller , Matthias Lindemann , Pia Weißenhorn

Music learners can greatly benefit from tools that accurately detect errors in their practice. Existing approaches typically compare audio recordings to music scores using heuristics or learnable models. This paper introduces LadderSym, a…

Learning a mapping between two unrelated domains-such as image and audio, without any supervision is a challenging task. In this work, we propose a distance-preserving generative adversarial model to translate images of human faces into an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-25 Chelhwon Kim , Andrew Port , Mitesh Patel

Understanding the structural characteristics of harmony is essential for an effective use of music as a communication medium. Of the three expressive axes of music (melody, rhythm, harmony), harmony is the foundation on which the emotional…

Multimedia · Computer Science 2020-01-13 Maria Rojo González , Simone Santini

Sound source proximity and distance estimation are of great interest in many practical applications, since they provide significant information for acoustic scene analysis. As both tasks share complementary qualities, ensuring efficient…

Sound · Computer Science 2021-07-27 Daniel Aleksander Krause , Archontis Politis , Annamaria Mesaros