English
Related papers

Related papers: Musical Audio Similarity with Self-supervised Conv…

200 papers

Machine Learning models are being utilized extensively to drive recommender systems, which is a widely explored topic today. This is especially true of the music industry, where we are witnessing a surge in growth. Besides a large chunk of…

Information Retrieval · Computer Science 2023-09-26 Rahul Singh , Pranav Kanuparthi

Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced…

Sound · Computer Science 2026-05-27 Abhinaba Roy , Junyi Liang , Dorien Herremans

We consider a novel task of automatically generating text descriptions of music. Compared with other well-established text generation tasks such as image caption, the scarcity of well-paired music and text datasets makes it a much more…

Sound · Computer Science 2022-09-07 Peining Zhang , Junliang Guo , Linli Xu , Mu You , Junming Yin

Music genre classification is an area that utilizes machine learning models and techniques for the processing of audio signals, in which applications range from content recommendation systems to music recommendation systems. In this…

Sound · Computer Science 2024-05-27 Keoikantse Mogonediwa

Automatic cover detection -- the task of finding in a audio dataset all covers of a query track -- has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to…

Machine Learning · Computer Science 2020-04-10 Guillaume Doras , Geoffroy Peeters

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

We introduce the problem of learning affective correspondence between audio (music) and visual data (images). For this task, a music clip and an image are considered similar (having true correspondence) if they have similar emotion content.…

Multimedia · Computer Science 2019-04-18 Gaurav Verma , Eeshan Gunesh Dhekane , Tanaya Guha

Being able to predict whether a song can be a hit has impor- tant applications in the music industry. Although it is true that the popularity of a song can be greatly affected by exter- nal factors such as social and commercial influences,…

Sound · Computer Science 2017-04-06 Li-Chia Yang , Szu-Yu Chou , Jen-Yu Liu , Yi-Hsuan Yang , Yi-An Chen

Information retrieval from brain responses to auditory and visual stimuli has shown success through classification of song names and image classes presented to participants while recording EEG signals. Information retrieval in the form of…

Sound · Computer Science 2022-07-29 Adolfo G. Ramirez-Aristizabal , Chris Kello

Recent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers.…

Multimedia · Computer Science 2024-05-16 Jiajie Teng , Huiyu Duan , Yucheng Zhu , Sijing Wu , Guangtao Zhai

Deep learning models are typically evaluated to measure and compare their performance on a given task. The metrics that are commonly used to evaluate these models are standard metrics that are used for different tasks. In the field of music…

Sound · Computer Science 2022-04-05 Carlos Hernandez-Olivan , Jorge Abadias Puyuelo , Jose R. Beltran

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic constraints that apply…

With the exponential growth of video content, the need for automated video highlight detection to extract key moments or highlights from lengthy videos has become increasingly pressing. This technology has the potential to enhance user…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Zahidul Islam , Sujoy Paul , Mrigank Rochan

Recent years have witnessed the success of deep learning on the visual sound separation task. However, existing works follow similar settings where the training and testing datasets share the same musical instrument categories, which to…

Multimedia · Computer Science 2022-03-28 Xinchi Zhou , Dongzhan Zhou , Wanli Ouyang , Hang Zhou , Ziwei Liu , Di Hu

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Sangho Lee , Jiwan Chung , Youngjae Yu , Gunhee Kim , Thomas Breuel , Gal Chechik , Yale Song

Recent advancements in web-based audio systems have enabled sufficiently accurate timing control and real-time sound processing capabilities. Numerous specialized music tools, as well as digital audio workstations, are now accessible from…

Sound · Computer Science 2019-05-17 Xavier Favory , Xavier Serra

Identification and extraction of singing voice from within musical mixtures is a key challenge in source separation and machine audition. Recently, deep neural networks (DNN) have been used to estimate 'ideal' binary masks for carefully…

Sound · Computer Science 2015-04-21 Andrew J. R. Simpson , Gerard Roma , Mark D. Plumbley

Cross-modal retrieval aims to retrieve data in one modality by a query in another modality, which has been a very interesting research issue in the field of multimedia, information retrieval, and computer vision, and database. Most existing…

Multimedia · Computer Science 2021-05-06 Donghuo Zeng , Yi Yu , Keizo Oyama

In this study, the notion of perceptual features is introduced for describing general music properties based on human perception. This is an attempt at rethinking the concept of features, in order to understand the underlying human…

Information Retrieval · Computer Science 2014-04-01 Anders Friberg , Erwin Schoonderwaldt , Anton Hedblad , Marco Fabiani , Anders Elowsson

Recently, we proposed a self-attention based music tagging model. Different from most of the conventional deep architectures in music information retrieval, which use stacked 3x3 filters by treating music spectrograms as images, the…

Sound · Computer Science 2019-11-12 Minz Won , Sanghyuk Chun , Xavier Serra
‹ Prev 1 8 9 10 Next ›