English
Related papers

Related papers: Time-Frequency Scattering Accurately Models Audito…

200 papers

The label noise transition matrix, characterizing the probabilities of a training instance being wrongly annotated, is crucial to designing popular solutions to learning with noisy labels. Existing works heavily rely on finding "anchor…

Machine Learning · Computer Science 2021-07-15 Zhaowei Zhu , Yiwen Song , Yang Liu

Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Aemon Yat Fei Chiu , Yujia Xiao , Qiuqiang Kong , Tan Lee

Automatic cover detection -- the task of finding in a audio dataset all covers of a query track -- has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to…

Machine Learning · Computer Science 2020-04-10 Guillaume Doras , Geoffroy Peeters

Flattery is an important aspect of human communication that facilitates social bonding, shapes perceptions, and influences behavior through strategic compliments and praise, leveraging the power of speech to build rapport effectively. Its…

Musical score following is the real-time mapping of a performance to corresponding locations in a musical score. Score following can be used in a variety of applications including automatic page turning and real-time accompaniment. This…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Josephine Cowley

Time Scale Modification (TSM) is a well-researched field; however, no effective objective measure of quality exists. This paper details the creation, subjective evaluation, and analysis of a dataset for use in the development of an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-17 Timothy Roberts , Kuldip K. Paliwal

Distinct striation patterns are observed in the spectrograms of speech and music. This motivated us to propose three novel time-frequency features for speech-music classification. These features are extracted in two stages. First, a preset…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-06 Mrinmoy Bhattacharjee , S. R. M. Prasanna , Prithwijit Guha

Perceptual similarity representations enable music retrieval systems to determine which songs sound most similar to listeners. State-of-the-art approaches based on task-specific training via self-supervised metric learning show promising…

Sound · Computer Science 2026-01-28 Arhan Vohra , Taketo Akama

We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply. Our theoretical…

Machine Learning · Computer Science 2025-02-20 Hyeonsu Jeong , Hye Won Chung

The aim of this paper is to study the feasibility of time-reversal methods in a non homogeneous elastic medium, from data recorded in an acoustic medium. We aim to determine, from partial aperture boundary measurements, the presence and…

Numerical Analysis · Mathematics 2020-10-28 Franck Assous , Moshe Lin

We describe an inference task in which a set of timestamped event observations must be clustered into an unknown number of temporal sequences with independent and varying rates of observations. Various existing approaches to multi-object…

Artificial Intelligence · Computer Science 2013-09-20 Dan Stowell , Mark D. Plumbley

Many practices have been presented in music generation recently. While stylistic music generation using deep learning techniques has became the main stream, these models still struggle to generate music with high musicality, different…

Sound · Computer Science 2021-05-12 Shuqi Dai , Xichu Ma , Ye Wang , Roger B. Dannenberg

We present an audio-to-analysis pipeline that produces composer-level information-theoretic profiles : reflecting compositional vocabulary as it emerges from aggregated performances : from raw recordings, built on a transcription layer…

Sound · Computer Science 2026-05-11 Fred Jalbert-Desforges

In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity…

Sound · Computer Science 2025-04-15 Weixuan Yuan , Qadeer Khan , Vladimir Golkov

Irregularly sampled time series data are common in a variety of fields. Many typical methods for drawing insight from data fail in this case. Here we attempt to generalize methods for clustering trajectories to irregularly and sparsely…

Quantitative Methods · Quantitative Biology 2021-09-01 Gary K. Nave , Swati Padhee , Amanuel Alambo , Tanvi Banerjee , Nirmish Shah , Daniel M. Abrams

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

A MIDI based approach for music recognition is proposed and implemented in this paper. Our Clarinet music retrieval system is designed to search piano MIDI files with high recall and speed. We design a novel melody extraction algorithm that…

Information Retrieval · Computer Science 2023-01-02 Kshitij Alwadhi , Rohan Sharma , Siddhant Sharma

How do different musical traditions achieve tonal coherence? Most computational measures to date have analysed tonal coherence in terms of a single dimension, whereas a multi-dimensional analyses have not been sufficiently explored. We…

Sound · Computer Science 2026-03-31 Weilun Xu , Edward Hall , Martin Rohrmeier

Pitch and meter are two fundamental music features for symbolic music generation tasks, where researchers usually choose different encoding methods depending on specific goals. However, the advantages and drawbacks of different encoding…

Sound · Computer Science 2023-02-01 Yuqiang Li , Shengchen Li , George Fazekas

Existing music generation models are mostly language-based, neglecting the frequency continuity property of notes, resulting in inadequate fitting of rare or never-used notes and thus reducing the diversity of generated samples. We argue…

Sound · Computer Science 2024-08-06 Shipei Liu , Xiaoya Fan , Guowei Wu