English
Related papers

Related papers: Metrical-accent Aware Vocal Onset Detection in Pol…

200 papers

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voice in acapella…

Sound · Computer Science 2023-12-19 Jeremy Cochoy

This paper introduces a novel method for quantifying vowel overlap. There is a tension in previous work between using multivariate measures, such as those derived from empirical distributions, and the ability to control for unbalanced data…

Computation and Language · Computer Science 2024-06-25 Irene Smith , Morgan Sonderegger , The Spade Consortium

This report proposes a polyphonic sound event detection (SED) method for the DCASE 2020 Challenge Task 4. The proposed SED method is based on semi-supervised learning to deal with the different combination of training datasets such as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-03 Nam Kyun Kim , Hong Kook Kim

This paper introduces a novel recurrent model for music composition that is tailored to the structure of polyphonic music. We propose an efficient new conditional probabilistic factorization of musical scores, viewing a score as a…

Sound · Computer Science 2019-11-28 John Thickstun , Zaid Harchaoui , Dean P. Foster , Sham M. Kakade

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

This paper presents a polyphonic pitch tracking system able to extract both framewise and note-based estimates from audio. The system uses several artificial neural networks in a deep layered learning setup. First, cascading networks are…

Sound · Computer Science 2019-03-19 Anders Elowsson

This paper proposes a weakly-supervised machine learning-based approach aiming at a tool to alert patients about possible respiratory diseases. Various types of pathologies may affect the respiratory system, potentially leading to severe…

Sound · Computer Science 2023-12-05 Michele Cozzatti , Federico Simonetta , Stavros Ntalampiras

Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-13 Kavya Ranjan Saxena , Vipul Arora

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervised speech models. Prior work on singing phonation has relied…

Sound · Computer Science 2026-02-17 Aju Ani Justus , Ruchit Agrawal , Sudarsana Reddy Kadiri , Shrikanth Narayanan

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-07 Yixuan Zhang , Heming Wang , DeLiang Wang

Singing voice separation and vocal pitch estimation are pivotal tasks in music information retrieval. Existing methods for simultaneous extraction of clean vocals and vocal pitches can be classified into two categories: pipeline methods and…

Sound · Computer Science 2024-03-20 Haojie Wei , Xueke Cao , Wenbo Xu , Tangpeng Dan , Yueguo Chen

Imprecise vowel articulation can be observed in people with Parkinson's disease (PD). Acoustic features measuring vowel articulation have been demonstrated to be effective indicators of PD in its assessment. Standard clinical vowel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-18 Yuanyuan Liu , Nelly Penttilä , Tiina Ihalainen , Juulia Lintula , Rachel Convey , Okko Räsänen

Sound event detection (SED) aims at identifying audio events (audio tagging task) in recordings and then locating them temporally (localization task). This last task ends with the segmentation of the frame-level class predictions, that…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-25 Leo Cances , Patrice Guyot , Thomas Pellegrini

We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation…

Sound · Computer Science 2025-06-06 Hien Ohnaka , Yuma Shirahata , Byeongseon Park , Ryuichi Yamamoto

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on…

Sound · Computer Science 2023-12-15 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

The onset of a musical note is the earliest time at which a note can be reliably detected. Detection of these musical onsets pose challenges in the presence of ornamentation such as vibrato, bending, and if the attack of the note transient…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

We propose a novel method to model hierarchical metrical structures for both symbolic music and audio signals in a self-supervised manner with minimal domain knowledge. The model trains and inferences on beat-aligned music signals and…

Sound · Computer Science 2023-01-26 Junyan Jiang , Gus Xia

Recent years has witnessed an increase in technologies that use speech for the sensing of the health of the talker. This survey paper proposes a general taxonomy of the technologies and a broad overview of current progress and challenges.…

Neurons and Cognition · Quantitative Biology 2024-08-12 Aki Härmä , Bert den Brinker , Ulf Grossekathofer , Okke Ouweltjes , Srikanth Nallanthighal , Sidharth Abrol , Vibhu Sharma

Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, beats, and downbeats…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Zhanhong He , Hanyu Meng , David Huang , Roberto Togneri