中文
相关论文

相关论文: Metrical-accent Aware Vocal Onset Detection in Pol…

200 篇论文

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voice in acapella…

声音 · 计算机科学 2023-12-19 Jeremy Cochoy

This paper introduces a novel method for quantifying vowel overlap. There is a tension in previous work between using multivariate measures, such as those derived from empirical distributions, and the ability to control for unbalanced data…

计算与语言 · 计算机科学 2024-06-25 Irene Smith , Morgan Sonderegger , The Spade Consortium

This report proposes a polyphonic sound event detection (SED) method for the DCASE 2020 Challenge Task 4. The proposed SED method is based on semi-supervised learning to deal with the different combination of training datasets such as…

音频与语音处理 · 电气工程与系统科学 2020-07-03 Nam Kyun Kim , Hong Kook Kim

This paper introduces a novel recurrent model for music composition that is tailored to the structure of polyphonic music. We propose an efficient new conditional probabilistic factorization of musical scores, viewing a score as a…

声音 · 计算机科学 2019-11-28 John Thickstun , Zaid Harchaoui , Dean P. Foster , Sham M. Kakade

Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…

音频与语音处理 · 电气工程与系统科学 2023-03-27 Khandokar Md. Nayem , Donald S. Williamson

This paper presents a polyphonic pitch tracking system able to extract both framewise and note-based estimates from audio. The system uses several artificial neural networks in a deep layered learning setup. First, cascading networks are…

声音 · 计算机科学 2019-03-19 Anders Elowsson

This paper proposes a weakly-supervised machine learning-based approach aiming at a tool to alert patients about possible respiratory diseases. Various types of pathologies may affect the respiratory system, potentially leading to severe…

声音 · 计算机科学 2023-12-05 Michele Cozzatti , Federico Simonetta , Stavros Ntalampiras

Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Kavya Ranjan Saxena , Vipul Arora

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

We present voice2mode, a method for classification of four singing phonation modes (breathy, neutral (modal), flow, and pressed) using embeddings extracted from large self-supervised speech models. Prior work on singing phonation has relied…

声音 · 计算机科学 2026-02-17 Aju Ani Justus , Ruchit Agrawal , Sudarsana Reddy Kadiri , Shrikanth Narayanan

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their…

音频与语音处理 · 电气工程与系统科学 2023-12-07 Yixuan Zhang , Heming Wang , DeLiang Wang

Singing voice separation and vocal pitch estimation are pivotal tasks in music information retrieval. Existing methods for simultaneous extraction of clean vocals and vocal pitches can be classified into two categories: pipeline methods and…

声音 · 计算机科学 2024-03-20 Haojie Wei , Xueke Cao , Wenbo Xu , Tangpeng Dan , Yueguo Chen

Imprecise vowel articulation can be observed in people with Parkinson's disease (PD). Acoustic features measuring vowel articulation have been demonstrated to be effective indicators of PD in its assessment. Standard clinical vowel…

音频与语音处理 · 电气工程与系统科学 2021-08-18 Yuanyuan Liu , Nelly Penttilä , Tiina Ihalainen , Juulia Lintula , Rachel Convey , Okko Räsänen

Sound event detection (SED) aims at identifying audio events (audio tagging task) in recordings and then locating them temporally (localization task). This last task ends with the segmentation of the frame-level class predictions, that…

音频与语音处理 · 电气工程与系统科学 2019-06-25 Leo Cances , Patrice Guyot , Thomas Pellegrini

We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation…

声音 · 计算机科学 2025-06-06 Hien Ohnaka , Yuma Shirahata , Byeongseon Park , Ryuichi Yamamoto

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on…

声音 · 计算机科学 2023-12-15 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

The onset of a musical note is the earliest time at which a note can be reliably detected. Detection of these musical onsets pose challenges in the presence of ornamentation such as vibrato, bending, and if the attack of the note transient…

音频与语音处理 · 电气工程与系统科学 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

We propose a novel method to model hierarchical metrical structures for both symbolic music and audio signals in a self-supervised manner with minimal domain knowledge. The model trains and inferences on beat-aligned music signals and…

声音 · 计算机科学 2023-01-26 Junyan Jiang , Gus Xia

Recent years has witnessed an increase in technologies that use speech for the sensing of the health of the talker. This survey paper proposes a general taxonomy of the technologies and a broad overview of current progress and challenges.…

Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, beats, and downbeats…

音频与语音处理 · 电气工程与系统科学 2026-02-04 Zhanhong He , Hanyu Meng , David Huang , Roberto Togneri