English
Related papers

Related papers: Real-Time Pitch/F0 Detection Using Spectrogram Ima…

200 papers

Pitch detection is a fundamental problem in speech processing as F0 is used in a large number of applications. Recent articles have proposed deep learning for robust pitch tracking. In this paper, we consider voicing detection as a…

Sound · Computer Science 2019-03-06 Thomas Drugman , Goeric Huybrechts , Viacheslav Klimkov , Alexis Moinet

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. F0 estimation is important in various speech processing and music…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Satwinder Singh , Ruili Wang , Yuanhang Qiu

This paper presents a polyphonic pitch tracking system able to extract both framewise and note-based estimates from audio. The system uses several artificial neural networks in a deep layered learning setup. First, cascading networks are…

Sound · Computer Science 2019-03-19 Anders Elowsson

In this paper, we present a novel approach for contour detection with Convolutional Neural Networks. A multi-scale CNN learning framework is designed to automatically learn the most relevant features for contour patch detection. Our method…

Computer Vision and Pattern Recognition · Computer Science 2017-05-10 Teck Wee Chua , Li Shen

Pitch or fundamental frequency (f0) extraction is a fundamental problem studied extensively for its potential applications in speech and clinical applications. In literature, explicit mode specific (modal speech or singing voice or…

Sound · Computer Science 2019-04-23 Pradeep Rengaswamy , Gurunath Reddy M , Krothapalli Sreenivasa Rao

This paper introduces a novel method to separate noisy speech into low or high frequency frames, in order to improve fundamental frequency (F0) estimation accuracy. In this proposal, the target signal is analyzed by means of the ensemble…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 A. Queiroz , R. Coelho

In this paper, we propose an accurate edge detector using richer convolutional features (RCF). Since objects in nature images have various scales and aspect ratios, the automatically learned rich hierarchical representations by CNNs are…

Computer Vision and Pattern Recognition · Computer Science 2019-07-04 Yun Liu , Ming-Ming Cheng , Xiaowei Hu , Kai Wang , Xiang Bai

We address a challenging fine-grain classification problem: recognizing a font style from an image of text. In this task, it is very easy to generate lots of rendered font examples but very hard to obtain real-world labeled images. This…

Computer Vision and Pattern Recognition · Computer Science 2015-04-02 Zhangyang Wang , Jianchao Yang , Hailin Jin , Eli Shechtman , Aseem Agarwala , Jonathan Brandt , Thomas S. Huang

The fundamental frequency (F0) represents pitch in speech that determines prosodic characteristics of speech and is needed in various tasks for speech analysis and synthesis. Despite decades of research on this topic, F0 estimation at low…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-03 Akihiro Kato , Tomi Kinnunen

Modern day audio signal classification techniques lack the ability to classify low feature audio signals in the form of spectrographic temporal frequency data representations. Additionally, currently utilized techniques rely on full diverse…

Sound · Computer Science 2024-10-30 Noel Elias

Pulse shape discrimination plays a key role in improving the signal-to-background ratio in NEOS analysis by removing fast neutrons. Identifying particles by looking at the tail of the waveform has been an effective and plausible approach…

In the past decade, Convolutional Neural Networks (CNNs) have been demonstrated successful for object detections. However, the size of network input is limited by the amount of memory available on GPUs. Moreover, performance degrades when…

Computer Vision and Pattern Recognition · Computer Science 2017-06-28 Zibo Meng , Xiaochuan Fan , Xin Chen , Min Chen , Yan Tong

In this paper we introduce a new method for text detection in natural images. The method comprises two contributions: First, a fast and scalable engine to generate synthetic images of text in clutter. This engine overlays synthetic text to…

Computer Vision and Pattern Recognition · Computer Science 2016-04-25 Ankush Gupta , Andrea Vedaldi , Andrew Zisserman

In this article we explore how the different semantics of spectrograms' time and frequency axes can be exploited for musical tempo and key estimation using Convolutional Neural Networks (CNN). By addressing both tasks with the same network…

Sound · Computer Science 2019-03-27 Hendrik Schreiber , Meinard Müller

Many edge and contour detection algorithms give a soft-value as an output and the final binary map is commonly obtained by applying an optimal threshold. In this paper, we propose a novel method to detect image contours from the extracted…

Computer Vision and Pattern Recognition · Computer Science 2021-05-12 Zahra Mousavi Kouzehkanan , Reshad Hosseini , Babak Nadjar Araabi

Fundamental frequency (F0) has long been treated as the physical definition of "pitch" in phonetic analysis. But there have been many demonstrations that F0 is at best an approximation to pitch, both in production and in perception: pitch…

Sound · Computer Science 2022-12-14 Danni Ma , Neville Ryant , Mark Liberman

Tracking the fundamental frequency (f0) of a monophonic instrumental performance is effectively a solved problem with several solutions achieving 99% accuracy. However, the related task of automatic music transcription requires a further…

Sound · Computer Science 2023-11-16 Xavier Riley , Simon Dixon

The identification of siren sounds in urban soundscapes is a crucial safety aspect for smart vehicles and has been widely addressed by means of neural networks that ensure robustness to both the diversity of siren signals and the strong and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-16 Stefano Damiano , Thomas Dietzen , Toon van Waterschoot

Despite much research, traditional methods to pitch prediction are still not perfect. With the emergence of neural networks (NNs), researchers hope to create a NN-based pitch predictor that outperforms traditional methods. Three pitch…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Anja Kroon

This paper addresses the extraction of multiple F0 values from polyphonic and a cappella vocal performances using convolutional neural networks (CNNs). We address the major challenges of ensemble singing, i.e., all melodic sources are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Helena Cuesta , Brian McFee , Emilia Gómez
‹ Prev 1 2 3 10 Next ›