中文
相关论文

相关论文: Improving Neural Pitch Estimation with SWIPE Kerne…

200 篇论文

In the domain of music and sound processing, pitch extraction plays a pivotal role. Our research presents a specialized convolutional neural network designed for pitch extraction, particularly from the human singing voice in acapella…

声音 · 计算机科学 2023-12-19 Jeremy Cochoy

Transferability estimation has emerged as an important problem in transfer learning. A transferability estimation method takes as inputs a set of pre-trained models and decides which pre-trained model can deliver the best transfer learning…

机器学习 · 计算机科学 2024-05-06 Yunhui Guo

Deep neural networks can learn complex and abstract representations, that are progressively obtained by combining simpler ones. A recent trend in speech and speaker recognition consists in discovering these representations starting from raw…

音频与语音处理 · 电气工程与系统科学 2019-02-26 Mirco Ravanelli , Yoshua Bengio

Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Umut Isik , Ritwik Giri , Neerad Phansalkar , Jean-Marc Valin , Karim Helwani , Arvindh Krishnaswamy

Modern medical image segmentation methods primarily use discrete representations in the form of rasterized masks to learn features and generate predictions. Although effective, this paradigm is spatially inflexible, scales poorly to…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Yejia Zhang , Pengfei Gu , Nishchal Sapkota , Danny Z. Chen

A previous signal processing algorithm that aimed to enhance spectral changes (SCE) over time showed benefit for hearing-impaired (HI) listeners to recognize speech in background noise. In this work, the previous SCE was manipulated to…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Xiang Li , Xin Tian , Henry Luo , Jinyu Qian , Xihong Wu , Dingsheng Luo , Jing Chen

This paper presents a polyphonic pitch tracking system able to extract both framewise and note-based estimates from audio. The system uses several artificial neural networks in a deep layered learning setup. First, cascading networks are…

声音 · 计算机科学 2019-03-19 Anders Elowsson

Recent strides in low-latency spiking neural network (SNN) algorithms have drawn significant interest, particularly due to their event-driven computing nature and fast inference capability. One of the most efficient ways to construct a…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Chen Li , Bipin Rajendran

This paper investigates the deep learning based approaches for simultaneous wireless information and power transfer (SWIPT). The quality-of-service (QoS) constrained sum-rate maximization problems are, respectively, formulated for…

信号处理 · 电气工程与系统科学 2025-02-07 Hong Han , Yang Lu , Zihan Song , Ruichen Zhang , Wei Chen , Bo Ai , Dusit Niyato , Dong In Kim

We identify and address a fundamental limitation of sinusoidal representation networks (SIRENs), a class of implicit neural representations. SIRENs Sitzmann et al. (2020), when not initialized appropriately, can struggle at fitting signals…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Hemanth Chandravamsi , Dhanush V. Shenoy , Steven H. Frankel

Spiking Neural Networks (SNN) are more closely related to brain-like computation and inspire hardware implementation. This is enabled by small networks that give high performance on standard classification problems. In literature, typical…

神经与进化计算 · 计算机科学 2016-12-08 Anmol Biswas , Sidharth Prasad , Sandip Lashkare , Udayan Ganguly

Accurate wireless channel estimation is critical for next-generation wireless systems, enabling precise precoding for effective user separation, reduced interference across cells, and high-resolution sensing, among other benefits.…

信号处理 · 电气工程与系统科学 2026-03-03 Alireza Javid , Nuria González-Prelcic

We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the generated speech can be…

音频与语音处理 · 电气工程与系统科学 2021-02-17 Adrian Łańcucki

We present a novel multi-channel front-end based on channel shortening with theWeighted Prediction Error (WPE) method followed by a fixed MVDR beamformer used in combination with a recently proposed self-attention-based channel combination…

音频与语音处理 · 电气工程与系统科学 2022-03-29 Dushyant Sharma , Rong Gong , James Fosburgh , Stanislav Yu. Kruchinin , Patrick A. Naylor , Ljubomir Milanovic

A novel framework, called InterGridNet, is introduced, leveraging a shallow RawNet model for geolocation classification of Electric Network Frequency (ENF) signatures in the SP Cup 2016 dataset. During data preparation, recordings are…

The detection and estimation of sinusoids is a fundamental signal processing task for many applications related to sensing and communications. While algorithms have been proposed for this setting, quantization is a critical, but often…

信号处理 · 电气工程与系统科学 2022-10-05 Ryan Dreifuerst , Robert W. Heath

We present a general framework for training spiking neural networks (SNNs) to perform binary classification on multivariate time series, with a focus on step-wise prediction and high precision at low false alarm rates. The approach uses the…

机器学习 · 计算机科学 2025-11-24 James Ghawaly , Andrew Nicholson , Catherine Schuman , Dalton Diez , Aaron Young , Brett Witherspoon

Jitter and shimmer measurements have shown to be carriers of voice quality and prosodic information which enhance the performance of tasks like speaker recognition, diarization or automatic speech recognition (ASR). However, such features…

计算与语言 · 计算机科学 2021-12-22 Guillermo Cámbara , Jordi Luque , Mireia Farrús

Vehicular communication systems face significant challenges due to high mobility and rapidly changing environments, which affect the channel over which the signals travel. To address these challenges, neural network (NN)-based channel…

机器学习 · 计算机科学 2025-02-12 Simbarashe Aldrin Ngorima , Albert Helberg , Marelie H. Davel

We study large-scale kernel methods for acoustic modeling and compare to DNNs on performance metrics related to both acoustic modeling and recognition. Measuring perplexity and frame-level classification accuracy, kernel-based acoustic…