English
Related papers

Related papers: Neural Network Alternatives to Convolutive Audio M…

200 papers

The recently proposed semi-blind source separation (SBSS) method for nonlinear acoustic echo cancellation (NAEC) outperforms adaptive NAEC in attenuating the nonlinear acoustic echo. However, the multiplicative transfer function (MTF)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-18 Guoliang Cheng , Lele Liao , Kai Chen , Yuxiang Hu , Changbao Zhu , Jing Lu

Representing symbolic music with compound tokens, where each token consists of several different sub-tokens representing a distinct musical feature or attribute, offers the advantage of reducing sequence length. While previous research has…

Sound · Computer Science 2026-03-17 HaeJun Yoo , Hao-Wen Dong , Jongmin Jung , Dasaem Jeong

Improving the interpretability of deep neural networks has recently gained increased attention, especially when the power of deep learning is leveraged to solve problems in physics. Interpretability helps us understand a model's ability to…

Sound · Computer Science 2023-10-12 Karim Helwani , Erfan Soltanmohammadi , Michael M. Goodwin

Data augmentation has proven to be a promising prospect in improving the performance of deep learning models by adding variability to training data. In previous work with developing a noise robust acoustic-to-articulatory speech inversion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-02 Yashish M. Siriwardena , Ahmed Adel Attia , Ganesh Sivaraman , Carol Espy-Wilson

A method for musical audio synthesis using autoencoding neural networks is proposed. The autoencoder is trained to compress and reconstruct magnitude short-time Fourier transform frames. The autoencoder produces a spectrogram by activating…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-29 Joseph Colonel , Christopher Curro , Sam Keene

The idea of adversarial learning of regularization functionals has recently been introduced in the wider context of inverse problems. The intuition behind this method is the realization that it is not only necessary to learn the basic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-04 Martin Ludvigsen , Markus Grasmair

In this paper, we suggest a new parallel, non-causal and shallow waveform domain architecture for speech enhancement based on FFTNet, a neural network for generating high quality audio waveform. In contrast to other waveform based…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-09 Muhammed PV Shifas , Nagaraj Adiga , Vassilis Tsiaras , Yannis Stylianou

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a…

Sound · Computer Science 2016-08-24 Syu-Siang Wang , Alan Chern , Yu Tsao , Jeih-weih Hung , Xugang Lu , Ying-Hui Lai , Borching Su

In this paper, we propose a dictionary update method for Nonnegative Matrix Factorization (NMF) with high dimensional data in a spectral conversion (SC) task. Voice conversion has been widely studied due to its potential applications such…

Machine Learning · Statistics 2016-10-14 Chin-Cheng Hsu , Hsin-Te Hwang , Yi-Chiao Wu , Yu Tsao , Hsin-Min Wang

Speech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-15 Chengyu Zheng , Xiulian Peng , Yuan Zhang , Sriram Srinivasan , Yan Lu

Traditional convolutional layers extract features from patches of data by applying a non-linearity on an affine function of the input. We propose a model that enhances this feature extraction process for the case of sequential data, by…

Machine Learning · Statistics 2017-07-21 Gil Keren , Björn Schuller

We show how to incorporate information from labeled examples into nonnegative matrix factorization (NMF), a popular unsupervised learning algorithm for dimensionality reduction. In addition to mapping the data into a space of lower…

Machine Learning · Computer Science 2011-12-19 Youngmin Cho , Lawrence K. Saul

Nonnegative matrix factorization (NMF) has become a widely used tool for the analysis of high-dimensional data as it automatically extracts sparse and meaningful features from a set of nonnegative data vectors. We first illustrate this…

Machine Learning · Statistics 2014-12-10 Nicolas Gillis

Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize…

Sound · Computer Science 2026-03-13 Hyung-Seok Oh , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

Non-negative matrix factorisation (NMF) is a widely used tool for unsupervised learning and feature extraction, with applications ranging from genomics to text analysis and signal processing. Standard formulations of NMF are typically…

Machine Learning · Computer Science 2026-03-11 Elisabeth Sommer James , Asger Hobolth , Marta Pelizzola

Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Jiachen Lian , Alan W Black , Yijing Lu , Louis Goldstein , Shinji Watanabe , Gopala K. Anumanchipalli

While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches have observed a lot of renewed interest, including the extended…

Sound · Computer Science 2025-08-20 Sarthak Yadav , Sergios Theodoridis , Zheng-Hua Tan

Nonnegative matrix factorization (NMF) seeks a low-rank approximation $X \approx UV^T$ with nonnegative factors and is commonly solved using interior methods that enforce feasibility throughout optimization. We show that such…

Machine Learning · Computer Science 2026-05-20 Qiujing Lu , Tonmoy Monsoor , Ehsan Ebrahimzadeh , Kartik Sharma , Vwani Roychowdhury

Distributed microphone arrays composed of multiple subarrays enable blind source separation over a wide spatial area. Directly applying fast multichannel nonnegative matrix factorization (FastMNMF) to all subarrays can exploit observations…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-20 Hirotaka Nishikori , Nobutaka Ito , Kouei Yamaoka , Norihiro Takamune , Hiroshi Saruwatari

Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-10 Zhuohuang Zhang , Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Dong Yu