English
Related papers

Related papers: Lightweight and interpretable neural modeling of a…

200 papers

Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To harden models against…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-19 Davis Liang , Zhiheng Huang , Zachary C. Lipton

A cascaded speech translation model relies on discrete and non-differentiable transcription, which provides a supervision signal from the source side and helps the transformation between source speech and target text. Such modeling suffers…

Computation and Language · Computer Science 2020-11-25 Parnia Bahar , Tobias Bieschke , Ralf Schlüter , Hermann Ney

We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is achieved by combining two distinct models. The first model,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-13 Kevin Kilgour , Beat Gfeller , Qingqing Huang , Aren Jansen , Scott Wisdom , Marco Tagliasacchi

End-to-end trainable models have reached the performance of traditional handcrafted compression techniques on videos and images. Since the parameters of these models are learned over large training sets, they are not optimal for any given…

Image and Video Processing · Electrical Eng. & Systems 2022-10-12 Oussama Jourairi , Muhammet Balcilar , Anne Lambert , François Schnitzler

Deep learning approaches for black-box modelling of audio effects have shown promise, however, the majority of existing work focuses on nonlinear effects with behaviour on relatively short time-scales, such as guitar amplifiers and…

Sound · Computer Science 2023-05-11 Marco Comunità , Christian J. Steinmetz , Huy Phan , Joshua D. Reiss

Selecting appropriate inductive biases is an essential step in the design of machine learning models, especially when working with audio, where even short clips may contain millions of samples. To this end, we propose the combolutional…

Sound · Computer Science 2025-08-06 Cameron Churchwell , Minje Kim , Paris Smaragdis

In this study, we introduce a method for estimating sound fields in reverberant environments using a conditional invertible neural network (CINN). Sound field reconstruction can be hindered by experimental errors, limited spatial data,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Xenofon Karakonstantis , Efren Fernandez-Grande , Peter Gerstoft

The mapping from sound to neural activity that underlies hearing is highly non-linear. The first few stages of this mapping in the cochlea have been modelled successfully, with biophysical models built by hand and, more recently, with DNN…

Neurons and Cognition · Quantitative Biology 2026-01-23 Lloyd Pellatt , Fotios Drakopoulos , Shievanie Sabesan , Nicholas A. Lesica

Several approaches have been developed to capture the complexity and nonlinearity of human growth. One widely used is the Super Imposition by Translation and Rotation (SITAR) model, which has become popular in studies of adolescent growth.…

Machine Learning · Statistics 2025-05-15 María Alejandra Hernández , Oscar Rodriguez , Dae-Jin Lee

Neural networks have become ubiquitous with guitar distortion effects modelling in recent years. Despite their ability to yield perceptually convincing models, they are susceptible to frequency aliasing when driven by high frequency and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-19 Alistair Carson , Alec Wright , Stefan Bilbao

Recently, the information content (IC) of predictions from a Generative Infinite-Vocabulary Transformer (GIVT) has been used to model musical expectancy and surprisal in audio. We investigate the effectiveness of such modelling using IC…

Sound · Computer Science 2025-08-08 Mathias Rose Bjare , Stefan Lattner , Gerhard Widmer

Applications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relatively unstructured and variable. To address these…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-21 Christian J. Steinmetz , Jordi Pons , Santiago Pascual , Joan Serrà

Instrumental variable (IV) regression is a standard strategy for learning causal relationships between confounded treatment and outcome variables from observational data by utilizing an instrumental variable, which affects the outcome only…

Machine Learning · Computer Science 2023-06-28 Liyuan Xu , Yutian Chen , Siddarth Srinivasan , Nando de Freitas , Arnaud Doucet , Arthur Gretton

We propose an audio effects processing framework that learns to emulate a target electric guitar tone from a recording. We train a deep neural network using an adversarial approach, with the goal of transforming the timbre of a guitar, into…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-21 Alec Wright , Vesa Välimäki , Lauri Juvela

Standard fine-tuning of pre-trained audio models couples representation learning with classifier training, which can obscure the true quality of the learned representations. In this work, we advocate for a disentangled two-stage framework…

Sound · Computer Science 2025-09-23 Yang Wang , Qibin Liang , Chenghao Xiao , Yizhi Li , Noura Al Moubayed , Chenghua Lin

Designing high-performance substrate-integrated waveguide (SIW) filters with both closely spaced and widely separated resonances is challenging. Consequently, there is a growing need for robust methods that reduce reliance on time-consuming…

Machine Learning · Computer Science 2025-10-21 Mohammad Mashayekhi , Kamran Salehian , Abbas Ozgoli , Saeed Abdollahi , Abdolali Abdipour , Ahmed A. Kishk

Applications of deep learning for audio effects often focus on modeling analog effects or learning to control effects to emulate a trained audio engineer. However, deep learning approaches also have the potential to expand creativity…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-07 Christian J. Steinmetz , Joshua D. Reiss

In all applications in digital communications, it is crucial for an estimator to be unbiased. Although so-called soft feedback is widely employed in many different fields of engineering, typically the biased estimate is used. In this paper,…

Information Theory · Computer Science 2018-02-21 Susanne Sparrer , Robert F. H. Fischer

This paper addresses the performance of bit-interleaved coded multiple beamforming (BICMB) [1], [2] with imperfect knowledge of beamforming vectors. Most studies for limited-rate channel state information at the transmitter (CSIT) assume…

Information Theory · Computer Science 2008-05-08 Ersin Sengul , Hong Ju Park , Ender Ayanoglu

Photoacoustic computed tomography (PACT) is a promising imaging modality that combines the advantages of optical contrast with ultrasound detection. Utilizing ultrasound transducers with larger surface areas can improve detection…