English
Related papers

Related papers: Manifold learning-supported estimation of relative…

200 papers

In this work, we propose a new recurrent autoencoder architecture, termed Feedback Recurrent AutoEncoder (FRAE), for online compression of sequential data with temporal dependency. The recurrent structure of FRAE is designed to efficiently…

Machine Learning · Computer Science 2020-02-18 Yang Yang , Guillaume Sautière , J. Jon Ryu , Taco S Cohen

This paper investigates the viability of Wave Field Synthesis (WFS) for enhancing auditory immersion in VR-based cognitive research. While Virtual Reality (VR) offers significant advantages for studying human perception and behavior,…

Human-Computer Interaction · Computer Science 2025-07-08 Benjamin Kahl

In this research, we present an interface based on Variational Autoencoders trained with a wide range of natural sounds for the innovative creation of Foley effects. The model can transfer new sound features to prerecorded audio or…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-25 Mateo Cámara , José Luis Blanco

Acoustic scene classification and related tasks have been dominated by Convolutional Neural Networks (CNNs). Top-performing CNNs use mainly audio spectograms as input and borrow their architectural design primarily from computer vision. A…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-09 Khaled Koutini , Hamid Eghbal-zadeh , Gerhard Widmer

Radio tomographic imaging (RTI) is an emerging technology for localization of physical objects in a geographical area covered by wireless networks. With attenuation measurements collected at spatially distributed sensors, RTI capitalizes on…

Signal Processing · Electrical Eng. & Systems 2020-08-26 Donghoon Lee , Georgios B. Giannakis

We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-16 Tyler Vuong , Yangyang Xia , Richard M. Stern

Previous research has shown that the principal singular vectors of a pre-trained model's weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-10 Zhe Li , Man-wai Mak , Mert Pilanci , Hung-yi Lee , Helen Meng

Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. Such embeddings can form the basis for speech search, indexing and discovery systems when conventional speech recognition is not possible. In…

Computation and Language · Computer Science 2021-02-08 Herman Kamper , Yevgen Matusevych , Sharon Goldwater

Recent advancements in learning Discrete Representations as opposed to continuous ones have led to state of art results in tasks that involve Language, Audio and Vision. Some latent factors such as words, phonemes and shapes are better…

Machine Learning · Computer Science 2020-04-14 Iordanis Fostiropoulos

Given an input sound signal and a target virtual sound source, sound spatialisation algorithms manipulate the signal so that a listener perceives it as though it were emitted from the target source. There exist several established…

Sound · Computer Science 2017-11-28 Ali Tarzan , Marco Alunno , Paolo Bientinesi

Recognition of speech, and in particular the ability to generalize and learn from small sets of labelled examples like humans do, depends on an appropriate representation of the acoustic input. We formulate the problem of finding robust…

New system for i-vector speaker recognition based on variational autoencoder (VAE) is investigated. VAE is a promising approach for developing accurate deep nonlinear generative models of complex data. Experiments show that VAE provides…

Sound · Computer Science 2017-05-26 Timur Pekhovsky , Maxim Korenevsky

Interpretability of learning algorithms is crucial for applications involving critical decisions, and variable importance is one of the main interpretation tools. Shapley effects are now widely used to interpret both tree ensembles and…

Machine Learning · Statistics 2022-02-03 Clément Bénard , Gérard Biau , Sébastien da Veiga , Erwan Scornet

The Modulation Transfer Function (MTF) is an important image quality metric typically used in the automotive domain. However, despite the fact that optical quality has an impact on the performance of computer vision in vehicle automation,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Daniel Jakab , Eoin Martino Grua , Brian Micheal Deegan , Anthony Scanlan , Pepijn Van De Ven , Ciarán Eising

Acoustic velocity vectors (AVVs) are related to the human's perception of sound at low frequencies and are widely used in Ambisonics. This paper proposes a spatial sound field reproduction algorithm called velocity matching, which…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Jiarui Wang , Thushara Abhayapala , Jihui Aimee Zhang , Prasanga Samarasinghe

Tone Transfer is a novel deep-learning technique for interfacing a sound source with a synthesizer, transforming the timbre of audio excerpts while keeping their musical form content. Due to its good audio quality results and continuous…

Sound · Computer Science 2023-10-10 Franco Caspe , Andrew McPherson , Mark Sandler

The modulation transfer function (MTF) is widely used to characterise the performance of optical systems. Measuring it is costly and it is thus rarely available for a given lens specimen. Instead, MTFs based on simulations or, at best, MTFs…

Computer Vision and Pattern Recognition · Computer Science 2018-05-07 Matthias Bauer , Valentin Volchkov , Michael Hirsch , Bernhard Schölkopf

The spatial information of sound plays a crucial role in various situations, ranging from daily activities to advanced engineering technologies. To fully utilize its potential, numerous research studies on spatial audio signal processing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-14 Natsuki Ueno , Shoichi Koyama

Value-function (VF) approximation is a central problem in Reinforcement Learning (RL). Classical non-parametric VF estimation suffers from the curse of dimensionality. As a result, parsimonious parametric models have been adopted to…

Machine Learning · Computer Science 2024-05-29 Sergio Rozada , Santiago Paternain , Antonio G. Marques

Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are…

Sound · Computer Science 2025-02-20 Pavol Harar , Roswitha Bammer , Anna Breger , Monika Dörfler , Zdenek Smekal