Related papers: Perceptual evaluation of listener envelopment usin…
A recently developed technique known as analogue transformation acoustics has allowed the extension of the transformational paradigm to general spacetime transformations under which the acoustic equations are not form invariant. In this…
The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…
Spatial attributes of room acoustics have been widely studied using microphone and loudspeaker arrays. However, systems that combine both arrays, referred to as multiple-input multiple-output (MIMO) systems, have only been studied to a…
Acoustic reverberation is one of the most relevant factors that hampers the localization of a sound source inside a room. To date, several approaches have been proposed to deal with it, but have not always been evaluated under realistic…
This article presents an interactive system for stage acoustics experimentation including considerations for hearing one's own and others' instruments. The quality of real-time auralization systems for psychophysical experiments on music…
Speech enhancement aims to improve the perceptual quality of the speech signal by suppression of the background noise. However, excessive suppression may lead to speech distortion and speaker information loss, which degrades the performance…
Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo…
In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…
While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present…
We show that intermittency of noiselike emission, after propagation through a scattering medium, affects the distribution of noise in the observed correlation function. Intermittency also affects correlation of noise among channels of the…
We describe a new method for estimating the direction of sound in a reverberant environment from basic principles of sound propagation. The method utilizes SNR-adaptive features from time-delay and energy of the directional components after…
Zero-shot speaker adaptation aims to clone an unseen speaker's voice without any adaptation time and parameters. Previous researches usually use a speaker encoder to extract a global fixed speaker embedding from reference speech, and…
The bubbles involved in sonochemistry and other applications of cavitation oscillate inertially. A correct estimation of the wave attenuation in such bubbly media requires a realistic estimation of the power dissipated by the oscillation of…
Recently, we witnessed a tremendous effort to conquer the realm of acoustics as a possible playground to test with sound waves topologically protected wave propagation. Acoustics differ substantially from photonic and electronic systems…
Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…
Acoustic streaming is an ubiquitous phenomenon resulting from time-averaged nonlinear dynamics in oscillating fluids. In this theoretical study, we show that acoustic streaming can be suppressed by two orders of magnitude in major regions…
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more identities is expensive, challenging, and often limited by…
Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating…
A wave propagating through a scattering medium typically yields a complex temporal field distribution. Over the years, a number of procedures have emerged to shape the temporal profile of the field in order to temporally focus its energy on…
Any audio recording encapsulates the unique fingerprint of the associated acoustic environment, namely the background noise and reverberation. Considering the scenario of a room equipped with a fixed smart speaker device with one or more…