Related papers: On Time Delay Interpolation for Improved Acoustic …
Sound source localization (SSL) demonstrates remarkable results in controlled settings but struggles in real-world deployment due to dual imbalance challenges: intra-task imbalance arising from long-tailed direction-of-arrival (DoA)…
Stereo matching provides depth estimation from binocular images for downstream applications. These applications mostly take video streams as input and require temporally consistent depth maps. However, existing methods mainly focus on the…
We introduce a novel direct calibration algorithm to address phase delay, gain, and offset mismatches in Analog-to-Digital Converter (ADC) time interleaving systems. These mismatches, common in high-speed data acquisition, degrade system…
Spatial and temporal delays in a wireless multi-antenna system, paired with an orthogonal frequency division multiplexing (OFDM) waveform, can be utilized to estimate the Angle of Arrival (AoA) and Time of Arrival (ToA) of scatterers in the…
Acoustic local positioning systems (ALPSs) are an interesting alternative for indoor positioning due to certain advantages over other approaches, including their relatively high accuracy, low cost, and room-level signal propagation.…
Speaker identification typically involves three stages. First, a front-end speaker embedding model is trained to embed utterance and speaker profiles. Second, a scoring function is applied between a runtime utterance and each speaker…
Dispersion scan is a self-referenced measurement technique for ultrashort pulses. Similar to frequency-resolved optical gating, the dispersion scan technique records the dependence of nonlinearly generated spectra as a function of a…
In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and…
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more identities is expensive, challenging, and often limited by…
Distributed antenna arrays have been proposed for many applications ranging from space-based observatories to automated vehicles. Achieving good performance in distributed antenna systems requires stringent synchronization at the wavelength…
Multiple wireless sensing tasks, e.g., radar detection for driver safety, involve estimating the "channel" or relationship between signal transmitted and received. In this work, we focus on a certain channel model known as the delay-doppler…
Computational time reversal imaging can be used to locate the position of multiple scatterers in a known background medium. Here, we discuss a sparse approximation method for computational time-reversal imaging. The method is formulated…
In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have…
Different machines can exhibit diverse frequency patterns in their emitted sound. This feature has been recently explored in anomaly sound detection and reached state-of-the-art performance. However, existing methods rely on the manual or…
The Wigner-Smith (WS) time delay matrix relates a lossless system's scattering matrix to its frequency derivative. First proposed in the realm of quantum mechanics to characterize time delays experienced by particles during a collision,…
Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only…
Time delay interferometry (TDI) is a key technique employed in gravitational wave (GW) space missions to mitigate laser frequency noise by combining multiple laser links and establishing an equivalent equal arm interferometry. The null…
This article presents a new approach for the wireless clock synchronization of Decawave ultra-wideband transceivers based on the time difference of arrival. The presented techniques combine the time-of-arrival and time-difference-of-arrival…
Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…
Binaural target sound extraction (TSE) aims to extract a desired sound from a binaural mixture of arbitrary sounds while preserving the spatial cues of the desired sound. Indeed, for many applications, the target sound signal and its…