English
Related papers

Related papers: Analytical Exploration of Spatial Audio Cues: A Di…

200 papers

Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations that are difficult to interpret. Inspired by…

In an ultrasonic array system, increasing the aperture size to achieve a high resolution requires more transmit and receive channels, thus making it essential to have an analysis technique that can reconstruct the shape and physical…

Optics · Physics 2025-08-21 Kai Yabumoto , Takayoshi Yumii , Kenjiro Kimura

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Xudong Xu , Dejan Markovic , Jacob Sandakly , Todd Keebler , Steven Krenn , Alexander Richard

In these last years, many studies have focalized on the design of reliable underwater acoustic communication systems. However, the ocean acoustic communication channel exhibits strong amplitude and phase fluctuations and the phenomena of…

Signal Processing · Electrical Eng. & Systems 2019-07-16 Yasin Yousif Al-Aboosi , Hussein A. Abdulnabi

Estimating scattering parameters of heterogeneous media from images is a severely under-constrained and challenging problem. Most of the existing approaches model BSSRDF either through an analysis-by-synthesis approach, approximating…

With the recent advancements of data driven approaches using deep neural networks, music source separation has been formulated as an instrument-specific supervised problem. While existing deep learning models implicitly absorb the spatial…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-16 Darius Petermann , Minje Kim

This investigation is concerned with the 2D acoustic scattering problem of a plane wave propagating in a non-lossy fluid host and soliciting a linear, isotropic, macroscopically-homogeneous, lossy, flat-plane layer in which the mass density…

Applied Physics · Physics 2019-05-01 Armand Wirgin

Multi-channel multi-talker speech recognition presents formidable challenges in the realm of speech processing, marked by issues such as background noise, reverberation, and overlapping speech. Overcoming these complexities requires…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-09 Yiwen Shao

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Radio frequency (RF) signals have been proved to be flexible for human silhouette segmentation (HSS) under complex environments. Existing studies are mainly based on a one-shot approach, which lacks a coherent projection ability from the RF…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Penghui Wen , Kun Hu , Dong Yuan , Zhiyuan Ning , Changyang Li , Zhiyong Wang

Time-frequency scattering is a mathematical transformation of sound waves. Its core purpose is to mimick the way the human auditory system extracts information from its environment. In the context of improving the artificial intelligence of…

Sound · Computer Science 2019-05-22 Vincent Lostanlen

Head Related Transfer Functions (HRTFs) play a crucial role in creating immersive spatial audio experiences. However, HRTFs differ significantly from person to person, and traditional methods for estimating personalized HRTFs are expensive,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-08 Vivek Jayaram , Ira Kemelmacher-Shlizerman , Steven M. Seitz

This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Gyeong-Tae Lee , Hyeonuk Nam , Yong-Hwa Park

While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this…

Sound · Computer Science 2026-03-30 Harunori Kawano , Takeshi Sasaki

We propose and demonstrate a generative deep learning approach for the shape recognition of an arbitrary object from its acoustic scattering properties. The strategy exploits deep neural networks to learn the mapping between the latent…

Sound · Computer Science 2022-07-13 W. W. Ahmed , M. Farhat , P. -Y. Chen , X. Zhang , Y. Wu

We develop a new scattering-based framework for the holographic encryption of analog and digital signals. The proposed methodology, termed "differential sensing", involves encryption of a wavefield image by means of two hard-to-guess,…

Optics · Physics 2024-09-10 Mohammadrasoul Taghavi , Edwin A. Marengo

In their everyday life, the speech recognition performance of human listeners is influenced by diverse factors, such as the acoustic environment, the talker and listener positions, possibly impaired hearing, and optional hearing devices.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-02 Marc René Schädler

In this paper, we consider acoustic or electromagnetic scattering in two dimensions from an infinite three-layer medium with thousands of wavelength-size dielectric particles embedded in the middle layer. Such geometries are typical of…

Numerical Analysis · Mathematics 2015-06-22 Jun Lai , Motoki Kobayashi , Leslie Greengard

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

3D reconstruction techniques such as LiDAR scanning and photogrammetry have made it practical to build detailed geometric models of real-world environments. Such reconstructed models can potentially serve as the foundation for wireless…

Networking and Internet Architecture · Computer Science 2026-05-27 Haofan Lu , Yadi Cao , Wanghao Yi , Omid Abari