中文
相关论文

相关论文: Wavelet-based spatial audio framework

200 篇论文

In this paper, we propose a new method for the construction of multi-dimensional, wavelet-like families of affine frames, commonly referred to as framelets, with specific directional characteristics, small and compact support in space,…

信息论 · 计算机科学 2019-09-13 Nikolaos Atreas , Nikolaos Karantzas , Manos Papadakis , Theodoros Stavropoulos

We find that the EPE evaluation metrics of RAFT-stereo converge inconsistently in the low and high frequency regions, resulting high frequency degradation (e.g., edges and thin objects) during the iterative process. The underlying reason…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Xiaobao Wei , Jiawei Liu , Dongbo Yang , Junda Cheng , Changyong Shu , Wei Wang

Spatial audio enhances immersion by reproducing 3D sound fields, with Ambisonics offering a scalable format for this purpose. While first-order Ambisonics (FOA) notably facilitates hardware-efficient acquisition and storage of sound fields…

音频与语音处理 · 电气工程与系统科学 2026-03-31 Amit Milstein , Nir Shlezinger , Boaz Rafaely

Sound in indoor spaces forms a complex wavefield due to multiple scattering encountered by the sound. Indoor acoustic communication involving multiple sources and receivers thus inevitably suffers from cross-talks. Here, we demonstrate the…

声音 · 计算机科学 2024-02-13 Hongkuan Zhang , Qiyuan Wang , Mathias Fink , Guancong Ma

This work develops a spherical-multipole expansion of Goldstein's acoustic analogy, for the prediction of tonal noise from rotating propellers. The acoustic field is expressed through spherical multipoles, which separate source integrals…

流体动力学 · 物理学 2026-03-20 Felice Fruncillo , Paolo Luchini , Flavio Giannetti

One of the biggest challenges of acoustic scene classification (ASC) is to find proper features to better represent and characterize environmental sounds. Environmental sounds generally involve more sound sources while exhibiting less…

声音 · 计算机科学 2019-04-11 Hongwei Song , Jiqing Han , Shiwen Deng

This work describes a novel data-driven latent space inference framework built on paired autoencoders to handle observational inconsistencies when solving inverse problems. Our approach uses two autoencoders, one for the parameter space and…

机器学习 · 计算机科学 2026-01-19 Emma Hart , Bas Peters , Julianne Chung , Matthias Chung

In this paper we address the problems of modeling the acoustic space generated by a full-spectrum sound source and of using the learned model for the localization and separation of multiple sources that simultaneously emit sparse-spectrum…

声音 · 计算机科学 2015-02-06 Antoine Deleforge , Florence Forbes , Radu Horaud

The present paper aims to complete an earlier paper where the acoustic world was introduced. This is accomplished by analyzing the interactions which occur between the inhomogeneities of the acoustic medium, which are induced by the…

综合物理 · 物理学 2016-12-02 Ion Simaciu , Gheorghe Dumitrescu , Zoltan Borsos , Mariana Bradac

This work proposes a neural network to extensively exploit spatial information for multichannel joint speech separation, denoising and dereverberation, named SpatialNet. In the short-time Fourier transform (STFT) domain, the proposed…

声音 · 计算机科学 2023-12-25 Changsheng Quan , Xiaofei Li

We present a new method for the analysis of images, a fundamental task in observational astronomy. It is based on the linear decomposition of each object in the image into a series of localised basis functions of different shapes, which we…

天体物理学 · 物理学 2008-11-26 Alexandre Refregier

In this work, we address the challenge of generalizable audio deepfake detection (ADD) across diverse speech synthesis paradigms-including conventional text-to-speech (TTS) systems and modern diffusion or flow-matching (FM) based…

音频与语音处理 · 电气工程与系统科学 2025-11-17 Farhan Sheth , Girish , Mohd Mujtaba Akhtar , Muskaan Singh

Modern autonomous systems are driving the critical need for next-generation adaptive materials and structures with embodied intelligence, i.e., the embodiment of memory, perception, learning, and decision-making within the mechanical…

应用物理 · 物理学 2025-11-18 Yuning Zhang , K. W. Wang

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation,…

音频与语音处理 · 电气工程与系统科学 2025-06-04 Huadai Liu , Tianyi Luo , Kaicheng Luo , Qikai Jiang , Peiwen Sun , Jialei Wang , Rongjie Huang , Qian Chen , Wen Wang , Xiangtai Li , Shiliang Zhang , Zhijie Yan , Zhou Zhao , Wei Xue

Many representation systems on the sphere have been proposed in the past, such as spherical harmonics, wavelets, or curvelets. Each of these data representations is designed to extract a specific set of features, and choosing the best fixed…

天体物理仪器与方法 · 物理学 2019-01-23 Florent Sureau , Felix Voigtlaender , Malte Wust , Jean-Luc Starck , Gitta Kutyniok

Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZen leverage these rich single-channel representations, but…

音频与语音处理 · 电气工程与系统科学 2026-01-06 Marc Deegen , Tobias Gburrek , Tobias Cord-Landwehr , Thilo von Neumann , Jiangyu Han , Lukáš Burget , Reinhold Haeb-Umbach

Sound field reconstruction (SFR) augments the information of a sound field captured by a microphone array. Conventional SFR methods using basis function decomposition are straightforward and computationally efficient, but may require more…

音频与语音处理 · 电气工程与系统科学 2024-02-15 Fei Ma , Sipei Zhao , Ian S. Burnett

Model calibration and debiasing are fundamental yet operationally expensive challenges in large-scale recommendation systems. Existing approaches treat them as separate problems requiring distinct infrastructure: post-hoc calibration…

信息检索 · 计算机科学 2026-04-28 Hailing Cheng , Yafang Yang , Hemeng Tao , Fengyu Zhang

First-order Ambisonics (FOA) is a standard spatial audio format based on spherical harmonic decomposition. Its zeroth- and first-order components capture the sound pressure and particle velocity, respectively. Recently, physics-informed…

声音 · 计算机科学 2026-03-25 Yoshiki Masuyama , Francois G. Germain , Gordon Wichern , Chiori Hori , Jonathan Le Roux

In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Zheqi Dai , Guangyan Zhang , Haolin He , Xiquan Li , Jingyu Li , Chunyat Wu , Yiwen Guo , Qiuqiang Kong