中文
相关论文

相关论文: Wavelet-based spatial audio framework

200 篇论文

The acoustic wave-propagation without mean flow and heat flux can be described in terms of velocity and pressure by the compressible nonlinear Navier-Stokes equations, where boundary layers appear at walls due to the viscosity and a…

偏微分方程分析 · 数学 2017-01-10 Anastasia Thoens-Zueva , Kersten Schmidt , Adrien Semin

Recently it has been established that asymptotic incoherence can be used to facilitate subsampling, in order to optimize reconstruction quality, in a variety of continuous compressed sensing problems, and the coherence structure of certain…

信息论 · 计算机科学 2016-10-25 Alex Jones , Ben Adcock , Anders Hansen

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into…

This paper presents SHTNet, a lightweight spherical harmonic transform (SHT) based framework, which is designed to address cross-array generalization challenges in multi-channel automatic speech recognition (ASR) through three key…

音频与语音处理 · 电气工程与系统科学 2025-10-22 Xiangzhu Kong , Huang Hao , Zhijian Ou

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Magnonics is a field of science that addresses the physical properties of spin waves and utilizes them for data processing. Scalability down to atomic dimensions, operations in the GHz-to-THz frequency range, utilization of nonlinear and…

应用物理 · 物理学 2023-11-01 A. V. Chumak , P. Kabos , M. Wu , C. Abert , C. Adelmann , A. Adeyeye , J. Åkerman , F. G. Aliev , A. Anane , A. Awad , C. H. Back , A. Barman , G. E. W. Bauer , M. Becherer , E. N. Beginin , V. A. S. V. Bittencourt , Y. M. Blanter , P. Bortolotti , I. Boventer , D. A. Bozhko , S. A. Bunyaev , J. J. Carmiggelt , R. R. Cheenikundil , F. Ciubotaru , S. Cotofana , G. Csaba , O. V. Dobrovolskiy , C. Dubs , M. Elyasi , K. G. Fripp , H. Fulara , I. A. Golovchanskiy , C. Gonzalez-Ballestero , P. Graczyk , D. Grundler , P. Gruszecki , G. Gubbiotti , K. Guslienko , A. Haldar , S. Hamdioui , R. Hertel , B. Hillebrands , T. Hioki , A. Houshang , C. -M. Hu , H. Huebl , M. Huth , E. Iacocca , M. B. Jungfleisch , G. N. Kakazei , A. Khitun , R. Khymyn , T. Kikkawa , M. Kläui , O. Klein , J. W. Kłos , S. Knauer , S. Koraltan , M. Kostylev , M. Krawczyk , I. N. Krivorotov , V. V. Kruglyak , D. Lachance-Quirion , S. Ladak , R. Lebrun , Y. Li , M. Lindner , R. Macêdo , S. Mayr , G. A. Melkov , S. Mieszczak , Y. Nakamura , H. T. Nembach , A. A. Nikitin , S. A. Nikitov , V. Novosad , J. A. Otalora , Y. Otani , A. Papp , B. Pigeau , P. Pirro , W. Porod , F. Porrati , H. Qin , B. Rana , T. Reimann , F. Riente , O. Romero-Isart , A. Ross , A. V. Sadovnikov , A. R. Safin , E. Saitoh , G. Schmidt , H. Schultheiss , K. Schultheiss , A. A. Serga , S. Sharma , J. M. Shaw , D. Suess , O. Surzhenko , K. Szulc , T. Taniguchi , M. Urbánek , K. Usami , A. B. Ustinov , T. van der Sar , S. van Dijken , V. I. Vasyuchka , R. Verba , S. Viola Kusminskiy , Q. Wang , M. Weides , M. Weiler , S. Wintz , S. P. Wolski , X. Zhang

The paper presents a versatile library of analytic and quasi-analytic complex-valued wavelet packets (WPs) which originate from discrete splines of arbitrary orders. The real parts of the quasi-analytic WPs are the regular spline-based…

数值分析 · 数学 2019-07-04 Amir Averbuch , Pekka Neittaanmaki , Valery Zheludev

Spatial audio methods are gaining a growing interest due to the spread of immersive audio experiences and applications, such as virtual and augmented reality. For these purposes, 3D audio signals are often acquired through arrays of…

音频与语音处理 · 电气工程与系统科学 2022-12-16 Eleonora Grassucci , Gioia Mancini , Christian Brignone , Aurelio Uncini , Danilo Comminiello

We introduce a framework for designing multi-scale, adaptive, shift-invariant frames and bi-frames for representing signals. The new framework, called AdaFrame, improves over dictionary learning-based techniques in terms of computational…

计算机视觉与模式识别 · 计算机科学 2015-07-20 Cheng Tai , Weinan E

Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths. Both endpoints, however, concentrate in thin spherical shells, and a Euclidean chord leaves those shells even…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Tuna Han Salih Meral , Kaan Oktay , Hidir Yesiltepe , Adil Kaan Akan , Pinar Yanardag

We present an indoor acoustic simulation framework that supports both ultrasonic and audible signaling. The framework opens the opportunity for fast indoor acoustic data generation and positioning development. The improved…

音频与语音处理 · 电气工程与系统科学 2023-06-22 Daan Delabie , Chesney Buyle , Bert Cox , Liesbet Van der Perre , Lieven De Strycker

A new construction of a directional continuous wavelet analysis on the sphere is derived herein. We adopt the harmonic scaling idea for the spherical dilation operator recently proposed by Sanz et al. but extend the analysis to a more…

天体物理学 · 物理学 2011-10-28 J. D. McEwen , M. P. Hobson , A. N. Lasenby

We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by…

Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving…

声音 · 计算机科学 2025-09-15 Marco Pasini , Stefan Lattner , George Fazekas

Optoacoustic image formation is conventionally based upon ultrasound time-of-flight readings from multiple detection positions. Herein, we exploit acoustic scattering to physically encode the position of optical absorbers in the acquired…

生物物理 · 物理学 2019-10-30 Xose Luis Dean-Ben , Ali Ozbek , Hernan Lopez-Schier , Daniel Razansky

Slow sound is a frequently exploited phenomenon that metamaterials can induce in order to permit wave energy compression, redirection, imaging, sound absorption and other special functionalities. Generally however such slow sound structures…

Characterizing sound field diffuseness has many practical applications, from room acoustics analysis to speech enhancement and sound field reproduction. In this paper we investigate how spherical microphone arrays (SMAs) can be used to…

声音 · 计算机科学 2016-07-04 Nicolas Epain , Craig T. Jin

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years.…

声音 · 计算机科学 2015-10-08 Bo Fan , Siu Wa Lee , Xiaohai Tian , Lei Xie , Minghui Dong

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice…

音频与语音处理 · 电气工程与系统科学 2018-10-02 Kou Tanaka , Takuhiro Kaneko , Nobukatsu Hojo , Hirokazu Kameoka