中文
相关论文

相关论文: FOA Tokenizer: Low-bitrate Neural Codec for First …

200 篇论文

Consider a microphone array, such as those present in Amazon Echos, conference phones, or self-driving cars. One of the goals of these arrays is to decode the angles in which acoustic signals arrive at them. This paper considers the problem…

声音 · 计算机科学 2021-09-28 Yu-Lin Wei , Romit Roy Choudhury

Selective fixed-filter active noise control (SFANC) is a novel approach capable of mitigating noise with varying frequency characteristics. It offers faster response and greater computational efficiency compared to traditional adaptive…

声音 · 计算机科学 2026-01-13 Boxiang Wang , Zhengding Luo , Haowen Li , Dongyuan Shi , Junwei Ji , Ziyi Yang , Woon-Seng Gan

While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice…

声音 · 计算机科学 2025-09-22 Ryan Collette , Ross Greenwood , Serena Nicoll

The feedback capacity of the stationary Gaussian additive noise channel has been open, except for the case where the noise is white. Here we find the feedback capacity of the stationary first-order moving average additive Gaussian noise…

信息论 · 计算机科学 2007-07-16 Young-Han Kim

While video-to-audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention to the spatial perception and immersive quality of the…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Yanan Wang , Linjie Ren , Zihao Li , Junyi Wang , Tian Gan

This paper presents the 3D soundfield synthesis of the pressure field radiated by directional acoustic sources using both the multimodal method and higher-order ambisonics (HOA). Ambisonics is a technique for encoding and reproducing…

经典物理 · 物理学 2024-09-20 Philippe Thorner , Eric Bavu , Jean-Baptiste Doc , Christophe Langrenne

Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model…

声音 · 计算机科学 2022-10-28 Kwanghee Choi , Eun Jung Yeo

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from…

音频与语音处理 · 电气工程与系统科学 2025-02-28 Yidi Jiang , Qian Chen , Shengpeng Ji , Yu Xi , Wen Wang , Chong Zhang , Xianghu Yue , ShiLiang Zhang , Haizhou Li

Automatic pronunciation assessment (APA) manages to quantify the pronunciation proficiency of a second language (L2) learner in a language. Prevailing approaches to APA normally leverage neural models trained with a regression loss…

音频与语音处理 · 电气工程与系统科学 2023-10-05 Bi-Cheng Yan , Hsin-Wei Wang , Yi-Cheng Wang , Jiun-Ting Li , Chi-Han Lin , Berlin Chen

Sound field decomposition predicts waveforms in arbitrary directions using signals from a limited number of microphones as inputs. Sound field decomposition is fundamental to downstream tasks, including source localization, source…

声音 · 计算机科学 2022-10-25 Qiuqiang Kong , Shilei Liu , Junjie Shi , Xuzhou Ye , Yin Cao , Qiaoxi Zhu , Yong Xu , Yuxuan Wang

Off-road semantic segmentation suffers from thick, inconsistent boundaries, sparse supervision for rare classes, and pervasive label noise. Designs that fuse only at low resolution blur edges and propagate local errors, whereas maintaining…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Seongkyu Choi , Jhonghyun An

Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however,…

声音 · 计算机科学 2026-01-29 Luca Cerovaz , Michele Mancusi , Emanuele Rodolà

Most widely-used modern audio codecs, such as Ogg Vorbis and MP3, as well as more recent "neural" codecs like Meta's Encodec or the Descript Audio Codec are based on block-coding; audio is divided into overlapping, fixed-size "frames" which…

声音 · 计算机科学 2025-05-12 John Vinyard

This paper presents a new neural speech compression method that is practical in the sense that it operates at low bitrate, introduces a low latency, is compatible in computational complexity with current mobile devices, and provides a…

音频与语音处理 · 电气工程与系统科学 2022-03-10 Reza Lotfidereshgi , Philippe Gournay

In this paper we present an experimental study on the performance of spatial Interference Alignment (IA) in indoor wireless local area network scenarios that use Orthogonal Frequency Division Multiplexing (OFDM) according to the…

Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-token prediction and…

声音 · 计算机科学 2025-09-10 Dimitrios Bralios , Paris Smaragdis , Jonah Casebeer

Dynamical stabilizer codes may offer a practical route to large-scale quantum computation. Such codes are defined by a schedule of error-detecting measurements, which allows for flexibility in their construction. In this work, we ask how…

Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the…

音频与语音处理 · 电气工程与系统科学 2024-09-20 Andreas Brendel , Nicola Pia , Kishan Gupta , Lyonel Behringer , Guillaume Fuchs , Markus Multrus

This study investigates the use of non-linear unsupervised dimensionality reduction techniques to compress a music dataset into a low-dimensional representation which can be used in turn for the synthesis of new sounds. We systematically…

音频与语音处理 · 电气工程与系统科学 2019-05-27 Fanny Roche , Thomas Hueber , Samuel Limier , Laurent Girin

Latent representation learning has been an active field of study for decades in numerous applications. Inspired among others by the tokenization from Natural Language Processing and motivated by the research of a simple data representation,…

信号处理 · 电气工程与系统科学 2024-09-26 Benoît Giniès , Xiaoyu Bie , Olivier Fercoq , Gaël Richard