English
Related papers

Related papers: SpatialCodec: Neural Spatial Speech Coding

200 papers

In this paper, we introduce the novel use of linear spatial precoding based on fixed and known parameters of multiple-input multiple-output (MIMO) channels to improve the performance of space-time coded MIMO systems. We derive linear…

Information Theory · Computer Science 2007-07-13 Tharaka A. Lamahewa , Rodney A. Kennedy , Thushara D. Abhayapala , Van K. Nguyen

In recent years, resolution adaptation based on deep neural networks has enabled significant performance gains for conventional (2D) video codecs. This paper investigates the effectiveness of spatial resolution resampling in the context of…

Image and Video Processing · Electrical Eng. & Systems 2022-02-28 Angeliki Katsenou , Fan Zhang , David Bull

This study presents a system for sound source localization in time domain using a deep residual neural network. Data from the linear 8 channel microphone array with 3 cm spacing is used by the network for direction estimation. We propose to…

Sound · Computer Science 2018-08-21 Dmitry Suvorov , Ge Dong , Roman Zhukov

This paper proposes a neural network based speech separation method using spatially distributed microphones. Unlike with traditional microphone array settings, neither the number of microphones nor their spatial arrangement is known in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Dongmei Wang , Zhuo Chen , Takuya Yoshioka

Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-20 Andreas Brendel , Nicola Pia , Kishan Gupta , Lyonel Behringer , Guillaume Fuchs , Markus Multrus

In the rapidly evolving fields of virtual and augmented reality, accurate spatial audio capture and reproduction are essential. For these applications, Ambisonics has emerged as a standard format. However, existing methods for encoding…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-27 Yhonatan Gayer , Vladimir Tourbabin , Zamir Ben-Hur , Jacob Donley , Boaz Rafaely

With recent rapid growth of large language models (LLMs), discrete speech tokenization has played an important role for injecting speech into LLMs. However, this discretization gives rise to a loss of information, consequently impairing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Zhichao Huang , Chutong Meng , Tom Ko

The emerging brain-inspired computing paradigm known as hyperdimensional computing (HDC) has been proven to provide a lightweight learning framework for various cognitive tasks compared to the widely used deep learning-based approaches.…

Emerging Technologies · Computer Science 2021-06-23 Geethan Karunaratne , Manuel Le Gallo , Michael Hersche , Giovanni Cherubini , Luca Benini , Abu Sebastian , Abbas Rahimi

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and…

Sound · Computer Science 2024-01-17 Ashutosh Pandey , Buye Xu

The advent of neuralmorphic spike cameras has garnered significant attention for their ability to capture continuous motion with unparalleled temporal resolution.However, this imaging attribute necessitates considerable resources for binary…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Kexiang Feng , Chuanmin Jia , Siwei Ma , Wen Gao

In this paper, techniques for improving multichannel lossless coding are examined. A method is proposed for the simultaneous coding of two or more different renderings (mixes) of the same content. The signal model uses both past samples of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-31 Toni Hirvonen , Mahmoud Namazi

When using artificial neural networks for multichannel speech enhancement, filtering is often achieved by estimating a complex-valued mask that is applied to all or one reference channel of the input signal. The estimation of this mask is…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Annika Briegleb , Walter Kellermann

We propose an efficient method to estimate source power spectral densities (PSDs) in a multi-source reverberant environment using a spherical microphone array. The proposed method utilizes the spatial correlation between the spherical…

Sound · Computer Science 2018-05-21 Abdullah Fahim , Prasanga N. Samarasinghe , Thushara D. Abhayapala

The past few years have seen remarkable progress in the decoding of speech from brain activity, primarily driven by large single-subject datasets. However, due to individual variation, such as anatomy, and differences in task design and…

Machine Learning · Computer Science 2025-06-03 Dulhan Jayalath , Gilad Landau , Brendan Shillingford , Mark Woolrich , Oiwi Parker Jones

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this paper, we propose a multi-modal DSR model by leveraging neural…

Sound · Computer Science 2024-06-25 Xueyuan Chen , Dongchao Yang , Dingdong Wang , Xixin Wu , Zhiyong Wu , Helen Meng

We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-02 Mikko Heikkinen , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

The spatial auditory attention decoding (Sp-AAD) technology aims to determine the direction of auditory attention in multi-talker scenarios via neural recordings. Despite the success of recent Sp-AAD algorithms, their performance is…

Human-Computer Interaction · Computer Science 2024-07-10 Zelin Qiu , Jianjun Gu , Dingding Yao , Junfeng Li

Recently, an end-to-end two-dimensional sound source localization algorithm with ad-hoc microphone arrays formulates the sound source localization problem as a classification problem. The algorithm divides the target indoor space into a set…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-18 Linfeng Feng , Yijun Gong , Xiao-Lei Zhang

Spatial frequency analysis and transforms serve a central role in most engineered image and video lossy codecs, but are rarely employed in neural network (NN)-based approaches. We propose a novel NN-based image coding framework that…

Image and Video Processing · Electrical Eng. & Systems 2023-01-04 Hyomin Choi , Fabien Racape , Shahab Hamidi-Rad , Mateen Ulhaq , Simon Feltman

Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction…

Sound · Computer Science 2025-09-10 Dimitrios Bralios , Jonah Casebeer , Paris Smaragdis