English
Related papers

Related papers: AffectCodec: Emotion-Preserving Neural Speech Code…

200 papers

Empathetic speech dialogue requires not only understanding linguistic content but also perceiving rich paralinguistic information such as prosody, tone, and emotional intensity for affective understandings. Existing speech-to-speech large…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Zhuoyue Gao , Xiaohui Wang , Xiaocui Yang , Wen Zhang , Daling Wang , Shi Feng , Yifei Zhang

Neural networks have proven to be a formidable tool to tackle the problem of speech coding at very low bit rates. However, the design of a neural coder that can be operated robustly under real-world conditions remains a major challenge.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-08 Nicola Pia , Kishan Gupta , Srikanth Korse , Markus Multrus , Guillaume Fuchs

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with…

Discretization of semantic features enables interoperability between semantic and digital communication systems, showing significant potential for practical applications. The fundamental difficulty in digitizing semantic features stems from…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jianqiao Chen , Tingting Zhu , Huishi Song , Nan Ma , Xiaodong Xu

Understanding brain function, constructing computational models and engineering neural prosthetics require assessing two problems, namely encoding and decoding, but their relation remains controversial. For decades, the encoding problem has…

Neurons and Cognition · Quantitative Biology 2017-01-16 Hugo Gabriel Eyherabide

Deep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural speech codec with low…

Sound · Computer Science 2022-02-16 Xue Jiang , Xiulian Peng , Chengyu Zheng , Huaying Xue , Yuan Zhang , Yan Lu

Auditory front-end is an integral part of a spiking neural network (SNN) when performing auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike…

Sound · Computer Science 2019-09-05 Zihan Pan , Yansong Chua , Jibin Wu , Malu Zhang , Haizhou Li , Eliathamby Ambikairajah

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives:…

Generative image codecs aim to optimize perceptual quality, producing realistic and detailed reconstructions. However, they often overlook a key property of human vision: our tendency to focus on particular aspects of a visual scene (e.g.,…

Image and Video Processing · Electrical Eng. & Systems 2026-04-02 Lucas Relic , Roberto Azevedo , Yang Zhang , Stephan Mandt , Markus Gross , Christopher Schroers

Neural audio codecs have made significant strides in efficiently mapping raw audio waveforms into discrete token representations, which are foundational for contemporary audio generative models. However, most existing codecs are optimized…

Deploying emotion recognition systems in real-world environments where devices must be small, low-power, and private remains a significant challenge. This is especially relevant for applications such as tension monitoring, conflict…

While many current neural speech codecs achieve impressive reconstructed speech quality, they often neglect latency and complexity considerations, limiting their practical deployment in downstream tasks such as real-time speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-18 En-Wei Zhang , Hui-Peng Du , Xiao-Hang Jiang , Yang Ai , Zhen-Hua Ling

Emotion recognition from speech is a challenging task. Re-cent advances in deep learning have led bi-directional recur-rent neural network (Bi-RNN) and attention mechanism as astandard method for speech emotion recognition, extractingand…

Sound · Computer Science 2021-06-09 Zixuan Peng , Yu Lu , Shengfeng Pan , Yunfeng Liu

Most of the prevalent approaches in speech prosody modeling rely on learning global style representations in a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which…

Fundamental rate-distortion-perception (RDP) trade-offs arise in applications requiring maintained perceptual quality of reconstructed data, such as neural image compression. When compressed data is transmitted over public communication…

Information Theory · Computer Science 2026-04-23 Gustaf Åhlgren , Onur Günlü

A novel coding strategy for block-based compressive sens-ing named spatially directional predictive coding (SDPC) is proposed, which efficiently utilizes the intrinsic spatial cor-relation of natural images. At the encoder, for each block…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 Jian Zhang , Debin Zhao , Feng Jiang

In order to efficiently transmit and store speech signals, speech codecs create a minimally redundant representation of the input signal which is then decoded at the receiver with the best possible perceptual quality. In this work we…

Speech Emotion Captioning (SEC) leverages large audio-language models to generate rich, context-aware affective descriptions from speech. However, real-world deployment remains challenging due to the substantial computational demands on…

Sound · Computer Science 2026-03-13 Xiangyuan Xue , Jiajun Lu , Yan Gao , Gongping Huang , Ting Dang , Hong Jia

Diffusion-based generative image compression has demonstrated remarkable potential for achieving realistic reconstruction at ultra-low bitrates. The key to unlocking this potential lies in making the entire compression process…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Xihua Sheng , Lingyu Zhu , Tianyu Zhang , Dong Liu , Shiqi Wang , Jing Wang

The rapid advancement of generative models has led to the synthesis of real-fake ambiguous voices. To erase the ambiguity, embedding watermarks into the frequency-domain features of synthesized voices has become a common routine. However,…

Cryptography and Security · Computer Science 2025-06-24 Yue Li , Weizhi Liu , Dongdong Lin , Hui Tian , Hongxia Wang