English
Related papers

Related papers: SpecTokenizer: A Lightweight Streaming Codec in th…

200 papers

We introduce STAR (Stream Transduction with Anchor Representations), a novel Transformer-based model designed for efficient sequence-to-sequence transduction over streams. STAR dynamically segments input streams to create compressed anchor…

Computation and Language · Computer Science 2025-05-22 Weiting Tan , Yunmo Chen , Tongfei Chen , Guanghui Qin , Haoran Xu , Heidi C. Zhang , Benjamin Van Durme , Philipp Koehn

Sentence compression is a Natural Language Processing (NLP) task aimed at shortening original sentences and preserving their key information. Its applications can benefit many fields e.g. one can build tools for language education. However,…

Computation and Language · Computer Science 2020-09-24 Weiwei Hou , Hanna Suominen , Piotr Koniusz , Sabrina Caldwell , Tom Gedeon

Convolutional neural networks (CNN) and Transformer have wildly succeeded in multimedia applications. However, more effort needs to be made to harmonize these two architectures effectively to satisfy speech enhancement. This paper aims to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Xinmeng Xu , Weiping Tu , Yuhong Yang

Having a sequence-to-sequence model which can operate in an online fashion is important for streaming applications such as Voice Search. Neural transducer is a streaming sequence-to-sequence model, but has shown a significant degradation in…

Computation and Language · Computer Science 2017-12-06 Tara N. Sainath , Chung-Cheng Chiu , Rohit Prabhavalkar , Anjuli Kannan , Yonghui Wu , Patrick Nguyen , Zhifeng Chen

Convolutional neural networks (CNNs) are commonplace in high-performing solutions to many real-world problems, such as audio classification. CNNs have many parameters and filters, with some having a larger impact on the performance than…

Sound · Computer Science 2023-05-08 James A King , Arshdeep Singh , Mark D. Plumbley

Transformers have drawn attention in the MIR field for their remarkable performance shown in natural language processing and computer vision. However, prior works in the audio processing domain mostly use Transformer as a temporal feature…

Sound · Computer Science 2021-10-26 Wei-Tsung Lu , Ju-Chiang Wang , Minz Won , Keunwoo Choi , Xuchen Song

Spiking neural networks (SNNs) are the third generation of neural networks and can explore both rate and temporal coding for energy-efficient event-driven computation. However, the decision accuracy of existing SNN designs is contingent…

Neural and Evolutionary Computing · Computer Science 2020-02-25 Changqing Xu , Wenrui Zhang , Yu Liu , Peng Li

Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in…

Sound · Computer Science 2024-07-31 Youqiang Zheng , Weiping Tu , Li Xiao , Xinmeng Xu

Deep learning has shown impressive performance in semantic segmentation, but it is still unaffordable for resource-constrained mobile devices. While offloading computation tasks is promising, the high traffic demands overwhelm the limited…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Xuedou Xiao , Juecheng Zhang , Wei Wang , Jianhua He , Qian Zhang

This paper considers the joint compression and enhancement problem for speech signal in the presence of noise. Recently, the SoundStream codec, which relies on end-to-end joint training of an encoder-decoder pair and a residual vector…

Sound · Computer Science 2025-09-03 Jiayi Huang , Zeyu Yan , Wenbin Jiang , He Wang , Fei Wen

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as speech coding and large language models (LLM). However, most…

Sound · Computer Science 2025-07-01 Youqiang Zheng , Weiping Tu , Yueteng Kang , Jie Chen , Yike Zhang , Li Xiao , Yuhong Yang , Long Ma

This work introduces TTS-Transducer - a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-16 Vladimir Bataev , Subhankar Ghosh , Vitaly Lavrukhin , Jason Li

Spectral band replication (SBR) enables bit-efficient coding by generating high-frequency bands from the low-frequency ones. However, it only utilizes coarse spectral features upon a subband-wise signal replication, limiting adaptability to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-29 Woongjib Choi , Byeong Hyeon Kim , Hyungseob Lim , Inseon Jang , Hong-Goo Kang

Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction…

Sound · Computer Science 2025-09-10 Dimitrios Bralios , Jonah Casebeer , Paris Smaragdis

Enhancing coded speech suffering from far-end acoustic background noise, quantization noise, and potentially transmission errors, is a challenging task. In this work we propose two postprocessing approaches applying convolutional neural…

Audio and Speech Processing · Electrical Eng. & Systems 2019-01-25 Ziyue Zhao , Huijun Liu , Tim Fingscheidt

Tokenized visual representations have shown promise in image compression, yet their extension to video remains underexplored due to the challenges posed by complex temporal dynamics and stringent bit rate constraints. In this paper, we…

Image and Video Processing · Electrical Eng. & Systems 2025-11-20 Lebin Zhou , Cihan Ruan , Nam Ling , Zhenghao Chen , Wei Wang , Wei Jiang

Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. However, speech is…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Hankun Wang , Yiwei Guo , Chongtian Shao , Bohan Li , Kai Yu

Neural audio codecs are widely used for audio compression and can be integrated into token-based language models. Traditional codecs preserve acoustic details well but lack semantic information. Recent hybrid codecs attempt to incorporate…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-09 Kaiyuan Zhang , Mohan Shi , Eray Eren , Natarajan Balaji Shankar , Zilai Wang , Abeer Alwan

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for…

Sound · Computer Science 2022-03-29 Jun Chen , Zilin Wang , Deyi Tuo , Zhiyong Wu , Shiyin Kang , Helen Meng

This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen languages during pre-training, (ii) whether speech-only pre-trained…

Sound · Computer Science 2026-01-21 Shih-Heng Wang , Jiatong Shi , Jinchuan Tian , Haibin Wu , Shinji Watanabe