中文
相关论文

相关论文: SecoustiCodec: Cross-Modal Aligned Streaming Singl…

200 篇论文

Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generative reconstruction objectives, contrastive approaches remain…

机器学习 · 计算机科学 2026-05-15 Hanxun Huang , Qizhou Wang , Xingjun Ma , Cihang Xie , Christopher Leckie , Sarah Erfani

Diffusion-based image compression has shown remarkable potential for achieving ultra-low bitrate coding (less than 0.05 bits per pixel) with high realism, by leveraging the generative priors of large pre-trained text-to-image diffusion…

图像与视频处理 · 电气工程与系统科学 2025-06-30 Tianyu Zhang , Xin Luo , Li Li , Dong Liu

Speech Bandwidth Extension improves clarity and intelligibility by restoring/inferring appropriate high-frequency content for low-bandwidth speech. Existing methods often rely on spectrogram or waveform modeling, which can incur higher…

声音 · 计算机科学 2026-03-04 Bowen Zhang , Junchuan Zhao , Ian McLoughlin , Ye Wang , A S Madhukumar

Gloss-free Sign Language Translation (SLT) has advanced rapidly, achieving strong performances without relying on gloss annotations. However, these gains have often come with increased model complexity and high computational demands,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 JianHe Low , Ozge Mercanoglu Sincan , Richard Bowden

Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual quality and signal…

音频与语音处理 · 电气工程与系统科学 2024-09-20 Zhikang Niu , Sanyuan Chen , Long Zhou , Ziyang Ma , Xie Chen , Shujie Liu

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and generation; and (2) how to build an audio foundation model that…

声音 · 计算机科学 2026-02-12 Dongchao Yang , Yuanyuan Wang , Dading Chong , Songxiang Liu , Xixin Wu , Helen Meng

Residual Vector Quantization (RVQ) has become a dominant approach in neural speech and audio coding, providing high-fidelity compression. However, speech coding presents additional challenges due to real-world noise, which degrades…

声音 · 计算机科学 2025-06-23 Yunkee Chae , Kyogu Lee

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicability across downstream…

机器学习 · 计算机科学 2025-07-10 Wonjin Jung , Sungil Kang , Dong-Yeon Cho

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the current speech generative models are still struggling regarding…

声音 · 计算机科学 2023-10-13 Xinfa Zhu , Yuanjun Lv , Yi Lei , Tao Li , Wendi He , Hongbin Zhou , Heng Lu , Lei Xie

Sparse coding is an unsupervised learning algorithm that learns a succinct high-level representation of the inputs given only unlabeled data; it represents each input as a sparse linear combination of a set of basis functions. Originally…

机器学习 · 计算机科学 2012-06-26 Roger Grosse , Rajat Raina , Helen Kwong , Andrew Y. Ng

Interpreting neural activity through meaningful latent representations remains a complex and evolving challenge at the intersection of neuroscience and artificial intelligence. We investigate the potential of multimodal foundation models to…

计算与语言 · 计算机科学 2025-04-22 Yijun Liu

Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly utilized in generation fields as intermediate…

声音 · 计算机科学 2023-05-09 Dongchao Yang , Songxiang Liu , Rongjie Huang , Jinchuan Tian , Chao Weng , Yuexian Zou

Discrete representation has emerged as a powerful tool in task-oriented semantic communication (ToSC), offering compact, interpretable, and efficient representations well-suited for low-power edge intelligence scenarios. Its inherent…

信号处理 · 电气工程与系统科学 2025-08-07 Anbang Zhang , Shuaishuai Guo , Chenyuan Feng , Hongyang Du , Haojin Li , Chen Sun , Haijun Zhang

The emergence of multi-codebook neutral audio codecs such as Residual Vector Quantization (RVQ) and Group Vector Quantization (GVQ) has significantly advanced Large-Language-Model (LLM) based Text-to-Speech (TTS) systems. These codecs are…

声音 · 计算机科学 2025-05-26 Rui Wang , Qianguo Sun , Tianrong Chen , Zhiyun Zeng , Junlong Wu , Jiaxing Zhang

We introduce Perception Encoder Audiovisual, PE-AV, a new family of encoders for audio and video understanding trained with scaled contrastive learning. Built on PE, PE-AV makes several key contributions to extend representations to audio,…

Neural audio coding has emerged as a vivid research direction by promising good audio quality at very low bitrates unachievable by classical coding techniques. Here, end-to-end trainable autoencoder-like models represent the state of the…

音频与语音处理 · 电气工程与系统科学 2024-09-20 Andreas Brendel , Nicola Pia , Kishan Gupta , Lyonel Behringer , Guillaume Fuchs , Markus Multrus

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often…

音频与语音处理 · 电气工程与系统科学 2024-09-19 Edresson Casanova , Ryan Langman , Paarth Neekhara , Shehzeen Hussain , Jason Li , Subhankar Ghosh , Ante Jukić , Sang-gil Lee

Semantic communications (SCs) aim to transmit only the essential information required to perform given tasks, thereby improving communication efficiency. Deep learning-based joint source-channel coding (deep JSCC) has emerged as a promising…

信号处理 · 电气工程与系统科学 2026-04-07 Eunhye Hong , Taewoo Park , Yongjune Kim

Speech disorders such as dysarthria and anarthria can severely impair the patient's ability to communicate verbally. Speech decoding brain-computer interfaces (BCIs) offer a potential alternative by directly translating speech intentions…

人机交互 · 计算机科学 2025-05-27 Hongbin Wang , Zhihong Jia , Yuanzhong Shen , Ziwei Wang , Siyang Li , Kai Shu , Feng Hu , Dongrui Wu

In low-bitrate speech coding, end-to-end speech coding networks aim to learn compact yet expressive features and a powerful decoder in a single network. A challenging problem as such results in unwelcome complexity increase and inferior…

音频与语音处理 · 电气工程与系统科学 2023-11-16 Haici Yang , Inseon Jang , Minje Kim