中文
相关论文

相关论文: WavCube: Unifying Speech Representation for Unders…

200 篇论文

The crux of single-channel speech separation is how to encode the mixture of signals into such a latent embedding space that the signals from different speakers can be precisely separated. Existing methods for speech separation either…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Zengwei Yao , Wenjie Pei , Fanglin Chen , Guangming Lu , David Zhang

In this paper, we propose a technique to alleviate the quality degradation caused by collapsed speech segments sometimes generated by the WaveNet vocoder. The effectiveness of the WaveNet vocoder for generating natural speech from acoustic…

音频与语音处理 · 电气工程与系统科学 2018-08-10 Yi-Chiao Wu , Kazuhiro Kobayashi , Tomoki Hayashi , Patrick Lumban Tobing , Tomoki Toda

The rapid advancement in self-supervised representation learning has highlighted its potential to leverage unlabeled data for learning rich visual representations. However, the existing techniques, particularly those employing different…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Sana Ayromlou , Vahid Reza Khazaie , Fereshteh Forghani , Arash Afkanpour

Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Alexander H. Liu , Sang-gil Lee , Chao-Han Huck Yang , Yuan Gong , Yu-Chiang Frank Wang , James R. Glass , Rafael Valle , Bryan Catanzaro

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

图形学 · 计算机科学 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Self-supervised learning (SSL) speech representations learned from large amounts of diverse, mixed-quality speech data without transcriptions are gaining ground in many speech technology applications. Prior work has shown that SSL is an…

音频与语音处理 · 电气工程与系统科学 2023-07-12 Siyang Wang , Gustav Eje Henter , Joakim Gustafson , Éva Székely

Information in speech can be categorized into two groups: Content (what is being said, such as linguistics) and Other (how it is expressed such as information about speaker and paralinguistic features). Current self-supervised learning…

计算与语言 · 计算机科学 2025-02-19 Hemant Yadav , Rajiv Ratn Shah , Sunayana Sitaram

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity…

Self-supervised models have revolutionized speech processing, achieving new levels of performance in a wide variety of tasks with limited resources. However, the inner workings of these models are still opaque. In this paper, we aim to…

声音 · 计算机科学 2024-06-25 Yassine El Kheir , Ahmed Ali , Shammur Absar Chowdhury

Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive…

音频与语音处理 · 电气工程与系统科学 2024-01-26 Heming Wang , Eric W. Healy , DeLiang Wang

Recent speech enhancement (SE) models increasingly leverage self-supervised learning (SSL) representations for their rich semantic information. Typically, intermediate features are aggregated into a single representation via a lightweight…

声音 · 计算机科学 2026-02-02 Seungu Han , Sungho Lee , Kyogu Lee

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement.…

Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability…

音频与语音处理 · 电气工程与系统科学 2025-08-15 Abhijit Sinha , Harishankar Kumar , Mohit Joshi , Hemant Kumar Kathania , Shrikanth Narayanan , Sudarsana Reddy Kadiri

Self-supervised learning (SSL) speech models, which can serve as powerful upstream models to extract meaningful speech representations, have achieved unprecedented success in speech representation learning. However, their effectiveness on…

声音 · 计算机科学 2023-02-01 Tung-Yu Wu , Chen-An Li , Tzu-Han Lin , Tsu-Yuan Hsu , Hung-Yi Lee

Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL)…

音频与语音处理 · 电气工程与系统科学 2025-09-01 Abhijit Sinha , Hemant Kumar Kathania , Sudarsana Reddy Kadiri , Shrikanth Narayanan

Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SSL models focus on leveraging continuous representations as…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Zehan Li , Yan Yang , Xueqing Li , Jian Kang , Xiao-Lei Zhang , Jie Li

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into…

We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes…

音频与语音处理 · 电气工程与系统科学 2024-06-26 Joun Yeop Lee , Myeonghun Jeong , Minchan Kim , Ji-Hyun Lee , Hoon-Young Cho , Nam Soo Kim

Speech discrete representation has proven effective in various downstream applications due to its superior compression rate of the waveform, fast convergence during training, and compatibility with other modalities. Discrete units extracted…

声音 · 计算机科学 2024-06-17 Jiatong Shi , Xutai Ma , Hirofumi Inaguma , Anna Sun , Shinji Watanabe

Speaker representation learning is crucial for voice recognition systems, with recent advances in self-supervised approaches reducing dependency on labeled data. Current two-stage iterative frameworks, while effective, suffer from…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Danwei Cai , Zexin Cai , Ze Li , Ming Li