English
Related papers

Related papers: USM-VC: Mitigating Timbre Leakage with Universal S…

200 papers

Although semantic communications have exhibited satisfactory performance for a large number of tasks, the impact of semantic noise and the robustness of the systems have not been well investigated. Semantic noise refers to the misleading…

Signal Processing · Electrical Eng. & Systems 2023-04-20 Qiyu Hu , Guangyi Zhang , Zhijin Qin , Yunlong Cai , Guanding Yu , Geoffrey Ye Li

The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which…

Sound · Computer Science 2024-09-30 Youngjae Kim , Yejin Jeon , Gary Geunbae Lee

Recent works on voice conversion (VC) focus on preserving the rhythm and the intonation as well as the linguistic content. To preserve these features from the source, we decompose current non-parallel VC systems into two encoders and one…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Kang-wook Kim , Seung-won Park , Junhyeok Lee , Myun-chul Joe

Word-piece models (WPMs) are commonly used subword units in state-of-the-art end-to-end automatic speech recognition (ASR) systems. For multilingual ASR, due to the differences in written scripts across languages, multilingual WPMs bring…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Chao Zhang , Bo Li , Tara N. Sainath , Trevor Strohman , Shuo-yiin Chang

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice…

Sound · Computer Science 2023-04-04 Haozhe Zhang , Zexin Cai , Xiaoyi Qin , Ming Li

Task-oriented semantic communications have achieved significant performance gains. However, the employed deep neural networks in semantic communications have to be updated when the task is changed or multiple models need to be stored for…

Signal Processing · Electrical Eng. & Systems 2024-06-11 Guangyi Zhang , Qiyu Hu , Zhijin Qin , Yunlong Cai , Guanding Yu , Xiaoming Tao

Generative semantic communication models are reshaping semantic communication frameworks by moving beyond pixel-wise optimization to align with human perception. However, many existing approaches prioritize image-level perceptual quality,…

Signal Processing · Electrical Eng. & Systems 2025-04-08 Kailang Ye , Mingze Gong , Shuoyao Wang , Daquan Feng

Uniform Meaning Representation (UMR) is a novel graph-based semantic representation which captures the core meaning of a text, with flexibility incorporated into the annotation schema such that the breadth of the world's languages can be…

Computation and Language · Computer Science 2026-05-20 Emma Markle , Javier Gutierrez Bach , Shira Wein

Quantization in SSL speech models (e.g., HuBERT) improves compression and performance in tasks like language modeling, resynthesis, and text-to-speech but often discards prosodic and paralinguistic information (e.g., emotion, prominence).…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Nicholas Sanders , Yuanchao Li , Korin Richmond , Simon King

Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capture long-range context. However, due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2019-09-30 Lin Song , Yanwei Li , Zeming Li , Gang Yu , Hongbin Sun , Jian Sun , Nanning Zheng

Voice conversion (VC) is a task to transform a person's voice to different style while conserving linguistic contents. Previous state-of-the-art on VC is based on sequence-to-sequence (seq2seq) model, which could mislead linguistic…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-28 Tae-Ho Kim , Sungjae Cho , Shinkook Choi , Sejik Park , Soo-Young Lee

We present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network…

Singing Voice Separation (SVS) tries to separate singing voice from a given mixed musical signal. Recently, many U-Net-based models have been proposed for the SVS task, but there were no existing works that evaluate and compare various…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-09 Woosung Choi , Minseok Kim , Jaehwa Chung , Daewon Lee , Soonyoung Jung

As a type of biometric identification, a speaker identification (SID) system is confronted with various kinds of attacks. The spoofing attacks typically imitate the timbre of the target speakers, while the adversarial attacks confuse the…

Sound · Computer Science 2023-09-06 Qing Wang , Jixun Yao , Li Zhang , Pengcheng Guo , Lei Xie

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities. Nevertheless, they suffer from multiple stability issues…

We present a novel typical-to-atypical voice conversion approach (DuTa-VC), which (i) can be trained with nonparallel data (ii) first introduces diffusion probabilistic model (iii) preserves the target speaker identity (iv) is aware of the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Helin Wang , Thomas Thebaud , Jesus Villalba , Myra Sydnor , Becky Lammers , Najim Dehak , Laureano Moro-Velazquez

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without…

As Vision-Language Models (VLMs) are increasingly deployed in split-DNN configurations--with visual encoders (e.g., ResNet, ViT) operating on user devices and sending intermediate features to the cloud--there is a growing privacy risk from…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Kedong Xiu , Sai Qian Zhang

Textless spoken language models (SLMs) are generative models of speech that do not rely on text supervision. Most textless SLMs learn to predict the next semantic token, a discrete representation of linguistic content, and rely on a…

Computation and Language · Computer Science 2025-10-23 Ju-Chieh Chou , Jiawei Zhou , Karen Livescu

Cross-lingual timbre and style generalizable text-to-speech (TTS) aims to synthesize speech with a specific reference timbre or style that is never trained in the target language. It encounters the following challenges: 1) timbre and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-28 Yahuan Cong , Haoyu Zhang , Haopeng Lin , Shichao Liu , Chunfeng Wang , Yi Ren , Xiang Yin , Zejun Ma