English
Related papers

Related papers: Estimating the Completeness of Discrete Speech Uni…

200 papers

Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from surrounding representations can inadvertently entangle speaker…

Sound · Computer Science 2024-04-02 Injune Hwang , Kyogu Lee

The coherent information concept is used to analyze a variety of simple quantum systems. Coherent information was calculated for the information decay in a two-level atom in the presence of an external resonant field, for the information…

Quantum Physics · Physics 2009-10-31 B. A. Grishanin , V. N. Zadkov

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and another trained on monolingual data. We explore the multi- and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-29 Jennifer Williams , Jason Fong , Erica Cooper , Junichi Yamagishi

Variational autoencoders have been widely applied for natural language generation, however, there are two long-standing problems: information under-representation and posterior collapse. The former arises from the fact that only the last…

Machine Learning · Computer Science 2021-06-17 Xianghong Fang , Haoli Bai , Zenglin Xu , Michael Lyu , Irwin King

Large-scale speech self-supervised learning (SSL) has emerged to the main field of speech processing, however, the problem of computational cost arising from its vast size makes a high entry barrier to academia. In addition, existing…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-04 Yeonghyeon Lee , Kangwook Jang , Jahyun Goo , Youngmoon Jung , Hoirin Kim

The primary characteristic of robust speaker representations is that they are invariant to factors of variability not related to speaker identity. Disentanglement of speaker representations is one of the techniques used to improve…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Raghuveer Peri , Haoqi Li , Krishna Somandepalli , Arindam Jati , Shrikanth Narayanan

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-01 Jiachen Lian , Chunlei Zhang , Dong Yu

We address the framework of analysing quantum metrology in the information-theoretic picture. Firstly we show how to extract the maximum amount of information in general via suitable state initialization of the probes at the beginning and a…

Quantum Physics · Physics 2018-10-02 Yi Peng , Heng Fan

We investigate so-called localisable information of bipartite states and a parallel notion of information deficit. Localisable information is defined as the amount of information that can be concentrated by means of classical communication…

Quantum Physics · Physics 2015-06-26 Barbara Synak , Karol Horodecki , Michal Horodecki

For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Tianchi Liu , Kong Aik Lee , Qiongqiong Wang , Haizhou Li

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Knowledge distillation deploys complex machine learning models in resource-constrained environments by training a smaller student model to emulate internal representations of a complex teacher model. However, the teacher's representations…

Preserving a patient's identity is a challenge for automatic, speech-based diagnosis of mental health disorders. In this paper, we address this issue by proposing adversarial disentanglement of depression characteristics and speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-08 Vijay Ravi , Jinhan Wang , Jonathan Flint , Abeer Alwan

Much research effort is being applied to the task of compressing the knowledge of self-supervised models, which are powerful, yet large and memory consuming. In this work, we show that the original method of knowledge distillation (and its…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Danilo de Oliveira , Timo Gerkmann

Self-supervised pre-trained models such as HuBERT and WavLM leverage unlabeled speech data for representation learning and offer significantly improve for numerous downstream tasks. Despite the success of these methods, their large memory…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-23 Yingying Gao , Shilei Zhang , Zihao Cui , Yanhan Xu , Chao Deng , Junlan Feng

Neural audio codecs (NACs), which use neural networks to generate compact audio representations, have garnered interest for their applicability to many downstream tasks -- especially quantized codecs due to their compatibility with large…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-13 Ryo Aihara , Yoshiki Masuyama , Gordon Wichern , François G. Germain , Jonathan Le Roux

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a…

Sound · Computer Science 2023-03-01 Dianwen Ng , Ruixi Zhang , Jia Qi Yip , Zhao Yang , Jinjie Ni , Chong Zhang , Yukun Ma , Chongjia Ni , Eng Siong Chng , Bin Ma

Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content information in…

Sound · Computer Science 2025-05-27 Giuseppe Ruggiero , Matteo Testa , Jurgen Van de Walle , Luigi Di Caro

We introduce DiceHuBERT, a knowledge distillation framework for compressing HuBERT, a widely used self-supervised learning (SSL)-based speech foundation model. Unlike existing distillation methods that rely on layer-wise and feature-wise…

Child-centered daylong recordings are essential for studying early language development, but existing speech models trained on clean adult data perform poorly due to acoustic and linguistic differences. We introduce BabyHuBERT, a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Théo Charlot , Tarek Kunze , Maxime Poli , Alejandrina Cristia , Emmanuel Dupoux , Marvin Lavechin