English
Related papers

Related papers: Estimating the Completeness of Discrete Speech Uni…

200 papers

Learning disentangled representations of textual data is essential for many natural language tasks such as fair classification, style transfer and sentence generation, among others. The existent dominant approaches in the context of text…

Artificial Intelligence · Computer Science 2021-05-07 Pierre Colombo , Chloe Clavel , Pablo Piantanida

Speech segmentation at both word and phoneme levels is crucial for various speech processing tasks. It significantly aids in extracting meaningful units from an utterance, thus enabling the generation of discrete elements. In this work we…

Machine Learning · Computer Science 2024-11-18 Simone Carnemolla , Salvatore Calcagno , Simone Palazzo , Daniela Giordano

We investigate notions of ambiguity and partial information in categorical distributional models of natural language. Probabilistic ambiguity has previously been studied using Selinger's CPM construction. This construction works well for…

Logic in Computer Science · Computer Science 2017-01-04 Dan Marsden

Humans have a remarkable ability to disentangle complex sensory inputs (e.g., image, text) into simple factors of variation (e.g., shape, color) without much supervision. This ability has inspired many works that attempt to solve the…

Machine Learning · Computer Science 2024-12-25 Kartik Ahuja , Divyat Mahajan , Vasilis Syrgkanis , Ioannis Mitliagkas

Quantum data hiding stores classical information in bipartite quantum states that are, in principle, perfectly distinguishable, yet remain almost indistinguishable without access to a quantum communication channel. Here, we investigate…

Quantum Physics · Physics 2025-11-07 Aby Philip , Alexander Streltsov

We propose a framework to analyze how multivariate representations disentangle ground-truth generative factors. A quantitative analysis of disentanglement has been based on metrics designed to compare how one variable explains each…

Machine Learning · Statistics 2022-02-11 Seiya Tokui , Issei Sato

With the advent of general-purpose speech representations from large-scale self-supervised models, applying a single model to multiple downstream tasks is becoming a de-facto approach. However, the pooling problem remains; the length of…

Machine Learning · Computer Science 2023-04-11 Jeongkyun Park , Kwanghee Choi , Hyunjun Heo , Hyung-Min Park

Communicating state machines provide a formal foundation for distributed computation. Unfortunately, they are Turing-complete and, thus, challenging to analyse. In this paper, we classify restrictions on channels which have been proposed to…

Formal Languages and Automata Theory · Computer Science 2022-08-12 Felix Stutz , Damien Zufferey

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives:…

We present the Zero Resource Speech Challenge 2020, which aims at learning speech representations from raw audio signals without any labels. It combines the data sets and metrics from two previous benchmarks (2017 and 2019) and features two…

Computation and Language · Computer Science 2020-10-14 Ewan Dunbar , Julien Karadayi , Mathieu Bernard , Xuan-Nga Cao , Robin Algayres , Lucas Ondel , Laurent Besacier , Sakriani Sakti , Emmanuel Dupoux

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Yuying Xie , Michael Kuhlmann , Frederik Rautenberg , Zheng-Hua Tan , Reinhold Haeb-Umbach

We propose the notion of process resource-breaking channels that break the resource for a quantum information processing task. We examine the same using quantum dense coding and teleportation protocols. We prove that the sets DBT (dense…

Quantum Physics · Physics 2025-01-28 Abhishek Muhuri , Ayan Patra , Rivu Gupta , Aditi Sen De

Speaker embedding has been a fundamental feature for speaker-related tasks such as verification, clustering, and diarization. Traditionally, speaker embeddings are represented as fixed vectors in high-dimensional space. This could lead to…

Sound · Computer Science 2022-06-28 Siqi Zheng , Hongbin Suo , Qian Chen

Communicating state machines provide a formal foundation for distributed computation. Unfortunately, they are Turing-complete and, thus, challenging to analyse. In this paper, we classify restrictions on channels which have been proposed to…

Formal Languages and Automata Theory · Computer Science 2022-09-22 Felix Stutz , Damien Zufferey

The problem of private data disclosure is studied from an information theoretic perspective. Considering a pair of dependent random variables $(X,Y)$, where $X$ and $Y$ denote the private and useful data, respectively, the following problem…

Information Theory · Computer Science 2021-01-25 Borzoo Rassouli , Deniz Gunduz

Speakers often have multiple ways to express the same meaning. The Uniform Information Density (UID) hypothesis suggests that speakers exploit this variability to maintain a consistent rate of information transmission during language…

Computation and Language · Computer Science 2025-10-27 Hailin Hao , Elsi Kaiser

Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank…

Audio and Speech Processing · Electrical Eng. & Systems 2018-01-12 Tak-Shing T. Chan , Yi-Hsuan Yang

Neural networks and other machine learning models compute continuous representations, while humans communicate with discrete symbols. Reconciling these two forms of communication is desirable to generate human-readable interpretations or to…

Machine Learning · Computer Science 2021-04-05 André F. T. Martins

Voice conversion (VC) is a task that transforms the source speaker's timbre, accent, and tones in audio into another one's while preserving the linguistic content. It is still a challenging work, especially in a one-shot setting.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-09 Da-Yi Wu , Yen-Hao Chen , Hung-Yi Lee

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Esther Sun , Abinay Reddy Naini , Carlos Busso