English
Related papers

Related papers: VQVAE Unsupervised Unit Discovery and Multi-scale …

200 papers

Recent advancements in Self-Supervised Learning (SSL) have shown promising results in Speaker Verification (SV). However, narrowing the performance gap with supervised systems remains an ongoing challenge. Several studies have observed that…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Victor Miara , Theo Lepage , Reda Dehak

Wastewater-based genomic surveillance has emerged as a powerful tool for population-level viral monitoring, offering comprehensive insights into circulating viral variants across entire communities. However, this approach faces significant…

Machine Learning · Computer Science 2025-12-04 Adele Chinda , Richmond Azumah , Hemanth Demakethepalli Venkateswara

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and another trained on monolingual data. We explore the multi- and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-29 Jennifer Williams , Jason Fong , Erica Cooper , Junichi Yamagishi

Current state-of-the-art open-vocabulary segmentation methods typically rely on image-mask-text triplet annotations for supervision. However, acquiring such detailed annotations is labour-intensive and poses scalability challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Zhaoqing Wang , Xiaobo Xia , Ziye Chen , Xiao He , Yandong Guo , Mingming Gong , Tongliang Liu

Speech synthesis systems powered by neural networks hold promise for multimedia production, but frequently face issues with producing expressive speech and seamless editing. In response, we present the Cross-Utterance Conditioned…

Sound · Computer Science 2024-09-20 Yang Li , Cheng Yu , Guangzhi Sun , Weiqin Zu , Zheng Tian , Ying Wen , Wei Pan , Chao Zhang , Jun Wang , Yang Yang , Fanglei Sun

This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-16 Aemon Yat Fei Chiu , Jingyu Li , Yusheng Tian , Guangyan Zhang , Tan Lee

An effective approach to non-parallel voice conversion (VC) is to utilize deep neural networks (DNNs), specifically variational auto encoders (VAEs), to model the latent structure of speech in an unsupervised manner. A previous study has…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Wen-Chin Huang , Hsin-Te Hwang , Yu-Huai Peng , Yu Tsao , Hsin-Min Wang

Recent progress in scaling up large language models has shown impressive capabilities in performing few-shot learning across a wide range of text-based tasks. However, a key limitation is that these language models fundamentally lack visual…

Machine Learning · Computer Science 2023-02-06 Hao Liu , Wilson Yan , Pieter Abbeel

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To…

Sound · Computer Science 2025-05-27 Jingguang Tian , Xinhui Hu , Xinkang Xu

Residual Vector Quantization (RVQ) has become a dominant approach in neural speech and audio coding, providing high-fidelity compression. However, speech coding presents additional challenges due to real-world noise, which degrades…

Sound · Computer Science 2025-06-23 Yunkee Chae , Kyogu Lee

Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Tommaso Boccato , Michal Olak , Matteo Ferrante

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss, noise pollution, and audio file corruption can occur due…

Sound · Computer Science 2025-12-02 Yaxuan Li , Han Jiang , Yifei Ma , Shihua Qin , Jonghye Woo , Fangxu Xing

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…

Computation and Language · Computer Science 2018-08-08 Yu-Hsuan Wang , Hung-yi Lee , Lin-shan Lee

Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different…

Sound · Computer Science 2025-04-17 Soobin Suh , Dabi Ahn , Heewoong Park , Jonghun Park

This paper describes our submitted systems to the 2022 ADD challenge withing the tracks 1 and 2. Our approach is based on the combination of a pre-trained wav2vec2 feature extractor and a downstream classifier to detect spoofed audio. This…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-04 Juan M. Martín-Doñas , Aitor Álvarez

We describe the Phonexia submission for the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21) in the unsupervised speaker verification track. Our solution was very similar to IDLab's winning submission for VoxSRC-20. An embedding…

Sound · Computer Science 2021-09-09 Josef Slavíček , Albert Swart , Michal Klčo , Niko Brümmer

Masked Autoencoders (MAEs) learn rich low-level representations from unlabeled data but require substantial labeled data to effectively adapt to downstream tasks. Conversely, Instance Discrimination (ID) emphasizes high-level semantics,…

Sound · Computer Science 2024-03-15 Afrina Tabassum , Dung Tran , Trung Dang , Ismini Lourentzou , Kazuhito Koishida

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-01 Jiachen Lian , Chunlei Zhang , Dong Yu

This research presents a novel approach to enhancing automatic speech recognition systems by integrating noise detection capabilities directly into the recognition architecture. Building upon the wav2vec2 framework, the proposed method…

Sound · Computer Science 2025-12-11 Karamvir Singh