English
Related papers

Related papers: Synthetic Speech Source Tracing using Metric Learn…

200 papers

We introduce Metis, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled speech data using…

Sound · Computer Science 2025-02-06 Yuancheng Wang , Jiachen Zheng , Junan Zhang , Xueyao Zhang , Huan Liao , Zhizheng Wu

Recent advances in Text-To-Speech (TTS) technology have enabled synthetic speech to mimic human voices with remarkable realism, raising significant security concerns. This underscores the need for traceable TTS models-systems capable of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Yuxiang Zhao , Yunchong Xiao , Yushen Chen , Zhikang Niu , Shuai Wang , Kai Yu , Xie Chen

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

Adversarial attack approaches to speaker identification either need high computational cost or are not very effective, to our knowledge. To address this issue, in this paper, we propose a novel generation-network-based approach, called…

Sound · Computer Science 2023-02-28 Jiadi Yao , Xing Chen , Xiao-Lei Zhang , Wei-Qiang Zhang , Kunde Yang

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work proposes a novel audio…

Sound · Computer Science 2025-06-04 Ajinkya Kulkarni , Sandipana Dowerah , Tanel Alumae , Mathew Magimai. -Doss

Speech synthesis technology has witnessed significant advancements in recent years, enabling the creation of natural and expressive synthetic speech. One area of particular interest is the generation of synthetic child speech, which…

Sound · Computer Science 2023-11-09 Rishabh Jain , Peter Corcoran

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech…

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Self-supervised learning (SSL) for speech representation has been successfully applied in various downstream tasks, such as speech and speaker recognition. More recently, speech SSL models have also been shown to be beneficial in advancing…

Computation and Language · Computer Science 2024-08-28 Takanori Ashihara , Takafumi Moriya , Kohei Matsuura , Tomohiro Tanaka , Yusuke Ijima , Taichi Asami , Marc Delcroix , Yukinori Honma

Many deep learning synthetic speech generation tools are readily available. The use of synthetic speech has caused financial fraud, impersonation of people, and misinformation to spread. For this reason forensic methods that can detect…

End-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio without employing the standard pipeline composed of a separately trained…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-22 Loren Lugosch , Brett Meyer , Derek Nowrouzezahrai , Mirco Ravanelli

Freely available and easy-to-use audio editing tools make it straightforward to perform audio splicing. Convincing forgeries can be created by combining various speech samples from the same person. Detection of such splices is important…

Sound · Computer Science 2024-05-06 Denise Moussa , Germans Hirsch , Christian Riess

Deepfake speech utterances can be forged by replacing one or more words in a bona fide utterance with semantically different words synthesized with speech-generative models. While a dedicated synthetic word detector could be developed, we…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Hoan My Tran , Xin Wang , Wanying Ge , Xuechen Liu , Junichi Yamagishi

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic information, and speaker…

State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-25 Théo Lepage , Réda Dehak

This study presents a system for sound source localization in time domain using a deep residual neural network. Data from the linear 8 channel microphone array with 3 cm spacing is used by the network for direction estimation. We propose to…

Sound · Computer Science 2018-08-21 Dmitry Suvorov , Ge Dong , Roman Zhukov

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker…

Computation and Language · Computer Science 2025-05-13 Junyi Peng , Takanori Ashihara , Marc Delcroix , Tsubasa Ochiai , Oldrich Plchot , Shoko Araki , Jan Černocký

Recently, pioneer work finds that speech pre-trained models can solve full-stack speech processing tasks, because the model utilizes bottom layers to learn speaker-related information and top layers to encode content-related information.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-17 Chengyi Wang , Yu Wu , Sanyuan Chen , Shujie Liu , Jinyu Li , Yao Qian , Zhenglu Yang

Voice conversion (VC) systems can transform audio to mimic another speaker's voice, thereby attacking speaker verification (SV) systems. However, ongoing studies on source speaker verification (SSV) are hindered by limited data availability…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Ze Li , Yuke Lin , Tian Yao , Hongbin Suo , Pengyuan Zhang , Yanzhen Ren , Zexin Cai , Hiromitsu Nishizaki , Ming Li