English
Related papers

Related papers: Light Convolutional Neural Network with Feature Ge…

200 papers

The increased adoption of digital assistants makes text-to-speech (TTS) synthesis systems an indispensable feature of modern mobile devices. It is hence desirable to build a system capable of generating highly intelligible speech in the…

Sound · Computer Science 2020-08-14 Dipjyoti Paul , Muhammed PV Shifas , Yannis Pantazis , Yannis Stylianou

In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Amirsina Torfi , Jeremy Dawson , Nasser M. Nasrabadi

AI-synthesized speech, also known as deepfake speech, has recently raised significant concerns due to the rapid advancement of speech synthesis and speech conversion techniques. Previous works often rely on distinguishing synthesizer…

Sound · Computer Science 2024-11-15 Kuiyuan Zhang , Zhongyun Hua , Yushu Zhang , Yifang Guo , Tao Xiang

Automatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-16 Haibin Wu , Xu Li , Andy T. Liu , Zhiyong Wu , Helen Meng , Hung-yi Lee

Voice spoofing attacks pose a significant threat to automated speaker verification systems. Existing anti-spoofing methods often simulate specific attack types, such as synthetic or replay attacks. However, in real-world scenarios, the…

Sound · Computer Science 2023-09-19 Awais Khan , Khalid Mahmood Malik , Shah Nawaz

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-16 Haibin Wu , Heng-Cheng Kuo , Naijun Zheng , Kuo-Hsuan Hung , Hung-Yi Lee , Yu Tsao , Hsin-Min Wang , Helen Meng

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Ivan Kukanov , Jun Wah Ng

Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-07 Lin Zhang , Xin Wang , Erica Cooper , Nicholas Evans , Junichi Yamagishi

This paper addresses source tracing in synthetic speech-identifying generative systems behind manipulated audio via speaker recognition-inspired pipelines. While prior work focuses on spoofing detection, source tracing lacks robust…

The rapid advancement of deep learning models that can generate and synthesis hyper-realistic videos known as Deepfakes and their ease of access to the general public have raised concern from all concerned bodies to their possible malicious…

Computer Vision and Pattern Recognition · Computer Science 2021-03-12 Deressa Wodajo , Solomon Atnafu

Generative adversarial networks have seen rapid development in recent years and have led to remarkable improvements in generative modelling of images. However, their application in the audio domain has received limited attention, and…

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this…

In this work we evaluate the utility of synthetic data for training automatic speech recognition (ASR). We use the ASR training data to train a text-to-speech (TTS) system similar to FastSpeech-2. With this TTS we reproduce the original…

Computation and Language · Computer Science 2024-10-29 Benedikt Hilmes , Nick Rossenbach , and Ralf Schlüter

Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in realistic crowded…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-29 Sherif Abdulatif , Karim Armanious , Karim Guirguis , Jayasankar T. Sajeev , Bin Yang

Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the coverage of these corpora. We believe the ability to handle…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-09 Li-Wei Chen , Shinji Watanabe , Alexander Rudnicky

Most current speech technology systems are designed to operate well even in the presence of multiple active speakers. However, most solutions assume that the number of co-current speakers is known. Unfortunately, this information might not…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-02 Midia Yousefi , John H. L. Hansen

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored…

ASV (automatic speaker verification) systems are intrinsically required to reject both non-target (e.g., voice uttered by different speaker) and spoofed (e.g., synthesised or converted) inputs. However, there is little consideration for how…

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these…

Sound · Computer Science 2025-10-13 Huu Tuong Tu , Huan Vu , cuong tien nguyen , Dien Hy Ngo , Nguyen Thi Thu Trang

Incorporating cross-speaker style transfer in text-to-speech (TTS) models is challenging due to the need to disentangle speaker and style information in audio. In low-resource expressive data scenarios, voice conversion (VC) can generate…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Lucas H. Ueda , Leonardo B. de M. M. Marques , Flávio O. Simões , Mário U. Neto , Fernando Runstein , Bianca Dal Bó , Paula D. P. Costa