English
Related papers

Related papers: Wanna hear your voice? A sample is all we need!

200 papers

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-20 Sefik Emre Eskimez , Takuya Yoshioka , Huaming Wang , Xiaofei Wang , Zhuo Chen , Xuedong Huang

Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Ze Li , Xiaoxiao Miao , Juan Liu , Ming Li

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction (TSE) or Target…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-20 Bang Zeng , Ming Li

Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal content separately,…

Sound · Computer Science 2025-11-18 Bingsong Bai , Yizhong Geng , Fengping Wang , Cong Wang , Puyuan Guo , Yingming Gao , Ya Li

We describe our submitted system for the ZeroSpeech Challenge 2019. The current challenge theme addresses the difficulty of constructing a speech synthesizer without any text or phonetic labels and requires a system that can (1) discover…

Computation and Language · Computer Science 2019-05-30 Andros Tjandra , Berrak Sisman , Mingyang Zhang , Sakriani Sakti , Haizhou Li , Satoshi Nakamura

In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often fall short,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-27 Tassadaq Hussain , Kia Dashtipour , Yu Tsao , Amir Hussain

Recent advancements in text-to-speech (TTS) technology have increased demand for personalized audio synthesis. Zero-shot voice cloning, a specialized TTS task, aims to synthesize a target speaker's voice using only a single audio sample and…

Sound · Computer Science 2025-06-03 Ming Meng , Ziyi Yang , Jian Yang , Zhenjie Su , Yonggui Zhu , Zhaoxin Fan

Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze several prominent models such as Whisper, Seamless, Wav2Vec,…

Computation and Language · Computer Science 2024-10-18 Abdul Waheed , Hanin Atwany , Bhiksha Raj , Rita Singh

Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content. Although recent systems have improved conversion quality, building zero-shot VC systems…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-23 Qixi Zheng , Yuxiang Zhao , Tianrui Wang , Wenxi Chen , Kele Xu , Yikang Li , Qinyuan Chen , Xipeng Qiu , Kai Yu , Xie Chen

In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-11 Yuepeng Jiang , Ziqian Ning , Shuai Wang , Chengjia Wang , Mengxiao Bi , Pengcheng Zhu , Zhonghua Fu , Lei Xie

Numerous methods have been proposed to enhance Keyword Spotting (KWS) in adult speech, but children's speech presents unique challenges for KWS systems due to its distinct acoustic and linguistic characteristics. This paper introduces a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Subham Kutum , Abhijit Sinha , Hemant Kumar Kathania , Sudarsana Reddy Kadiri , Mahesh Chandra Govil

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses significant challenges…

Computation and Language · Computer Science 2025-03-31 Heqing Zou , Fengmao Lv , Desheng Zheng , Eng Siong Chng , Deepu Rajan

Recent advancements in speech synthesis have enabled large language model (LLM)-based systems to perform zero-shot generation with controllable content, timbre, speaker identity, and emotion through input prompts. As a result, these models…

We consider hate speech detection through keyword spotting on radio broadcasts. One approach is to build an automatic speech recognition (ASR) system for the target low-resource language. We compare this to using acoustic word embedding…

Computation and Language · Computer Science 2023-06-02 Christiaan Jacobs , Nathanaël Carraz Rakotonirina , Everlyn Asiko Chimoto , Bruce A. Bassett , Herman Kamper

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been underexplored.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-20 Leying Zhang , Wangyou Zhang , Chenda Li , Yanmin Qian

Detecting hate speech, especially in low-resource languages, is a non-trivial challenge. To tackle this, we developed a tailored architecture based on frozen, pre-trained Transformers to examine cross-lingual zero-shot and few-shot…

Computation and Language · Computer Science 2020-04-30 Lukas Stappen , Fabian Brunn , Björn Schuller

Despite rapid progress in the voice style transfer (VST) field, recent zero-shot VST systems still lack the ability to transfer the voice style of a novel speaker. In this paper, we present HierVST, a hierarchical adaptive end-to-end…

Sound · Computer Science 2023-08-01 Sang-Hoon Lee , Ha-Yeong Choi , Hyung-Seok Oh , Seong-Whan Lee

Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Rui Liu , Pu Gao , Jiatian Xi , Berrak Sisman , Carlos Busso , Haizhou Li

In this work, we address the problem of binaural target-speaker extraction in the presence of multiple simultane-ous talkers. We propose a novel approach that leverages the individual listener's Head-Related Transfer Function (HRTF) to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-25 Yoav Ellinson , Sharon Gannot

Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can hardly perform zero-shot due to incomplete feature…

Sound · Computer Science 2024-11-18 Zihao Wang , Le Ma , Yongsheng Feng , Xin Pan , Yuhang Jin , Kejun Zhang