English
Related papers

Related papers: An Integrated Framework for Two-pass Personalized …

200 papers

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

This paper describes the TSUP team's submission to the ISCSLP 2022 conversational short-phrase speaker diarization (CSSD) challenge which particularly focuses on short-phrase conversations with a new evaluation metric called conversational…

Sound · Computer Science 2023-10-26 Bowen Pang , Huan Zhao , Gaosheng Zhang , Xiaoyue Yang , Yang Sun , Li Zhang , Qing Wang , Lei Xie

In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-20 Quan Wang , Hannah Muckenhirn , Kevin Wilson , Prashant Sridhar , Zelin Wu , John Hershey , Rif A. Saurous , Ron J. Weiss , Ye Jia , Ignacio Lopez Moreno

This paper describes our proposed integration system for the spoofing-aware speaker verification challenge. It consists of a robust spoofing-aware verification system that use the speaker verification and antispoofing embeddings extracted…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-05 Juan M. Martín-Doñas , Iván G. Torre , Aitor Álvarez , Joaquin Arellano

User-defined keyword spotting (KWS) is crucial for personalized voice interaction, yet existing methods face several challenges: (1) insufficient discriminability among confusable words, (2) performance inconsistency across speakers with…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-22 Zhiqi Ai , Han Cheng , Shiyi Mu , Xinnuo Li , Yongjin Zhou , Shugong Xu

The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Santi Prieto , Alfonso Ortega , Iván López-Espejo , Eduardo Lleida

This paper describes the Academia Sinica systems for the two tasks of Voice Conversion Challenge 2020, namely voice conversion within the same language (Task 1) and cross-lingual voice conversion (Task 2). For both tasks, we followed the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-07 Yu-Huai Peng , Cheng-Hung Hu , Alexander Kang , Hung-Shin Lee , Pin-Yuan Chen , Yu Tsao , Hsin-Min Wang

In this technical report we describe the IDLAB top-scoring submissions for the VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20) in the supervised and unsupervised speaker verification tracks. For the supervised verification tracks we…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

This report describes the SJTU-AISPEECH system for the Voxceleb Speaker Recognition Challenge 2022. For track1, we implemented two kinds of systems, the online system and the offline system. Different ResNet-based backbones and loss…

Sound · Computer Science 2022-09-21 Zhengyang Chen , Bing Han , Xu Xiang , Houjun Huang , Bei Liu , Yanmin Qian

Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbone networks for speaker embedding are designed to aggregate…

Sound · Computer Science 2022-03-18 Ruiteng Zhang , Jianguo Wei , Xugang Lu , Wenhuan Lu , Di Jin , Junhai Xu , Lin Zhang , Yantao Ji , Jianwu Dang

Data augmentation is conventionally used to inject robustness in Speaker Verification systems. Several recently organized challenges focus on handling novel acoustic environments. Deep learning based speech enhancement is a modern solution…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-29 Saurabh Kataria , Phani Sankar Nidadavolu , Jesús Villalba , Najim Dehak

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech.…

In this paper, we propose pass-phrase dependent background models (PBMs) for text-dependent (TD) speaker verification (SV) to integrate the pass-phrase identification process into the conventional TD-SV system, where a PBM is derived from a…

Computation and Language · Computer Science 2021-03-29 A. K. Sarkar , Zheng-Hua Tan

A Spoken dialogue system for an unseen language is referred to as Zero resource speech. It is especially beneficial for developing applications for languages that have low digital resources. Zero resource speech synthesis is the task of…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-11 Karthik Pandia D S , Anusha Prakash , Mano Ranjith Kumar , Hema A Murthy

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was performed on…

Computation and Language · Computer Science 2025-09-30 Bassam Matar , Mohamed Fayed , Ayman Khalafallah

In this paper, we propose MM-KWS, a novel approach to user-defined keyword spotting leveraging multi-modal enrollments of text and speech templates. Unlike previous methods that focus solely on either text or speech features, MM-KWS…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Zhiqi Ai , Zhiyong Chen , Shugong Xu

This report describes the NPU-HC speaker verification system submitted to the O-COCOSDA Multi-lingual Speaker Verification (MSV) Challenge 2022, which focuses on developing speaker verification systems for low-resource Asian languages. We…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-05 Yue Li , Li Zhang , Namin Wang , Jie Liu , Lei Xie

Single-stage text-to-speech models have been actively studied recently, and their results have outperformed two-stage pipeline systems. Although the previous single-stage model has made great progress, there is room for improvement in terms…

Sound · Computer Science 2023-08-01 Jungil Kong , Jihoon Park , Beomjeong Kim , Jeongmin Kim , Dohee Kong , Sangjin Kim

In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to transcribe the source speech into the underlying linguistic…

Sound · Computer Science 2021-07-21 Wen-Chin Huang , Tomoki Hayashi , Xinjian Li , Shinji Watanabe , Tomoki Toda

Capturing long-range dependency and modeling long temporal contexts is proven to benefit speaker verification tasks. In this paper, we propose the combination of the Hierarchical-Split block(HS-block) and the Depthwise Separable…

Sound · Computer Science 2021-12-15 Zhuo Li
‹ Prev 1 8 9 10 Next ›