中文
相关论文

相关论文: Text-Independent Speaker Verification Using Long S…

200 篇论文

The x-vector based deep neural network (DNN) embedding systems have demonstrated effectiveness for text-independent speaker verification. This paper presents a multi-task learning architecture for training the speaker embedding DNN with the…

音频与语音处理 · 电气工程与系统科学 2019-04-05 Lanhua You , Wu Guo , Lirong Dai , Jun Du

This document describes the Short-duration Speaker Verification (SdSV) Challenge 2021. The main goal of the challenge is to evaluate new technologies for text-dependent (TD) and text-independent (TI) speaker verification (SV) in a short…

音频与语音处理 · 电气工程与系统科学 2021-03-26 Hossein Zeinali , Kong Aik Lee , Jahangir Alam , Lukas Burget

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no…

音频与语音处理 · 电气工程与系统科学 2019-04-16 Takuya Yoshioka , Zhuo Chen , Changliang Liu , Xiong Xiao , Hakan Erdogan , Dimitrios Dimitriadis

We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows…

音频与语音处理 · 电气工程与系统科学 2025-03-25 Luan Vinícius Fiorio , Bruno Defraene , Johan David , Alex Young , Frans Widdershoven , Wim van Houtum , Ronald M. Aarts

Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and…

声音 · 计算机科学 2018-04-19 Yi Luo , Zhuo Chen , Nima Mesgarani

Speaker embedding is an important front-end module to explore discriminative speaker features for many speech applications where speaker information is needed. Current SOTA backbone networks for speaker embedding are designed to aggregate…

声音 · 计算机科学 2022-03-18 Ruiteng Zhang , Jianguo Wei , Xugang Lu , Wenhuan Lu , Di Jin , Junhai Xu , Lin Zhang , Yantao Ji , Jianwu Dang

Verifying the identity of a speaker is crucial in modern human-machine interfaces, e.g., to ensure privacy protection or to enable biometric authentication. Classical speaker verification (SV) approaches estimate a fixed-dimensional…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Ahmad Aloradi , Wolfgang Mack , Mohamed Elminshawi , Emanuël A. P. Habets

Long short-term memory(LSTM) units on sequence-based models are being used in translation, question-answering systems, classification tasks due to their capability of learning long-term dependencies. In Natural language generation, LSTM…

计算与语言 · 计算机科学 2020-05-04 Sivasurya Santhanam

This paper proposes a deep multi-speaker text-to-speech (TTS) model for spoofing speaker verification (SV) systems. The proposed model employs one network to synthesize time-downsampled mel-spectrograms from text input and another network…

音频与语音处理 · 电气工程与系统科学 2019-10-30 Mingrui Yuan , Zhiyao Duan

Recently several end-to-end speaker verification systems based on deep neural networks (DNNs) have been proposed. These systems have been proven to be competitive for text-dependent tasks as well as for text-independent tasks with short…

音频与语音处理 · 电气工程与系统科学 2018-01-09 Johan Rohdin , Anna Silnova , Mireia Diez , Oldrich Plchot , Pavel Matejka , Lukas Burget

This paper presents our submission to the Iranian division of the Text-Dependent Speaker Verification Challenge (TdSV) 2024. Conventional TdSV approaches typically jointly model speaker and linguistic features, requiring unsegmented inputs…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Seyed Ali Farokh , Hossein Zeinali

Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under…

音频与语音处理 · 电气工程与系统科学 2024-02-06 Kerlos Atia Abdalmalak , Ascensión Gallardo-Antol'in

A number of studies have successfully developed speaker verification or presentation attack detection systems. However, studies integrating the two tasks remain in the preliminary stages. In this paper, we propose two approaches for…

音频与语音处理 · 电气工程与系统科学 2020-09-29 Hye-jin Shim , Jee-weon Jung , Ju-ho Kim , Seung-bin Kim , Ha-Jin Yu

The aim of this paper is to investigate the benefit of combining both language and acoustic modelling for speaker diarization. Although conventional systems only use acoustic features, in some scenarios linguistic data contain high…

音频与语音处理 · 电气工程与系统科学 2025-01-31 Miquel India , Javier Hernando , José A. R. Fonollosa

In this work, we introduce metric learning (ML) to enhance the deep embedding learning for text-independent speaker verification (SV). Specifically, the deep speaker embedding network is trained with conventional cross entropy loss and…

音频与语音处理 · 电气工程与系统科学 2023-03-24 Yafeng Chen , Wu Guo , Jingjing Shi , Jiajun Qi , Tan Liu

Learning-based Text To Speech systems have the potential to generalize from one speaker to the next and thus require a relatively short sample of any new voice. However, this promise is currently largely unrealized. We present a method that…

机器学习 · 计算机科学 2018-02-21 Eliya Nachmani , Adam Polyak , Yaniv Taigman , Lior Wolf

As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization,…

声音 · 计算机科学 2026-05-15 KiHyun Nam , Jungwoo Heo , Siu Bae , Ha-Jin Yu , Joon Son Chung

Large Language Models (LLMs) are one of the most promising technologies for the next era of speech generation systems, due to their scalability and in-context learning capabilities. Nevertheless, they suffer from multiple stability issues…

Speaker-conditioned target speaker extraction systems rely on auxiliary information about the target speaker to extract the target speaker signal from a mixture of multiple speakers. Typically, a deep neural network is applied to isolate…

音频与语音处理 · 电气工程与系统科学 2021-04-12 Ragini Sinha , Marvin Tammen , Christian Rollwage , Simon Doclo

Transcribing and understanding multi-speaker conversations requires speech recognition, speaker attribution, and timestamp localization. While speech LLMs excel at single-speaker tasks, multi-speaker scenarios remain challenging due to…

音频与语音处理 · 电气工程与系统科学 2026-04-06 Zhennan Lin , Shuai Wang , Zhaokai Sun , Pengyuan Xie , Chuan Xie , Jie Liu , Qiang Zhang , Lei Xie