中文
相关论文

相关论文: Triplet loss based embeddings for forensic speaker…

200 篇论文

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker embeddings to enhance the performance of the phrasing model. We…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Dong Yang , Yuki Saito , Takaaki Saeki , Tomoki Koriyama , Wataru Nakata , Detai Xin , Hiroshi Saruwatari

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.…

Text-to-speech (TTS) acoustic models map linguistic features into an acoustic representation out of which an audible waveform is generated. The latest and most natural TTS systems build a direct mapping between linguistic and waveform…

声音 · 计算机科学 2019-09-24 David Álvarez , Santiago Pascual , Antonio Bonafonte

Accurately modeling idiomatic or non-compositional language has been a longstanding challenge in Natural Language Processing (NLP). This is partly because these expressions do not derive their meanings solely from their constituent words,…

计算与语言 · 计算机科学 2024-09-06 Wei He , Marco Idiart , Carolina Scarton , Aline Villavicencio

Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. The embeddings are commonly used to classify and discriminate…

音频与语音处理 · 电气工程与系统科学 2023-02-07 Adriana Stan

This paper proposes a speech rhythm-based method for speaker embeddings to model phoneme duration using a few utterances by the target speaker. Speech rhythm is one of the essential factors among speaker characteristics, along with acoustic…

声音 · 计算机科学 2024-02-13 Kenichi Fujita , Atsushi Ando , Yusuke Ijima

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to…

声音 · 计算机科学 2019-07-03 Miquel India , Pooyan Safari , Javier Hernando

Deep speaker embeddings have been shown effective for assessing cognitive impairments aside from their original purpose of speaker verification. However, the research found that speaker embeddings encode speaker identity and an array of…

音频与语音处理 · 电气工程与系统科学 2022-03-22 Dongseok Heo , Cheul Young Park , Jaemin Cheun , Myung Jin Ko

We present progress towards bilingual Text-to-Speech which is able to transform a monolingual voice to speak a second language while preserving speaker voice quality. We demonstrate that a bilingual speaker embedding space contains a…

计算与语言 · 计算机科学 2020-04-13 Soumi Maiti , Erik Marchi , Alistair Conkie

Text-independent speaker verification is an important artificial intelligence problem that has a wide spectrum of applications, such as criminal investigation, payment certification, and interest-based customer services. The purpose of…

音频与语音处理 · 电气工程与系统科学 2020-07-22 Jiwei Xu , Xinggang Wang , Bin Feng , Wenyu Liu

In speaker anonymization, speech recordings are modified in a way that the identity of the speaker remains hidden. While this technology could help to protect the privacy of individuals around the globe, current research restricts this by…

计算与语言 · 计算机科学 2024-10-08 Sarina Meyer , Florian Lux , Ngoc Thang Vu

The state-of-art approach for speaker verification consists of a neural network based embedding extractor along with a backend generative model such as the Probabilistic Linear Discriminant Analysis (PLDA). In this work, we propose a neural…

音频与语音处理 · 电气工程与系统科学 2020-05-26 Shreyas Ramoji , Prashant Krishnan , Sriram Ganapathy

This paper describes one objective function for learning semantically coherent feature embeddings in multi-output classification problems, i.e., when the response variables have dimension higher than one. In particular, we consider the…

计算机视觉与模式识别 · 计算机科学 2020-03-23 Hugo Proença , Ehsan Yaghoubi , Pendar Alirezazadeh

Speech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved when audio and…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We consider several class-agnostic semantic constraints that apply…

Self-supervised learning models for speech processing, such as wav2vec2, HuBERT, WavLM, and Whisper, generate embeddings that capture both linguistic and paralinguistic information, making it challenging to analyze tone independently of…

机器学习 · 计算机科学 2025-02-27 Hamdan Al Ahbabi , Gautier Marti , Saeed AlMarri , Ibrahim Elfadel

Automatic cover detection -- the task of finding in a audio dataset all covers of a query track -- has long been a challenging theoretical problem in MIR community. It also became a practical need for music composers societies requiring to…

机器学习 · 计算机科学 2020-04-10 Guillaume Doras , Geoffroy Peeters

Spoofing detection systems are typically trained using diverse recordings from multiple speakers, often assuming that the resulting embeddings are independent of speaker identity. However, this assumption remains unverified. In this paper,…

声音 · 计算机科学 2026-02-25 Anh-Tuan Dao , Driss Matrouf , Nicholas Evans

Training multilingual Neural Text-To-Speech (NTTS) models using only monolingual corpora has emerged as a popular way for building voice cloning based Polyglot NTTS systems. In order to train these models, it is essential to understand how…

音频与语音处理 · 电气工程与系统科学 2022-07-05 Ziyao Zhang , Alessio Falai , Ariadna Sanchez , Orazio Angelini , Kayoko Yanagisawa

Speaker embeddings are widely used in speaker verification systems and other applications where it is useful to characterise the voice of a speaker with a fixed-length vector. These embeddings tend to be treated as "black box" encodings,…

声音 · 计算机科学 2025-10-21 Mark Huckvale