中文
相关论文

相关论文: Triplet loss based embeddings for forensic speaker…

200 篇论文

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

声音 · 计算机科学 2021-02-11 Zeqian Li , Jacob Whitehill

Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inter-personal variability among speakers, and the complex…

计算机视觉与模式识别 · 计算机科学 2025-02-18 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Forensic audio analysis for speaker verification offers unique challenges due to location/scenario uncertainty and diversity mismatch between reference and naturalistic field recordings. The lack of real naturalistic forensic audio corpora…

音频与语音处理 · 电气工程与系统科学 2020-09-10 Zhenyu Wang , Wei Xia , John H. L. Hansen

Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic…

声音 · 计算机科学 2025-06-04 Aditya Srinivas Menon , Raj Prakash Gohil , Kumud Tripathi , Pankaj Wasnik

Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained to distinguish…

计算与语言 · 计算机科学 2022-05-20 Benjamin Kepecs , Homayoon Beigi

In forensic applications, it is very common that only small naturalistic datasets consisting of short utterances in complex or unknown acoustic environments are available. In this study, we propose a pipeline solution to improve speaker…

音频与语音处理 · 电气工程与系统科学 2020-09-22 Mufan Sang , Wei Xia , John H. L. Hansen

Audio deepfakes have reached a level of realism that makes it increasingly difficult to distinguish between human and artificial voices, which poses risks such as identity theft or spread of disinformation. Despite these concerns, research…

音频与语音处理 · 电气工程与系统科学 2025-12-11 Eugenia San Segundo , Aurora López-Jareño , Xin Wang , Junichi Yamagishi

Accented speech recognition and accent classification are relatively under-explored research areas in speech technology. Recently, deep learning-based methods and Transformer-based pretrained models have achieved superb performances in both…

计算与语言 · 计算机科学 2022-06-30 Qingcheng Zeng , Dading Chong , Peilin Zhou , Jie Yang

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically. In this work…

声音 · 计算机科学 2022-11-01 Luigi Attorresi , Davide Salvi , Clara Borrelli , Paolo Bestagini , Stefano Tubaro

To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features. Specifically, we combine spectral, spatial, and…

音频与语音处理 · 电气工程与系统科学 2021-04-12 Saurabh Kataria , Shi-Xiong Zhang , Dong Yu

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

音频与语音处理 · 电气工程与系统科学 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate…

音频与语音处理 · 电气工程与系统科学 2020-02-05 Erica Cooper , Cheng-I Lai , Yusuke Yasuda , Fuming Fang , Xin Wang , Nanxin Chen , Junichi Yamagishi

This paper addresses the problem of fake news detection in Spanish using Machine Learning techniques. It is fundamentally the same problem tackled for the English language; however, there is not a significant amount of publicly available…

计算与语言 · 计算机科学 2021-10-14 Kevin Martínez-Gallego , Andrés M. Álvarez-Ortiz , Julián D. Arias-Londoño

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas,…

声音 · 计算机科学 2025-07-18 Zhoulin Ji , Chenhao Lin , Hang Wang , Chao Shen

Different studies have shown the importance of visual cues throughout the speech perception process. In fact, the development of audiovisual approaches has led to advances in the field of speech technologies. However, although noticeable…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Automatic classification of disordered speech can provide an objective tool for identifying the presence and severity of speech impairment. Classification approaches can also help identify hard-to-recognize speech samples to teach ASR…

音频与语音处理 · 电气工程与系统科学 2021-07-09 Subhashini Venugopalan , Joel Shor , Manoj Plakal , Jimmy Tobin , Katrin Tomanek , Jordan R. Green , Michael P. Brenner

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the importance of presenting…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Speech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the…

音频与语音处理 · 电气工程与系统科学 2023-02-24 Qiongqiong Wang , Kong Aik Lee , Tianchi Liu

Speaker embeddings achieve promising results on many speaker verification tasks. Phonetic information, as an important component of speech, is rarely considered in the extraction of speaker embeddings. In this paper, we introduce phonetic…

声音 · 计算机科学 2018-06-15 Yi Liu , Liang He , Jia Liu , Michael T. Johnson

The development of privacy-preserving automatic speaker verification systems has been the focus of a number of studies with the intent of allowing users to authenticate themselves without risking the privacy of their voice. However, current…

音频与语音处理 · 电气工程与系统科学 2022-10-28 Francisco Teixeira , Alberto Abad , Bhiksha Raj , Isabel Trancoso