中文
相关论文

相关论文: Personal VAD: Speaker-Conditioned Voice Activity D…

200 篇论文

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are…

音频与语音处理 · 电气工程与系统科学 2022-03-22 Shehzeen Hussain , Van Nguyen , Shuhua Zhang , Erik Visser

Robust Voice Activity Detection (VAD) remains a challenging task, especially under noisy, diverse, and unseen acoustic conditions. Beyond algorithmic development, a key limitation in advancing VAD research is the lack of large-scale,…

Usable speech criteria are proposed to extract minimally corrupted speech for speaker identification (SID) in co-channel speech. In co-channel speech, either speaker can randomly appear as the stronger speaker or the weaker one at a time.…

声音 · 计算机科学 2013-01-03 Wajdi Ghezaiel , Amel Ben Slimane , Ezzedine Ben Braiek

Speech activity detection (SAD), which often rests on the fact that the noise is "more" stationary than speech, is particularly challenging in non-stationary environments, because the time variance of the acoustic scene makes it difficult…

音频与语音处理 · 电气工程与系统科学 2020-07-29 Jens Heitkaemper , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of…

计算机视觉与模式识别 · 计算机科学 2018-06-08 Amirsina Torfi , Jeremy Dawson , Nasser M. Nasrabadi

In many VoIP systems, Voice Activity Detection (VAD) is often used on VoIP traffic to suppress packets of silence in order to reduce the bandwidth consumption of phone calls. Unfortunately, although VoIP traffic is fully encrypted and…

密码学与安全 · 计算机科学 2023-06-02 Roy Laurens , Edo Christianto , Bruce Caulkins , Cliff C. Zou

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

计算机视觉与模式识别 · 计算机科学 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic…

音频与语音处理 · 电气工程与系统科学 2021-06-16 Rajeev Rikhye , Quan Wang , Qiao Liang , Yanzhang He , Ding Zhao , Yiteng , Huang , Arun Narayanan , Ian McGraw

The accuracy of automated speaker recognition is negatively impacted by change in emotions in a person's speech. In this paper, we hypothesize that speaker identity is composed of various vocal style factors that may be learned from…

音频与语音处理 · 电气工程与系统科学 2023-08-04 Morgan Sandler , Arun Ross

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

图像与视频处理 · 电气工程与系统科学 2022-09-27 Rahul Sharma , Shrikanth Narayanan

Direct speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice activity detector (VAD). Since VAD segmentation is not…

计算与语言 · 计算机科学 2020-08-06 Marco Gaido , Mattia Antonino Di Gangi , Matteo Negri , Mauro Cettolo , Marco Turchi

This paper proposes a multi-task learning network with phoneme-aware and channel-wise attentive learning strategies for text-dependent Speaker Verification (SV). In the proposed structure, the frame-level multi-task learning along with the…

声音 · 计算机科学 2021-06-28 Yan Liu , Zheng Li , Lin Li , Qingyang Hong

Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2vec 2.0. In this…

音频与语音处理 · 电气工程与系统科学 2023-05-10 Marie Kunešová , Zbyněk Zajíc

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks.…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Santi Prieto , Alfonso Ortega , Iván López-Espejo , Eduardo Lleida

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

计算与语言 · 计算机科学 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain

This paper addresses the problem of Target Activity Detection (TAD) for binaural listening devices. TAD denotes the problem of robustly detecting the activity of a target speaker in a harsh acoustic environment, which comprises interfering…

声音 · 计算机科学 2016-12-21 Daniel Gerber , Stefan Meier , Walter Kellermann

Detecting anchor's voice in live musical streams is an important preprocessing for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to…

声音 · 计算机科学 2020-11-03 Yuanbo Hou , Yi Deng , Bilei Zhu , Zejun Ma , Dick Botteldooren

We propose SpeakerNet - a new neural architecture for speaker recognition and speaker verification tasks. It is composed of residual blocks with 1D depth-wise separable convolutions, batch-normalization, and ReLU layers. This architecture…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Nithin Rao Koluguri , Jason Li , Vitaly Lavrukhin , Boris Ginsburg