中文
相关论文

相关论文: Joint speech and overlap detection: a benchmark ov…

200 篇论文

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even…

音频与语音处理 · 电气工程与系统科学 2023-06-02 Tobias Cord-Landwehr , Christoph Boeddeker , Cătălin Zorilă , Rama Doddipatla , Reinhold Haeb-Umbach

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speech activities of target and non-target speakers as clues,…

音频与语音处理 · 电气工程与系统科学 2025-06-17 Shota Horiguchi , Takanori Ashihara , Marc Delcroix , Atsushi Ando , Naohiro Tawara

Open-vocabulary segmentation is the task of segmenting anything that can be named in an image. Recently, large-scale vision-language modelling has led to significant advances in open-vocabulary segmentation, but at the cost of gargantuan…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Laurynas Karazija , Iro Laina , Andrea Vedaldi , Christian Rupprecht

Open-set learning and discovery (OSLD) is a challenging machine learning task in which samples from new (unknown) classes can appear at test time. It can be seen as a generalization of zero-shot learning, where the new classes are not known…

To better model the contextual information and increase the generalization ability of Speech Activity Detection (SAD) system, this paper leverages a multi-lingual Automatic Speech Recognition (ASR) system to perform SAD. Sequence…

声音 · 计算机科学 2021-04-13 Seyyed Saeed Sarfjoo , Srikanth Madikeri , Petr Motlicek

Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often involve further fine-tuning of VL-PTMs, which may distort their general knowledge and impair…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Lin Zhu , Yifeng Yang , Qinying Gu , Xinbing Wang , Chenghu Zhou , Nanyang Ye

Sound Event Localization and Detection (SELD) is crucial in spatial audio processing, enabling systems to detect sound events and estimate their 3D directions. Existing SELD methods use single- or dual-branch architectures: single-branch…

声音 · 计算机科学 2025-07-31 Hogeon Yu

Visual voice activity detection (V-VAD) uses visual features to predict whether a person is speaking or not. V-VAD is useful whenever audio VAD (A-VAD) is inefficient either because the acoustic signal is difficult to analyze or because it…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Sylvain Guy , Stéphane Lathuilière , Pablo Mesejo , Radu Horaud

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Shota Horiguchi , Takafumi Moriya , Atsushi Ando , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

Since its introduction in 2019, the whole end-to-end neural diarization (EEND) line of work has been addressing speaker diarization as a frame-wise multi-label classification problem with permutation-invariant training. Despite EEND showing…

声音 · 计算机科学 2023-10-23 Alexis Plaquet , Hervé Bredin

Mismatching problem between the source and target noisy corpora severely hinder the practical use of the machine-learning-based voice activity detection (VAD). In this paper, we try to address this problem in the transfer learning…

机器学习 · 计算机科学 2013-03-11 Xiao-Lei Zhang , Ji Wu

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Yong Liu , SongLi Wu , Sule Bai , Jiahao Wang , Yitong Wang , Yansong Tang

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corruption, an audio-visual…

音频与语音处理 · 电气工程与系统科学 2023-07-07 Guinan Li , Jiajun Deng , Mengzhe Geng , Zengrui Jin , Tianzi Wang , Shujie Hu , Mingyu Cui , Helen Meng , Xunying Liu

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any…

音频与语音处理 · 电气工程与系统科学 2023-10-04 Samuele Cornell , Jee-weon Jung , Shinji Watanabe , Stefano Squartini

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal…

Speech recognition systems are often highly domain dependent, a fact widely reported in the literature. However the concept of domain is complex and not bound to clear criteria. Hence it is often not evident if data should be considered to…

计算与语言 · 计算机科学 2015-09-23 Mortaza Doulaty , Oscar Saz , Thomas Hain

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification…

音频与语音处理 · 电气工程与系统科学 2020-03-16 Takenori Yoshimura , Tomoki Hayashi , Kazuya Takeda , Shinji Watanabe

For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach a preset maximum duration before segmentation, resulting in…

音频与语音处理 · 电气工程与系统科学 2023-05-23 Mohan Shi , Yuchun Shu , Lingyun Zuo , Qian Chen , Shiliang Zhang , Jie Zhang , Li-Rong Dai

Pre-trained Transformer-based speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem of domain mismatch…

计算与语言 · 计算机科学 2023-12-13 Mohammed Maqsood Shaik , Dietrich Klakow , Badr M. Abdullah

Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for transparency of…

音频与语音处理 · 电气工程与系统科学 2024-01-18 Théo Mariotte , Antonio Almudévar , Marie Tahon , Alfonso Ortega
‹ 上一页 1 8 9 10 下一页 ›