English
Related papers

Related papers: pyannote.audio: neural building blocks for speaker…

200 papers

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker…

Sound · Computer Science 2021-08-17 Joon Son Chung , Jaesung Huh , Arsha Nagrani , Triantafyllos Afouras , Andrew Zisserman

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-18 Mark R. Saddler , Andrew Francl , Jenelle Feather , Kaizhi Qian , Yang Zhang , Josh H. McDermott

While standard speaker diarization attempts to answer the question "who spoken when", most of relevant applications in reality are more interested in determining "who spoken what". Whether it is the conventional modularized approach or the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Yiling Huang , Weiran Wang , Guanlong Zhao , Hank Liao , Wei Xia , Quan Wang

Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. More relevant…

Sound · Computer Science 2024-06-18 Shoval Messica , Yossi Adi

Speaker clustering is an essential step in conventional speaker diarization systems and is typically addressed as an audio-only speech processing task. The language used by the participants in a conversation, however, carries additional…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-12 Nikolaos Flemotomos , Shrikanth Narayanan

As the French, European and worldwide populations are aging, there is a strong interest for new systems that guarantee a reliable and privacy preserving home monitoring for frailty prevention. This work is a part of a global environmental…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Yannis Tevissen , Dan Istrate , Vincent Zalc , Jérôme Boudy , Gérard Chollet , Frédéric Petitpont , Sami Boutamine

This research introduces an innovative voice-assisted debugging plugin for Python that transforms silent runtime errors into actionable audible diagnostics. By implementing a global exception hook architecture with pyttsx3 text-to-speech…

Automatic Speech Recognition and Text-to-Speech systems are primarily trained in a supervised fashion and require high-quality, accurately labeled speech datasets. In this work, we examine common problems with speech data and introduce a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-10 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg

We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which rely on laboriously engineered processing pipelines; these…

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

Sound · Computer Science 2021-04-26 Chau Luu , Peter Bell , Steve Renals

This paper presents a novel evaluation approach to text-based speaker diarization (SD), tackling the limitations of traditional metrics that do not account for any contextual information in text. Two new metrics are proposed, Text-based…

Computation and Language · Computer Science 2023-09-15 Chen Gong , Peilin Wu , Jinho D. Choi

Strong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously…

Sound · Computer Science 2023-06-07 Chin-Yi Cheng , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

In this paper, we propose an online speaker diarization system based on Relation Network, named RenoSD. Unlike conventional diariztion systems which consist of several independently-optimized modules, RenoSD implements…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-22 Xiang Li , Yucheng Zhao , Chong Luo , Wenjun Zeng

Audio fingerprinting is a technique used to identify and match audio recordings based on their unique characteristics. It involves creating a condensed representation of an audio signal that can be used to quickly compare and match against…

Sound · Computer Science 2023-05-03 Aarón López-García

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-05 Antonio Gomez

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Jixuan Wang , Xiong Xiao , Jian Wu , Ranjani Ramamurthy , Frank Rudzicz , Michael Brudno

Recently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis. However, current models always treat overlapped speaker diarization as a multi-label classification…

Sound · Computer Science 2022-11-21 Zhihao Du , Shiliang Zhang , Siqi Zheng , Zhijie Yan

We present a method for audio denoising that combines processing done in both the time domain and the time-frequency domain. Given a noisy audio clip, the method trains a deep neural network to fit this signal. Since the fitting is only…

Sound · Computer Science 2020-06-11 Michael Michelashvili , Lior Wolf

This paper presents our solution for the DL Sprint 4.0, addressing the dual challenges of Bengali Long-Form Speech Recognition (Task 1) and Speaker Diarization (Task 2). Processing long-form, multi-speaker Bengali audio introduces…

Sound · Computer Science 2026-03-06 Aurchi Chowdhury , Rubaiyat -E-Zaman , Sk. Ashrafuzzaman Nafees

Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, domain mismatch is still a challenging problem in this field. In this study, we propose a novel framework to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-24 Mufan Sang , Wei Xia , John H. L. Hansen