English
Related papers

Related papers: TokenSplit: Using Discrete Speech Representations …

200 papers

The popular frameworks for self-supervised learning of speech representations have largely focused on frame-level masked prediction of speech regions. While this has shown promising downstream task performance for speech recognition and…

Computation and Language · Computer Science 2025-07-22 Varun Krishna , Sriram Ganapathy

Self-supervised representation learning approaches have grown in popularity due to the ability to train models on large amounts of unlabeled data and have demonstrated success in diverse fields such as natural language processing, computer…

Machine Learning · Computer Science 2023-02-06 John Harvill , Jarred Barber , Arun Nair , Ramin Pishehvar

Recent advancements in speech-based topic segmentation have highlighted the potential of pretrained speech encoders to capture semantic representations directly from speech. Traditionally, topic segmentation has relied on a pipeline…

Computation and Language · Computer Science 2024-09-11 Sakshi Deo Shukla , Pavel Denisov , Tugtekin Turan

Recent advancements in speech synthesis witness significant benefits by leveraging discrete tokens extracted from self-supervised learning (SSL) models. Discrete tokens offer higher storage efficiency and greater operability in intermediate…

Sound · Computer Science 2024-06-21 Yuning Wu , Chunlei zhang , Jiatong Shi , Yuxun Tang , Shan Yang , Qin Jin

End-to-end simultaneous speech translation (SimulST) outputs translation while receiving the streaming speech inputs (a.k.a. streaming speech translation), and hence needs to segment the speech inputs and then translate based on the current…

Computation and Language · Computer Science 2023-11-13 Shaolei Zhang , Yang Feng

Recent research shows end-to-end ASR systems can recognize overlapped speech from multiple speakers. However, all published works have assumed no latency constraints during inference, which does not hold for most voice assistant…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-22 Ilya Sklyar , Anna Piunova , Yulan Liu

Multi-talker conversational speech processing has drawn many interests for various applications such as meeting transcription. Speech separation is often required to handle overlapped speech that is commonly observed in conversation.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-18 Wangyou Zhang , Zhuo Chen , Naoyuki Kanda , Shujie Liu , Jinyu Li , Sefik Emre Eskimez , Takuya Yoshioka , Xiong Xiao , Zhong Meng , Yanmin Qian , Furu Wei

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-18 Julian Neri , Sebastian Braun

In this paper, we introduce an unsupervised approach for Speech Segmentation, which builds on previously researched approaches, e.g., Speaker Diarization, while being applicable to an inclusive set of acoustic-semantic distinctions, paving…

Computation and Language · Computer Science 2025-01-08 Avishai Elmakies , Omri Abend , Yossi Adi

Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches often rely on pretrained encoders, semantic distillation,…

Discrete audio tokens derived from self-supervised learning models have gained widespread usage in speech generation. However, current practice of directly utilizing audio tokens poses challenges for sequence modeling due to the length of…

Sound · Computer Science 2024-01-17 Feiyu Shen , Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

Speech separation has been studied widely for single-channel close-talk microphone recordings over the past few years; developed solutions are mostly in frequency-domain. Recently, a raw audio waveform separation network (TasNet) is…

Sound · Computer Science 2019-07-25 Fahimeh Bahmaninezhad , Jian Wu , Rongzhi Gu , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, underscores the importance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Renana Opochinsky , Mordehay Moradi , Sharon Gannot

Previous methods for audio-image matching generally fall into one of two categories: pipeline models or End-to-End models. Pipeline models first transcribe speech and then encode the resulting text; End-to-End models encode speech directly.…

Sound · Computer Science 2024-08-21 Zhenyu Lu , Lakshay Sethi

In this paper, we propose a novel approach for the transcription of speech conversations with natural speaker overlap, from single channel speech recordings. The proposed model is a combination of a speaker diarization system and a hybrid…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-30 Srikanth Raj Chetupalli , Sriram Ganapathy

TV subtitles are a rich source of transcriptions of many types of speech, ranging from read speech in news reports to conversational and spontaneous speech in talk shows and soaps. However, subtitles are not verbatim (i.e. exact)…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-17 Jakob Poncelet , Hugo Van hamme

Speech separation has been successfully applied as a frontend processing module of conversation transcription systems thanks to its ability to handle overlapped speech and its flexibility to combine with downstream tasks such as automatic…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-06 Jian Wu , Zhuo Chen , Sanyuan Chen , Yu Wu , Takuya Yoshioka , Naoyuki Kanda , Shujie Liu , Jinyu Li

Speech restoration aims at restoring high quality speech in the presence of a diverse set of distortions. Although several deep learning paradigms have been studied for this task, the power of the recently emerging language models has not…

Sound · Computer Science 2024-06-05 Xu Li , Qirui Wang , Xiaoyu Liu

Data-driven speech processing models usually perform well with a large amount of text supervision, but collecting transcribed speech data is costly. Therefore, we propose SpeechCLIP, a novel framework bridging speech and text through images…

Computation and Language · Computer Science 2022-10-26 Yi-Jen Shih , Hsuan-Fu Wang , Heng-Jui Chang , Layne Berry , Hung-yi Lee , David Harwath

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited…