中文
相关论文

相关论文: Speech Representation Learning Revisited: The Nece…

200 篇论文

Data augmentations are known to improve robustness in speech-processing tasks. In this study, we summarize and compare different data augmentation strategies using S3PRL toolkit. We explore how HuBERT and wav2vec perform using different…

声音 · 计算机科学 2024-04-01 Mina Huh , Ruchira Ray , Corey Karnei

We present mHuBERT-147, the first general-purpose massively multilingual HuBERT speech representation model trained on 90K hours of clean, open-license data. To scale up the multi-iteration HuBERT approach, we use faiss-based clustering,…

计算与语言 · 计算机科学 2024-11-22 Marcely Zanon Boito , Vivek Iyer , Nikolaos Lagos , Laurent Besacier , Ioan Calapodescu

Transformer-based self-supervised models are trained as feature extractors and have empowered many downstream speech tasks to achieve state-of-the-art performance. However, both the training and inference process of these models may…

计算与语言 · 计算机科学 2021-05-04 Jinchuan Tian , Rongzhi Gu , Helin Wang , Yuexian Zou

Recently, the usefulness of self-supervised representation learning (SSRL) methods has been confirmed in various downstream tasks. Many of these models, as exemplified by HuBERT and WavLM, use pseudo-labels generated from spectral features…

声音 · 计算机科学 2023-10-09 Takashi Maekaku , Jiatong Shi , Xuankai Chang , Yuya Fujita , Shinji Watanabe

Unsupervised speech models are becoming ubiquitous in the speech and machine learning communities. Upstream models are responsible for learning meaningful representations from raw audio. Later, these representations serve as input to…

音频与语音处理 · 电气工程与系统科学 2023-05-10 Heitor Guimarães , Arthur Pimentel , Anderson Avila , Mehdi Rezagholizadeh , Tiago H. Falk

Recent years have witnessed great strides in self-supervised learning (SSL) on the speech processing. The SSL model is normally pre-trained on a great variety of unlabelled data and a large model size is preferred to increase the modeling…

音频与语音处理 · 电气工程与系统科学 2025-05-08 Yujin Wang , Changli Tang , Ziyang Ma , Zhisheng Zheng , Xie Chen , Wei-Qiang Zhang

How important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic \textit{information} in the modulation spectrum of…

音频与语音处理 · 电气工程与系统科学 2023-03-24 Samik Sadhu , Hynek Hermansky

Modern neural speech models benefit from having longer context, and many approaches have been proposed to increase the maximum context a model can use. However, few have attempted to measure how much context these models actually use, i.e.,…

声音 · 计算机科学 2025-05-29 Yen Meng , Sharon Goldwater , Hao Tang

Self-supervised learning enables the training of large neural models without the need for large, labeled datasets. It has been generating breakthroughs in several fields, including computer vision, natural language processing, biology, and…

计算与语言 · 计算机科学 2023-12-19 Luis Lugo , Valentin Vielzeuf

Despite being the best known objective for learning speech representations, the HuBERT objective has not been further developed and improved. We argue that it is the lack of an underlying principle that stalls the development, and, in this…

音频与语音处理 · 电气工程与系统科学 2026-01-05 Sung-Lin Yeh , Peter Bell , Hao Tang

AV-HuBERT, a multi-modal self-supervised learning model, has been shown to be effective for categorical problems such as automatic speech recognition and lip-reading. This suggests that useful audio-visual speech representations can be…

音频与语音处理 · 电气工程与系统科学 2023-06-02 I-Chun Chern , Kuo-Hsuan Hung , Yi-Ting Chen , Tassadaq Hussain , Mandar Gogate , Amir Hussain , Yu Tsao , Jen-Cheng Hou

The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks…

End-to-end Speech Translation is hindered by a lack of available data resources. While most of them are based on documents, a sentence-level version is available, which is however single and static, potentially impeding the usefulness of…

计算与语言 · 计算机科学 2023-11-02 Ioannis Tsiamas , José A. R. Fonollosa , Marta R. Costa-jussà

Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequences that are costly…

计算与语言 · 计算机科学 2025-03-04 Cheol Jun Cho , Nicholas Lee , Akshat Gupta , Dhruv Agarwal , Ethan Chen , Alan W Black , Gopala K. Anumanchipalli

This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is difficult to obtain, or…

音频与语音处理 · 电气工程与系统科学 2025-02-12 Louis Bahrman , Mathieu Fontaine , Gael Richard

In this study, we aim to explore efficient tuning methods for speech self-supervised learning. Recent studies show that self-supervised learning (SSL) can learn powerful representations for different speech tasks. However, fine-tuning…

音频与语音处理 · 电气工程与系统科学 2023-01-31 Zih-Ching Chen , Chin-Lun Fu , Chih-Ying Liu , Shang-Wen Li , Hung-yi Lee

Self-supervised learning (SSL) is a powerful technique for learning representations from unlabeled data. Transformer based models such as HuBERT, which consist a feature extractor and transformer layers, are leading the field in the speech…

音频与语音处理 · 电气工程与系统科学 2023-01-23 Zih-Ching Chen , Yu-Shun Sung , Hung-yi Lee

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a great number of…

音频与语音处理 · 电气工程与系统科学 2022-03-16 Zili Huang , Shinji Watanabe , Shu-wen Yang , Paola Garcia , Sanjeev Khudanpur

Audio-based automatic speech recognition (ASR) degrades significantly in noisy environments and is particularly vulnerable to interfering speech, as the model cannot determine which speaker to transcribe. Audio-visual speech recognition…

声音 · 计算机科学 2022-07-18 Bowen Shi , Wei-Ning Hsu , Abdelrahman Mohamed

Automating dysarthria assessments offers the opportunity to develop practical, low-cost tools that address the current limitations of manual and subjective assessments. Nonetheless, the small size of most dysarthria datasets makes it…

计算与语言 · 计算机科学 2024-03-26 Xavier F. Cadet , Ranya Aloufi , Sara Ahmadi-Abhari , Hamed Haddadi