English
Related papers

Related papers: MAST: Multiscale Audio Spectrogram Transformers

200 papers

Noise pollution significantly affects our daily life and urban development. Urban Sound Tagging (UST) has attracted much attention recently, which aims to analyze and monitor urban noise pollution. One weakness of the previous UST studies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-22 Jisheng Bai , Jianfeng Chen , Mou Wang

End-to-end automatic speech translation (AST) relies on data that combines audio inputs with text translation outputs. Previous work used existing large parallel corpora of transcriptions and translations in a knowledge distillation (KD)…

Computation and Language · Computer Science 2023-07-18 Rebekka Hubert , Artem Sokolov , Stefan Riezler

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to…

Sound · Computer Science 2022-02-14 Qinghua Wu , Quanbo Shen , Jian Luan , YuJun Wang

Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-30 Phurich Saengthong , Tomoya Nishida , Kota Dohi , Natsuo Yamashita , Yohei Kawaguchi

Self-attention mechanisms have enabled transformers to achieve superhuman-level performance on many speech-to-text (STT) tasks, yet the challenge of automatic prosodic segmentation has remained unsolved. In this paper we finetune Whisper, a…

Computation and Language · Computer Science 2025-02-28 Nathan Roll , Calbert Graham , Simon Todd

Speech tokenization is the task of representing speech signals as a sequence of discrete units. Such representations can be later used for various downstream tasks including automatic speech recognition, text-to-speech, etc. More relevant…

Sound · Computer Science 2024-06-18 Shoval Messica , Yossi Adi

Acoustic models in real-time speech recognition systems typically stack multiple unidirectional LSTM layers to process the acoustic frames over time. Performance improvements over vanilla LSTM architectures have been reported by prepending…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Maarten Van Segbroeck , Harish Mallidih , Brian King , I-Fan Chen , Gurpreet Chadha , Roland Maas

Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-14 Xu Tan , Xiao-Lei Zhang

Voice anonymization masks vocal traits while preserving linguistic content, which may still leak speaker-specific patterns. To assess and strengthen privacy evaluation, we propose a dual-stream attacker that fuses spectral and…

Sound · Computer Science 2026-03-17 Ridwan Arefeen , Xiaoxiao Miao , Rong Tong , Aik Beng Ng , Simon See , Timothy Liu

Automatic Speech Recognition (ASR) aims to convert human speech content into corresponding text. In conversational scenarios, effectively utilizing context can enhance its accuracy. Large Language Models' (LLMs) exceptional long-context…

Sound · Computer Science 2026-01-19 Bingshen Mu , Hexin Liu , Hongfei Xue , Kun Wei , Lei Xie

In speech translation, leveraging multimodal data to improve model performance and address limitations of individual modalities has shown significant effectiveness. In this paper, we harness the complementary strengths of speech and text,…

Computation and Language · Computer Science 2023-05-24 Wenbiao Yin , Zhicheng Liu , Chengqi Zhao , Tao Wang , Jian Tong , Rong Ye

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson

In this paper, we present SpecAugment++, a novel data augmentation method for deep neural networks based acoustic scene classification (ASC). Different from other popular data augmentation methods such as SpecAugment and mixup that only…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Helin Wang , Yuexian Zou , Wenwu Wang

Existing leading methods for spectral reconstruction (SR) focus on designing deeper or wider convolutional neural networks (CNNs) to learn the end-to-end mapping from the RGB image to its hyperspectral image (HSI). These CNN-based methods…

Computer Vision and Pattern Recognition · Computer Science 2022-04-19 Yuanhao Cai , Jing Lin , Zudi Lin , Haoqian Wang , Yulun Zhang , Hanspeter Pfister , Radu Timofte , Luc Van Gool

We propose Fast Language-Audio Pre-training (FLAP), a self-supervised approach that efficiently and effectively learns aligned audio and language representations through masking, contrastive learning and reconstruction. For efficiency, FLAP…

Sound · Computer Science 2023-11-06 Ching-Feng Yeh , Po-Yao Huang , Vasu Sharma , Shang-Wen Li , Gargi Gosh

We propose Masked-Attention Transformers for Surgical Instrument Segmentation (MATIS), a two-stage, fully transformer-based method that leverages modern pixel-wise attention mechanisms for instrument segmentation. MATIS exploits the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Nicolás Ayobi , Alejandra Pérez-Rondón , Santiago Rodríguez , Pablo Arbeláez

Supervised speech enhancement methods have been very successful. However, in practical scenarios, there is a lack of clean speech, and self-supervised learning-based (SSL) speech enhancement methods that offer comparable enhancement…

Sound · Computer Science 2026-02-03 Rajalaxmi Rajagopalan , Ritwik Giri , Zhiqiang Tang , Kyu Han

While FastSpeech2 aims to integrate aspects of speech such as pitch, energy, and duration as conditional inputs, it still leaves scope for richer representations. As a part of this work, we leverage representations from various…

Computation and Language · Computer Science 2023-08-03 Ramanan Sivaguru , Vasista Sai Lodagala , S Umesh

Recently, self-supervised learning (SSL) techniques have been introduced to solve the monaural speech enhancement problem. Due to the lack of using clean phase information, the enhancement performance is limited in most SSL methods.…

Sound · Computer Science 2021-12-22 Yi Li , Yang Sun , Syed Mohsen Naqvi
‹ Prev 1 8 9 10 Next ›