English
Related papers

Related papers: FastAST: Accelerating Audio Spectrogram Transforme…

200 papers

Acceleration methods for diffusion models (e.g., token merging or downsampling) typically optimize synthesis quality under reduced compute, yet often ignore discriminative capacity. We revisit token compression with a joint objective and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Jiacheng Liu , Shengkun Tang , Jiacheng Cui , Dongkuan Xu , Zhiqiang Shen

We present a modular approach to building cascade speech translation (AST) models that guarantees that the resulting model performs no worse than the 1-best cascade baseline while preserving state-of-the-art speech recognition (ASR) and…

Computation and Language · Computer Science 2024-07-26 Ciprian Chelba , Johan Schalkwyk

Here we propose FastFCA-AS, an accelerated algorithm for Full-rank spatial Covariance Analysis (FCA), which is a robust audio source separation method proposed by Duong et al. ["Under-determined reverberant audio source separation using a…

Sound · Computer Science 2018-05-25 Nobutaka Ito , Tomohiro Nakatani

The atomic norm provides a generalization of the $\ell_1$-norm to continuous parameter spaces. When applied as a sparse regularizer for line spectral estimation the solution can be obtained by solving a convex optimization problem. This…

Numerical Analysis · Mathematics 2019-06-24 Thomas Lundgaard Hansen , Tobias Lindstrøm Jensen

Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Bing Han , Chushu Zhou , Yifan Yang , Wei Wang , Chenda Li , Wangyou Zhang , Yanmin Qian

Transformers achieve promising performance in document understanding because of their high effectiveness and still suffer from quadratic computational complexity dependency on the sequence length. General efficient transformers are…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Mingliang Zhai , Yulin Li , Xiameng Qin , Chen Yi , Qunyi Xie , Chengquan Zhang , Kun Yao , Yuwei Wu , Yunde Jia

Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models. However, due to the naturally significant sequence length of audio, the…

Sound · Computer Science 2024-12-18 Kai Qiu , Xiang Li , Hao Chen , Jie Sun , Jinglu Wang , Zhe Lin , Marios Savvides , Bhiksha Raj

For deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation. To address the mismatch issue, numerous noise adaptation strategies have been derived. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-22 Chi-Chang Lee , Cheng-Hung Hu , Yu-Chen Lin , Chu-Song Chen , Hsin-Min Wang , Yu Tsao

In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data driven techniques, e.g., deep learning, where audio…

Sound · Computer Science 2018-06-29 Eduardo Fonseca , Rong Gong , Xavier Serra

Significant improvement has been achieved in automated audio captioning (AAC) with recent models. However, these models have become increasingly large as their performance is enhanced. In this work, we propose a knowledge distillation (KD)…

Sound · Computer Science 2024-07-22 Xuenan Xu , Haohe Liu , Mengyue Wu , Wenwu Wang , Mark D. Plumbley

End-to-end speech translation relies on data that pair source-language speech inputs with corresponding translations into a target language. Such data are notoriously scarce, making synthetic data augmentation by back-translation or…

Computation and Language · Computer Science 2023-06-12 Tsz Kin Lam , Shigehiko Schamoni , Stefan Riezler

Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) adopt square…

Sound · Computer Science 2025-09-01 Aditya Makineni , Baocheng Geng , Qing Tian

The field of steganography has experienced a surge of interest due to the recent advancements in AI-powered techniques, particularly in the context of multimodal setups that enable the concealment of signals within signals of a different…

Cryptography and Security · Computer Science 2023-03-16 Jaume Ros , Margarita Geleta , Jordi Pons , Xavier Giro-i-Nieto

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

Low-dose CT (LDCT) protocols reduce radiation exposure but increase image noise, compromising diagnostic confidence. Diffusion-based generative models have shown promise for LDCT denoising by learning image priors and performing iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Tomás de la Sotta , José M. Saavedra , Héctor Henríquez , Violeta Chang , Aline Xavier

Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Kevin Wilkinghoff , Zheng-Hua Tan

Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimodal features in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Tanvir Mahmud , Shentong Mo , Yapeng Tian , Diana Marculescu

Recent foundational models, SSAST, EAT, HuBERT, Qwen-Audio, and Audio Flamingo, achieve top-tier results across standard audio benchmarks but are limited by fixed input rates and durations, hindering their reusability. This paper introduces…

Sound · Computer Science 2025-11-25 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

MIDI velocity is crucial for capturing expressive dynamics in human performances. In practical scenarios, a music score with inaccurate velocities may be available alongside the performance audio (e.g., music education and free online…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Zhanhong He , Roberto Togneri , David Huang

This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than real-time on CPUs. FastSVC uses Conformer-based phoneme…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-25 Songxiang Liu , Yuewen Cao , Na Hu , Dan Su , Helen Meng
‹ Prev 1 3 4 5 6 7 10 Next ›