English
Related papers

Related papers: EfficientLEAF: A Faster LEarnable Audio Frontend o…

200 papers

Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Yi-Chiao Wu , Dejan Marković , Steven Krenn , Israel D. Gebru , Alexander Richard

We present AFEN (Audio Feature Ensemble Learning), a model that leverages Convolutional Neural Networks (CNN) and XGBoost in an ensemble learning fashion to perform state-of-the-art audio classification for a range of respiratory diseases.…

Sound · Computer Science 2024-05-10 Rahul Nadkarni , Emmanouil Nikolakakis , Razvan Marinescu

We present an efficient frequency-based neural representation termed PREF: a shallow MLP augmented with a phasor volume that covers significant border spectra than previous Fourier feature mapping or Positional Encoding. At the core is our…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Binbin Huang , Xinhao Yan , Anpei Chen , Shenghua Gao , Jingyi Yu

With the success of language pretraining, it is highly desirable to develop more efficient architectures of good scalability that can exploit the abundant unlabeled data at a lower cost. To improve the efficiency, we examine the…

Machine Learning · Computer Science 2020-06-08 Zihang Dai , Guokun Lai , Yiming Yang , Quoc V. Le

In this paper, we presents a low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed framework can be separated into three main steps: Front-end spectrogram extraction, back-end classification, and late…

Sound · Computer Science 2021-06-17 Lam Pham , Hieu Tang , Anahid Jalali , Alexander Schindler , Ross King

Convolutional frontends are a typical choice for Transformer-based automatic speech recognition to preprocess the spectrogram, reduce its sequence length, and combine local information in time and frequency similarly. However, the width and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-13 Belen Alastruey , Lukas Drude , Jahn Heymann , Simon Wiesler

Audio classification is considered as a challenging problem in pattern recognition. Recently, many algorithms have been proposed using deep neural networks. In this paper, we introduce a new attention-based neural network architecture…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-18 Haoye Lu , Haolong Zhang , Amit Nayak

The rapid adoption of Large Language Models (LLMs) has raised significant environmental concerns. Unlike the one-time cost of training, LLM inference occurs continuously and dominates the AI energy footprint. Yet most sustainability studies…

Machine Learning · Computer Science 2026-04-08 Hemang Jain , Shailender Goyal , Divyansh Pandey , Karthik Vaidhyanathan

On-device directional hearing requires audio source separation from a given direction while achieving stringent human-imperceptible latency requirements. While neural nets can achieve significantly better performance than traditional…

Sound · Computer Science 2021-12-14 Anran Wang , Maruchi Kim , Hao Zhang , Shyamnath Gollakota

Learning a good speaker embedding is important for many automatic speaker recognition tasks, including verification, identification and diarization. The embeddings learned by softmax are not discriminative enough for open-set verification…

Machine Learning · Computer Science 2019-08-13 Zhiyong Chen , Zongze Ren , Shugong Xu

We propose a learnable mel-frequency cepstral coefficient (MFCC) frontend architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features…

Sound · Computer Science 2021-02-23 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutions often come at the cost of significant performance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Xikun Lu , Yujian Ma , Xianquan Jiang , Xuelong Wang , Jinqiu Sang

The Forward-Forward (FF) Algorithm has been recently proposed to alleviate the issues of backpropagation (BP) commonly used to train deep neural networks. However, its current formulation exhibits limitations such as the generation of…

Machine Learning · Computer Science 2024-03-29 Andreas Papachristodoulou , Christos Kyrkou , Stelios Timotheou , Theocharis Theocharides

Today very few deep learning-based mobile augmented reality (MAR) applications are applied in mobile devices because they are significantly energy-guzzling. In this paper, we design an edge-based energy-aware MAR system that enables MAR…

Computer Vision and Pattern Recognition · Computer Science 2022-05-30 Haoxin Wang , BaekGyu Kim , Jiang Xie , Zhu Han

Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obtaining good…

Convolutional Neural Networks (CNN) are being increasingly used in computer vision for a wide range of classification and recognition problems. However, training these large networks demands high computational time and energy requirements;…

Neural and Evolutionary Computing · Computer Science 2017-11-13 Syed Shakib Sarwar , Priyadarshini Panda , Kaushik Roy

Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the…

Sound · Computer Science 2022-10-13 Piotr Kawa , Marcin Plata , Piotr Syga

Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existing representations…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Junjie Wen , Junlin He , Fei Ma , Jinqiang Cui

This paper proposes to perform unsupervised detection of bioacoustic events by pooling the magnitudes of spectrogram frames after per-channel energy normalization (PCEN). Although PCEN was originally developed for speech recognition, it…

We introduce EffiFusion-GAN (Efficient Fusion Generative Adversarial Network), a lightweight yet powerful model for speech enhancement. The model integrates depthwise separable convolutions within a multi-scale block to capture diverse…

Sound · Computer Science 2025-08-21 Bin Wen , Tien-Ping Tan
‹ Prev 1 4 5 6 7 8 10 Next ›