English
Related papers

Related papers: Align-ULCNet: Towards Low-Complexity and Robust Ac…

200 papers

Deep Learning-based end-to-end Automatic Speech Recognition (ASR) has made significant strides but still struggles with performance on out-of-domain samples due to domain shifts in real-world scenarios. Test-Time Adaptation (TTA) methods…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-04 Guan-Ting Lin , Wei-Ping Huang , Hung-yi Lee

This study presents UX-Net, a time-domain audio separation network (TasNet) based on a modified U-Net architecture. The proposed UX-Net works in real-time and handles either single or multi-microphone input. Inspired by the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-31 Kashyap Patel , Anton Kovalyov , Issa Panahi

We propose Mobile Audio Streaming Networks (MASnet) for efficient low-latency speech enhancement, which is particularly suitable for mobile devices and other applications where computational capacity is a limitation. MASnet processes…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Michał Romaniuk , Piotr Masztalski , Karol Piaskowski , Mateusz Matuszewski

Acoustic Echo Cancellation (AEC) plays a key role in speech interaction by suppressing the echo received at microphone introduced by acoustic reverberations from loudspeakers. Since the performance of linear adaptive filter (AF) would…

Sound · Computer Science 2021-06-02 Lu Ma , Song Yang , Yaguang Gong , Zhongqin Wu

This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing and retrieval of similar audio events using an audio example…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Sagar Dutta , Vipul Arora

Despite the large progress in supervised learning with neural networks, there are significant challenges in obtaining high-quality, large-scale and accurately labelled datasets. In such a context, how to learn in the presence of noisy…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Chen Feng , Georgios Tzimiropoulos , Ioannis Patras

Despite the remarkable progress made by learning based stereo matching algorithms, one key challenge remains unsolved. Current state-of-the-art stereo models are mostly based on costly 3D convolutions, the cubic computational complexity and…

Computer Vision and Pattern Recognition · Computer Science 2020-04-22 Haofei Xu , Juyong Zhang

Acoustic echo cancellation (AEC) in full-duplex communication systems eliminates acoustic feedback. However, nonlinear distortions induced by audio devices, background noise, reverberation, and double-talk reduce the efficiency of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Vinay Kothapally , Yong Xu , Meng Yu , Shi-Xiong Zhang , Dong Yu

Speech enhancement (SE) is crucial for reliable communication devices or robust speech recognition systems. Although conventional artificial neural networks (ANN) have demonstrated remarkable performance in SE, they require significant…

Sound · Computer Science 2023-07-28 Abir Riahi , Éric Plourde

Objective: Lung auscultation is a valuable tool in diagnosing and monitoring various respiratory diseases. However, lung sounds (LS) are significantly affected by numerous sources of contamination, especially when recorded in real-world…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Samiul Based Shuvo , Syed Samiul Alam , Taufiq Hasan

Distant speech recognition is a challenge, particularly due to the corruption of speech signals by reverberation caused by large distances between the speaker and microphone. In order to cope with a wide range of reverberations in…

Computation and Language · Computer Science 2016-08-18 Jeehye Lee , Myungin Lee , Joon-Hyuk Chang

Sleep-disordered breathing (SDB) is a serious and prevalent condition, and acoustic analysis via consumer devices (e.g. smartphones) offers a low-cost solution to screening for it. We present a novel approach for the acoustic identification…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-08 Hector E. Romero , Ning Ma , Guy J. Brown , Amy V. Beeston , Madina Hasan

Computer analysis of Lung Sound (LS) signals has been proposed in recent years as a tool to analyze the lungs' status but there have always been main challenges, including the contamination of LS with environmental noises, which come from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-21 Mozhde Firoozi Pouyani , Mansour Vali , Mohammad Amin Ghasemi

The DeepFilterNet (DFN) architecture was recently proposed as a deep learning model suited for hearing aid devices. Despite its competitive performance on numerous benchmarks, it still follows a `one-size-fits-all' approach, which aims to…

Drones are becoming increasingly important in search and rescue missions, and even military operations. While the majority of drones are equipped with camera vision capabilities, the realm of drone audition remains underexplored due to the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-11 Yihsuan Wu , Yukai Chiu , Michael Anthony , Mingsian R. Bai

We propose a new training algorithm, ScanMix, that explores semantic clustering and semi-supervised learning (SSL) to allow superior robustness to severe label noise and competitive robustness to non-severe label noise problems, in…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Ragav Sachdeva , Filipe R Cordeiro , Vasileios Belagiannis , Ian Reid , Gustavo Carneiro

This paper presents an acoustic echo canceler based on a U-Net convolutional neural network for single-talk and double-talk scenarios. U-Net networks have previously been used in the audio processing area for source separation problems…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-21 J. Silva-Rodríguez , M. F. Dolz , M. Ferrer , A. Castelló , V. Naranjo , G. Piñero

Federated edge learning (FEEL) enables wireless devices to collaboratively train a centralised model without sharing raw data, but repeated uplink transmission of model updates makes communication the dominant bottleneck. Over-the-air (OTA)…

Information Theory · Computer Science 2025-12-24 Antonio Tarizzo , Mohammad Kazemi , Deniz Gündüz

Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large…

Sound · Computer Science 2026-03-27 Yupei Li , Shuaijie Shao , Manuel Milling , Björn Schuller

The increasing prevalence of microphones in everyday devices and the growing reliance on online services have amplified the risk of acoustic side-channel attacks (ASCAs) targeting keyboards. This study explores deep learning techniques,…

Machine Learning · Computer Science 2025-02-20 Jin Hyun Park , Seyyed Ali Ayati , Yichen Cai