English
Related papers

Related papers: Pre-training Autoencoder for Acoustic Event Classi…

200 papers

Deep neural network-based systems have significantly improved the performance of speaker diarization tasks. However, end-to-end neural diarization (EEND) systems often struggle to generalize to scenarios with an unseen number of speakers,…

Sound · Computer Science 2023-09-14 Zhengyang Chen , Bing Han , Shuai Wang , Yanmin Qian

Wearable audio devices with active noise control (ANC) enhance listening comfort but often at the expense of situational awareness. However, this auditory isolation may mask crucial environmental cues, posing significant safety risks. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-09 Jun-Wei Yeow , Ee-Leng Tan , Santi Peksi , Zhen-Ting Ong , Woon-Seng Gan

Decoding speech from non-invasive brain signals, such as electroencephalography (EEG), has the potential to advance brain-computer interfaces (BCIs), with applications in silent communication and assistive technologies for individuals with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-30 Terrance Yu-Hao Chen , Yulin Chen , Pontus Soederhaell , Sadrishya Agrawal , Kateryna Shapovalenko

Acoustic Scene Classification (ASC) faces challenges in generalizing across recording devices, particularly when labeled data is limited. The DCASE 2024 Challenge Task 1 highlights this issue by requiring models to learn from small labeled…

Sound · Computer Science 2026-02-02 Peihong Zhang , Yuxuan Liu , Zhixin Li , Rui Sang , Yiqiang Cai , Yizhou Tan , Shengchen Li

We investigate the potential of autoencoders (AEs) for building a joint communication and sensing (JCAS) system that enables communication with one user while detecting multiple radar targets and estimating their positions. Foremost, we…

Signal Processing · Electrical Eng. & Systems 2023-01-25 Charlotte Muth , Laurent Schmalen

End-to-end learning of a communications system using the deep learning-based autoencoder concept has drawn interest in recent research due to its simplicity, flexibility and its potential of adapting to complex channel models and practical…

Information Theory · Computer Science 2020-01-22 Nuwanthika Rajapaksha , Nandana Rajatheva , Matti Latva-aho

Spiking Neural Networks (SNNs) offer energy efficient processing suitable for edge applications, but conventional sensor data must first be converted into spike trains for neuromorphic processing. Environmental sound, including urban…

Sound · Computer Science 2025-11-27 Andres Larroza , Javier Naranjo-Alcazar , Vicent Ortiz , Maximo Cobos , Pedro Zuccarello

Masked Autoencoders (MAE) have been popular paradigms for large-scale vision representation pre-training. However, MAE solely reconstructs the low-level RGB signals after the decoder and lacks supervision upon high-level semantics for the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Peng Gao , Renrui Zhang , Rongyao Fang , Ziyi Lin , Hongyang Li , Hongsheng Li , Qiao Yu

Environmental Sound Classification (ESC) is an important and challenging problem, and feature representation is a critical and even decisive factor in ESC. Feature representation ability directly affects the accuracy of sound…

Sound · Computer Science 2019-08-19 Tianhao Qiao , Shunqing Zhang , Zhichao Zhang , Shan Cao , Shugong Xu

This paper proposes an effective modelling of sound event spectra with a hidden data-size-imbalance, for improved Acoustic Event Detection (AED). The proposed method models each event as an aggregated representation of a few latent factors,…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-08 Chaitanya Narisetty , Tatsuya Komatsu , Reishi Kondo

Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Tommaso Boccato , Michal Olak , Matteo Ferrante

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then…

Sound · Computer Science 2020-12-18 Mostafa Sadeghi , Simon Leglaive , Xavier Alameda-PIneda , Laurent Girin , Radu Horaud

In this paper we present our system for the detection and classification of acoustic scenes and events (DCASE) 2020 Challenge Task 4: Sound event detection and separation in domestic environments. We introduce two new models: the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-12 Janek Ebbers , Reinhold Haeb-Umbach

In this letter, we propose a semantic communication scheme for wireless relay channels based on Autoencoder, named AESC, which encodes and decodes sentences from the semantic dimension. The Autoencoder module provides anti-noise performance…

Information Theory · Computer Science 2021-11-22 Xinlai Luo , Zhiyong Chen , Bin Xia , Jiangzhou Wang

We investigate the potential of adaptive blind equalizers based on variational inference for carrier recovery in optical communications. These equalizers are based on a low-complexity approximation of maximum likelihood channel estimation.…

Signal Processing · Electrical Eng. & Systems 2022-09-16 Vincent Lauinger , Fred Buchali , Laurent Schmalen

Acoustic Echo Cancellation (AEC) is an essential speech signal processing technology that removes echoes from microphone inputs to facilitate natural-sounding full-duplex communication. Currently, deep learning-based AEC methods primarily…

Sound · Computer Science 2024-12-30 Fei Zhao , Xueliang Zhang

We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a way that allows the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Wim Boes , Hugo Van hamme

Event cameras, with a high dynamic range exceeding $120dB$, significantly outperform traditional embedded cameras, robustly recording detailed changing information under various lighting conditions, including both low- and high-light…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Yunfan Lu , Xiaogang Xu , Hao Lu , Yanlin Qian , Pengteng Li , Huizai Yao , Bin Yang , Junyi Li , Qianyi Cai , Weiyu Guo , Hui Xiong

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags…

Sound · Computer Science 2020-10-28 Xavier Favory , Konstantinos Drossos , Tuomas Virtanen , Xavier Serra

Autoexposure (AE) is a critical step applied by camera systems to ensure properly exposed images. While current AE algorithms are effective in well-lit environments with constant illumination, these algorithms still struggle in environments…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 SaiKiran Tedla , Beixuan Yang , Michael S. Brown