English
Related papers

Related papers: Tiny Transformers for Environmental Sound Classifi…

200 papers

We propose a novel transfer learning method for speech emotion recognition allowing us to obtain promising results when only few training data is available. With as low as 125 examples per emotion class, we were able to reach a higher…

Machine Learning · Computer Science 2020-11-12 Jonathan Boigne , Biman Liyanage , Ted Östrem

In this paper, we show that ImageNet-Pretrained standard deep CNN models can be used as strong baseline networks for audio classification. Even though there is a significant difference between audio Spectrogram and standard ImageNet image…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Kamalesh Palanisamy , Dipika Singhania , Angela Yao

We present a machine learning based method for noise classification using a low-power and inexpensive IoT unit. We use Mel-frequency cepstral coefficients for audio feature extraction and supervised classification algorithms (that is,…

Sound · Computer Science 2018-09-05 Yasser Alsouda , Sabri Pllana , Arianit Kurti

Sound event detection (SED) has gained increasing attention with its wide application in surveillance, video indexing, etc. Existing models in SED mainly generate frame-level prediction, converting it into a sequence multi-label…

Sound · Computer Science 2021-11-15 Zhirong Ye , Xiangdong Wang , Hong Liu , Yueliang Qian , Rui Tao , Long Yan , Kazushige Ouchi

The explosion of IoT sensors in industrial, consumer and remote sensing use cases has come with unprecedented demand for computing infrastructure to transmit and to analyze petabytes of data. Concurrently, the world is slowly shifting its…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Emmanuel Azuh Mensah , Anderson Lee , Haoran Zhang , Yitong Shan , Kurtis Heimerl

In the context of the Internet of Things (IoT), sound sensing applications are required to run on embedded platforms where notions of product pricing and form factor impose hard constraints on the available computing power. Whereas…

Sound · Computer Science 2016-09-09 Siddharth Sigtia , Adam M. Stark , Sacha Krstulovic , Mark D. Plumbley

Multi-channel inputs offer several advantages over single-channel, to improve the robustness of on-device speech recognition systems. Recent work on multi-channel transformer, has proposed a way to incorporate such inputs into end-to-end…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-31 Feng-Ju Chang , Martin Radfar , Athanasios Mouchtaris , Maurizio Omologo

Label noise is ubiquitous in various machine learning scenarios such as self-labeling with model predictions and erroneous data annotation. Many existing approaches are based on heuristics such as sample losses, which might not be flexible…

Machine Learning · Computer Science 2022-12-29 Zhihao Wang , Zongyu Lin , Peiqi Liu , Guidong ZHeng , Junjie Wen , Xianxin Chen , Yujun Chen , Zhilin Yang

Background: Active noise cancellation has been a subject of research for decades. Traditional techniques, like the Fast Fourier Transform, have limitations in certain scenarios. This research explores the use of deep neural networks (DNNs)…

Sound · Computer Science 2024-06-03 Brandon Colelough , Andrew Zheng

DCentNet is a novel decentralized multistage signal classification approach designed for biomedical data from IoT wearable sensors, integrating early exit points (EEP) to enhance energy efficiency and processing speed. Unlike traditional…

Signal Processing · Electrical Eng. & Systems 2025-02-26 Xiaolin Li , Binhua Huang , Barry Cardiff , Deepu John

Software engineering of network-centric Artificial Intelligence (AI) and Internet of Things (IoT) enabled Cyber-Physical Systems (CPS) and services, involves complex design and validation challenges. In this paper, we propose a novel…

Software Engineering · Computer Science 2022-07-12 Armin Moin , Moharram Challenger , Atta Badii , Stephan Günnemann

Accented speech recognition and accent classification are relatively under-explored research areas in speech technology. Recently, deep learning-based methods and Transformer-based pretrained models have achieved superb performances in both…

Computation and Language · Computer Science 2022-06-30 Qingcheng Zeng , Dading Chong , Peilin Zhou , Jie Yang

In the last few years, an emerging trend in automatic speech recognition research is the study of end-to-end (E2E) systems. Connectionist Temporal Classification (CTC), Attention Encoder-Decoder (AED), and RNN Transducer (RNN-T) are the…

Computation and Language · Computer Science 2019-09-30 Jinyu Li , Rui Zhao , Hu Hu , Yifan Gong

Noise suppression and echo cancellation are critical in speech enhancement and essential for smart devices and real-time communication. Deployed in voice processing front-ends and edge devices, these algorithms must ensure efficient…

Sound · Computer Science 2023-11-28 Kaijun Tan , Benzhe Dai , Jiakui Li , Wenyu Mao

The development of highly accurate deep learning methods for indoor localization is often hindered by the unavailability of sufficient data measurements in the desired environment to perform model training. To overcome the challenge of…

Signal Processing · Electrical Eng. & Systems 2021-08-06 Mohamed I. AlHajri , Raed M. Shubair , Marwa Chafii

The attention-based Transformers have been increasingly applied to audio classification because of their global receptive field and ability to handle long-term dependency. However, the existing frameworks which are mainly extended from the…

Sound · Computer Science 2023-03-15 Xiaoyu Liu , Hanlin Lu , Jianbo Yuan , Xinyu Li

In hierarchical cognitive radio networks, edge or cloud servers utilize the data collected by edge devices for modulation classification, which, however, is faced with problems of the transmission overhead, data privacy, and computation…

Information Retrieval · Computer Science 2024-09-17 Chaowei He , Peihao Dong , Fuhui Zhou , Qihui Wu

Efficient audio feature extraction is critical for low-latency, resource-constrained speech recognition. Conventional preprocessing techniques, such as Mel Spectrogram, Perceptual Linear Prediction (PLP), and Learnable Spectrogram, achieve…

Sound · Computer Science 2025-10-28 Akshaya Rajesh , Pavithra Ananthasubramanian , Nagarajan Raghavan , Ankush Kumar

Identity recognition from ear images is an active field of research within the biometric community. The ability to capture ear images from a distance and in a covert manner makes ear recognition technology an appealing choice for…

Computer Vision and Pattern Recognition · Computer Science 2019-02-04 Žiga Emeršič , Dejan Štepec , Vitomir Štruc , Peter Peer

Hardware imperfections in RF transmitters introduce features that can be used to identify a specific transmitter amongst others. Supervised deep learning has shown good performance in this task but using datasets not applicable to real…

Signal Processing · Electrical Eng. & Systems 2019-05-21 Cyrille Morin , Leonardo Cardoso , Jakob Hoydis , Jean-Marie Gorce , Thibaud Vial