English
Related papers

Related papers: Temporal Convolution Network Based Onset Detection…

200 papers

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

Sound · Computer Science 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

This paper investigates different methods and various neural network architectures applicable in the time series classification domain. The data is obtained from a fleet of gas sensors that measure and track quantities such as oxygen and…

Machine Learning · Computer Science 2023-07-06 Mohamed Abouelnaga , Julien Vitay , Aida Farahani

We present in this paper a simple, yet efficient convolutional neural network (CNN) architecture for robust audio event recognition. Opposing to deep CNN architectures with multiple convolutional and pooling layers topped up with multiple…

Neural and Evolutionary Computing · Computer Science 2016-06-23 Huy Phan , Lars Hertel , Marco Maass , Alfred Mertins

In this work, we propose small footprint Convolutional Recurrent Neural Network models applied to the problem of wakeword detection and augment them with scaled dot product attention. We find that false accepts compared to Convolutional…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-26 Mohammad Omar Khursheed , Christin Jose , Rajath Kumar , Gengshen Fu , Brian Kulis , Santosh Kumar Cheekatmalla

In this paper, we propose an efficient and reproducible deep learning model for musical onset detection (MOD). We first review the state-of-the-art deep learning models for MOD, and identify their shortcomings and challenges: (i) the lack…

Sound · Computer Science 2018-06-20 Rong Gong , Xavier Serra

The challenges of polyphonic sound event detection (PSED) stem from the detection of multiple overlapping events in a time series. Recent efforts exploit Deep Neural Networks (DNNs) on Time-Frequency Representations (TFRs) of audio clips as…

Sound · Computer Science 2021-11-29 Wangkai Jin , Junyu Liu , Jianfeng Ren , Xiangjun Peng

Deep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020). This paper further extends…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Shubo Lv , Yanxin Hu , Shimin Zhang , Lei Xie

This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model performance across diverse conditions. We investigate the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-19 Davoud Shariat Panah , Alessandro Ragano , Dan Barry , Jan Skoglund , Andrew Hines

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual displays of an audio…

Computer Vision and Pattern Recognition · Computer Science 2017-06-23 M. Huzaifah

Recurrent Neural Networks (RNNs) have become the standard modeling technique for sequence data, and are used in a number of novel text-to-speech models. However, training a TTS model including RNN components has certain requirements for GPU…

Computation and Language · Computer Science 2023-04-18 Ziqi Liang

Automatic identification of animal species by their vocalization is an important and challenging task. Although many kinds of audio monitoring system have been proposed in the literature, they suffer from several disadvantages such as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-25 Weitao Xu , Xiang Zhang , Lina Yao , Wanli Xue , Bo Wei

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning…

Sound · Computer Science 2017-06-09 Sharath Adavanne , Pasi Pertilä , Tuomas Virtanen

In this study, we introduce a convolutional time-frequency-channel "Squeeze and Excitation" (tfc-SE) module to explicitly model inter-dependencies between the time-frequency domain and multiple channels. The tfc-SE module consists of two…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-06 Wei Xia , Kazuhito Koishida

Acoustic events often have a visual counterpart. Knowledge of visual information can aid the understanding of complex auditory scenes, even when only a stereo mixdown is available in the audio domain, \eg identifying which musicians are…

Neural and Evolutionary Computing · Computer Science 2017-06-30 A. Bazzica , J. C. van Gemert , C. C. S. Liem , A. Hanjalic

Convolutional Neural Networks (CNNs) have shown to be powerful medical image segmentation models. In this study, we address some of the main unresolved issues regarding these models. Specifically, training of these models on small medical…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Davood Karimi , Ali Gholipour

Large annotated lung sound databases are publicly available and might be used to train algorithms for diagnosis systems. However, it might be a challenge to develop a well-performing algorithm for small non-public data, which have only a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-03 Truc Nguyen , Franz Pernkopf

In the realm of construction safety, the detection of personal protective equipment, such as helmets, plays a critical role in preventing workplace injuries. This paper details the development and evaluation of convolutional neural networks…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Mujadded Al Rabbani Alif

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the application of…

Sound · Computer Science 2024-12-17 Idoko Agbo , Dr Hoda El-Sayed , M. D Kamruzzan Sarker

This paper presents an improved deep embedding learning method based on convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) Multi-scale convolution…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-15 Bin Gu , Wu Guo

State-of-the-art sound event detection (SED) methods usually employ a series of convolutional neural networks (CNNs) to extract useful features from the input audio signal, and then recurrent neural networks (RNNs) to model longer temporal…

‹ Prev 1 3 4 5 6 7 10 Next ›