English
Related papers

Related papers: End-to-End Polyphonic Sound Event Detection Using …

200 papers

Recently, the end-to-end approach that learns hierarchical representations from raw data using deep convolutional neural networks has been successfully explored in the image, text and speech domains. This approach was applied to musical…

Sound · Computer Science 2017-05-23 Jongpil Lee , Jiyoung Park , Keunhyoung Luke Kim , Juhan Nam

In this paper, we propose a temporal-frequential attention model for sound event detection (SED). Our network learns how to listen with two attention models: a temporal attention model and a frequential attention model. Proposed system…

Sound · Computer Science 2025-05-06 Yu-Han Shen , Ke-Xin He , Wei-Qiang Zhang

In this technical report, the systems we submitted for subtask 4 of the DCASE 2021 challenge, regarding sound event detection, are described in detail. These models are closely related to the baseline provided for this problem, as they are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-20 Wim Boes , Hugo Van hamme

End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations. In this paper we present preliminary work that shows the feasibility of training the first layers…

Sound · Computer Science 2017-12-04 Tycho Max Sylvester Tax , Jose Luis Diez Antich , Hendrik Purwins , Lars Maaløe

Object detection and segmentation represents the basis for many tasks in computer and machine vision. In biometric recognition systems the detection of the region-of-interest (ROI) is one of the most crucial steps in the overall processing…

Computer Vision and Pattern Recognition · Computer Science 2019-02-04 Žiga Emeršič , Luka Lan Gabriel , Vitomir Štruc , Peter Peer

Sound event detection (SED) and Acoustic scene classification (ASC) are two widely researched audio tasks that constitute an important part of research on acoustic scene analysis. Considering shared information between sound events and…

Sound · Computer Science 2022-09-14 Daniel Aleksander Krause , Annamaria Mesaros

Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, natural speech contains many fast-changing and relatively brief acoustic events,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-16 Xinmeng Xu , Weiping Tu , Yuhong Yang

Identifying musical instruments in polyphonic music recordings is a challenging but important problem in the field of music information retrieval. It enables music search by instrument, helps recognize musical genres, or can make music…

Sound · Computer Science 2016-12-28 Yoonchang Han , Jaehun Kim , Kyogu Lee

One of the challenges of the Optical Music Recognition task is to transcript the symbols of the camera-captured images into digital music notations. Previous end-to-end model which was developed as a Convolutional Recurrent Neural Network…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Aozhi Liu , Lipei Zhang , Yaqi Mei , Baoqiang Han , Zifeng Cai , Zhaohua Zhu , Jing Xiao

In recent years, deep learning systems have shown a concerning trend toward increased complexity and higher energy consumption. As researchers in this domain and organizers of one of the Detection and Classification of Acoustic Scenes and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-18 Francesca Ronchini , Romain Serizel

We propose an end-to-end affect recognition approach using a Convolutional Neural Network (CNN) that handles multiple languages, with applications to emotion and personality recognition from speech. We lay the foundation of a universal…

Computation and Language · Computer Science 2019-01-28 Dario Bertero , Onno Kampman , Pascale Fung

In Sound Event Detection (SED) systems, the lengths of median filters for post-processing have never been optimized during training due to several problems. No gradient is received by the lengths so they cannot be learned during…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-23 Fengnian Zhao , Ruwei Li , Xin Liu , Liwen Xu

Modern day audio signal classification techniques lack the ability to classify low feature audio signals in the form of spectrographic temporal frequency data representations. Additionally, currently utilized techniques rely on full diverse…

Sound · Computer Science 2024-10-30 Noel Elias

In this study, we introduce a convolutional time-frequency-channel "Squeeze and Excitation" (tfc-SE) module to explicitly model inter-dependencies between the time-frequency domain and multiple channels. The tfc-SE module consists of two…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-06 Wei Xia , Kazuhito Koishida

This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition on each event. Exploiting the lightweight nature of Temporal Convolution…

The brain electrical activity presents several short events during sleep that can be observed as distinctive micro-structures in the electroencephalogram (EEG), such as sleep spindles and K-complexes. These events have been associated with…

Signal Processing · Electrical Eng. & Systems 2020-10-06 Nicolás I. Tapia , Pablo A. Estévez

Crash events identification and prediction plays a vital role in understanding safety conditions for transportation systems. While existing systems use traffic parameters correlated with crash data to classify and train these models, we…

Sound · Computer Science 2022-03-14 Zubayer Islam , Mohamed Abdel-Aty

State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features. Recent studies attempted to extract speaker embeddings directly from raw waveforms and have shown…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-10 Ge Zhu , Fei Jiang , Zhiyao Duan

The choice of an optimal time-frequency resolution is usually a difficult but important step in tasks involving speech signal classification, e.g., speech anti-spoofing. The variations of the performance with different choices of…

Sound · Computer Science 2021-10-12 Wei Liu , Meng Sun , Xiongwei Zhang , Hugo Van hamme , Thomas Fang Zheng

In this paper we present our system for the detection and classification of acoustic scenes and events (DCASE) 2020 Challenge Task 4: Sound event detection and separation in domestic environments. We introduce two new models: the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-12 Janek Ebbers , Reinhold Haeb-Umbach