English
Related papers

Related papers: Acoustic Scene Classification Using Bilinear Pooli…

200 papers

Convolutional neural networks (CNNs) are widely used in computer vision. They can be used not only for conventional digital image material to recognize patterns, but also for feature extraction from digital imagery representing spectral and…

Sound · Computer Science 2025-09-16 Friedrich Wolf-Monheim

Sound Event Localization and Detection refers to the problem of identifying the presence of independent or temporally-overlapped sound sources, correctly identifying to which sound class it belongs, estimating their spatial directions while…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Daniel Arteaga , Andrés Pérez-López

This technical report describes the IOA team's submission for TASK1A of DCASE2019 challenge. Our acoustic scene classification (ASC) system adopts a data augmentation scheme employing generative adversary networks. Two major classifiers, 1D…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-17 Hangting Chen , Zuozhen Liu , Zongming Liu , Pengyuan Zhang , Yonghong Yan

Sound Event Localization and Detection (SELD) is a problem related to the field of machine listening whose objective is to recognize individual sound events, detect their temporal activity, and estimate their spatial location. Thanks to the…

Recent studies have put into question the commonly assumed shift invariance property of convolutional networks, showing that small shifts in the input can affect the output predictions substantially. In this paper, we analyze the benefits…

Sound · Computer Science 2021-07-23 Eduardo Fonseca , Andres Ferraro , Xavier Serra

This thesis presents and evaluates the Dilated Convolution with Learnable Spacings (DCLS) method. Through various supervised learning experiments in the fields of computer vision, audio, and speech processing, the DCLS method proves to…

Machine Learning · Computer Science 2024-08-14 Ismail Khalfaoui-Hassani

In acoustic signal processing, the target signals usually carry semantic information, which is encoded in a hierarchal structure of short and long-term contexts. However, the background noise distorts these structures in a nonuniform way.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-26 Tassadaq Hussain , Wei-Chien Wang , Mandar Gogate , Kia Dashtipour , Yu Tsao , Xugang Lu , Adeel Ahsan , Amir Hussain

Acoustic scene classification and related tasks have been dominated by Convolutional Neural Networks (CNNs). Top-performing CNNs use mainly audio spectograms as input and borrow their architectural design primarily from computer vision. A…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-09 Khaled Koutini , Hamid Eghbal-zadeh , Gerhard Widmer

Hyperspectral Image Classification (HSIC) is a difficult task due to high inter and intra-class similarity and variability, nested regions, and overlapping. 2D Convolutional Neural Networks (CNN) emerged as a viable network whereas, 3D CNNs…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 Muhammad Ahmad

Hyperspectral image classification (HSIC) is a challenging task due to high spectral dimensionality, complex spectral-spatial correlations, and limited labeled training samples. Although transformer-based models have shown strong potential…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Farhan Ullah , Irfan Ullah , Khalil Khan , Giovanni Pau , JaKeoung Koo

Automated histopathological image analysis plays a vital role in computer-aided diagnosis of various diseases. Among developed algorithms, deep learning-based approaches have demonstrated excellent performance in multiple tasks, including…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Nima Torbati , Anastasia Meshcheryakova , Ramona Woitek , Diana Mechtcheriakova , Amirreza Mahbod

Time delay estimation is essential in Acoustic Source Localization (ASL) systems. One of the most used techniques for this purpose is the Generalized Cross Correlation (GCC) between a pair of signals and its use in Steered Response Power…

Signal Processing · Electrical Eng. & Systems 2020-06-11 Juan Manuel Vera-Diaz , Daniel Pizarro , Javier Macias-Guarasa

Polyphonic sound event detection and direction-of-arrival estimation require different input features from audio signals. While sound event detection mainly relies on time-frequency patterns, direction-of-arrival estimation relies on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-17 Thi Ngoc Tho Nguyen , Douglas L. Jones , Woon-Seng Gan

Recent studies have been revisiting whole words as the basic modelling unit in speech recognition and query applications, instead of phonetic units. Such whole-word segmental systems rely on a function that maps a variable-length speech…

Computation and Language · Computer Science 2016-01-11 Herman Kamper , Weiran Wang , Karen Livescu

In this paper, ensembles of classifiers that exploit several data augmentation techniques and four signal representations for training Convolutional Neural Networks (CNNs) for audio classification are presented and tested on three freely…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-18 Loris Nanni , Gianluca Maguolo , Sheryl Brahnam , Michelangelo Paci

Environmental sound classification (ESC) is a challenging problem due to the complexity of sounds. The ESC performance is heavily dependent on the effectiveness of representative features extracted from the environmental sounds. However,…

Sound · Computer Science 2019-07-05 Zhichao Zhang , Shugong Xu , Tianhao Qiao , Shunqing Zhang , Shan Cao

State-of-the-art sound event detection (SED) methods usually employ a series of convolutional neural networks (CNNs) to extract useful features from the input audio signal, and then recurrent neural networks (RNNs) to model longer temporal…

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Since the advent of Deep Learning (DL), Speech Enhancement (SE) models have performed well under a variety of noise conditions. However, such systems may still introduce sonic artefacts, sound unnatural, and restrict the ability for a user…

It is a practical research topic how to deal with multi-device audio inputs by a single acoustic scene classification system with efficient design. In this work, we propose Residual Normalization, a novel feature normalization method that…

Sound · Computer Science 2021-11-15 Byeonggeun Kim , Seunghan Yang , Jangho Kim , Simyung Chang