English
Related papers

Related papers: City classification from multiple real-world sound…

200 papers

Sound event detection (SED) is a task to detect sound events in an audio recording. One challenge of the SED task is that many datasets such as the Detection and Classification of Acoustic Scenes and Events (DCASE) datasets are weakly…

Sound · Computer Science 2020-08-25 Qiuqiang Kong , Yong Xu , Wenwu Wang , Mark D. Plumbley

With ever-increasing number of car-mounted electric devices and their complexity, audio classification is increasingly important for the automotive industry as a fundamental tool for human-device interactions. Existing approaches for audio…

Sound · Computer Science 2018-04-11 Myounggyu Won , Haitham Alsaadan , Yongsoon Eun

Computer-based scene understanding has influenced fields ranging from urban planning to autonomous vehicle performance, yet little is known about how well these technologies work across social differences. We investigate the biases of deep…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Michelle R. Greene , Mariam Josyula , Wentao Si , Jennifer A. Hart

The task of classifying videos of natural dynamic scenes into appropriate classes has gained lot of attention in recent years. The problem especially becomes challenging when the camera used to capture the video is dynamic. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2015-09-01 Aalok Gangopadhyay , Shivam Mani Tripathi , Ishan Jindal , Shanmuganathan Raman

In this paper we present ensembles of classifiers for automated animal audio classification, exploiting different data augmentation techniques for training Convolutional Neural Networks (CNNs). The specific animal audio classification…

Machine Learning · Computer Science 2020-03-17 Loris Nanni , Gianluca Maguolo , Michelangelo Paci

The problem of identifying voice commands has always been a challenge due to the presence of noise and variability in speed, pitch, etc. We will compare the efficacies of several neural network architectures for the speech recognition…

Machine Learning · Statistics 2020-11-25 Sanjay Krishna Gouda , Salil Kanetkar , David Harrison , Manfred K Warmuth

Recent deep learning models have demonstrated strong capabilities for classifying text and non-text components in natural images. They extract a high-level feature computed globally from a whole image component (patch), where the cluttered…

Computer Vision and Pattern Recognition · Computer Science 2016-05-04 Tong He , Weilin Huang , Yu Qiao , Jian Yao

In this paper, we propose a sub-utterance unit selection framework to remove acoustic segments in audio recordings that carry little information for acoustic scene classification (ASC). Our approach is built upon a universal set of acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Hu Hu , Sabato Marco Siniscalchi , Yannan Wang , Xue Bai , Jun Du , Chin-Hui Lee

We present a work on low-complexity acoustic scene classification (ASC) with multiple devices, namely the subtask A of Task 1 of the DCASE2021 challenge. This subtask focuses on classifying audio samples of multiple devices with a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Yanxiong Li , Wenchang Cao , Wei Xie , Qisheng Huang , Wenfeng Pang , Qianhua He

Convolutional neural networks (CNNs) are commonplace in high-performing solutions to many real-world problems, such as audio classification. CNNs have many parameters and filters, with some having a larger impact on the performance than…

Sound · Computer Science 2023-05-08 James A King , Arshdeep Singh , Mark D. Plumbley

Reading text in the wild is a challenging task in the field of computer vision. Existing approaches mainly adopted Connectionist Temporal Classification (CTC) or Attention models based on Recurrent Neural Network (RNN), which is…

Computer Vision and Pattern Recognition · Computer Science 2017-09-14 Yunze Gao , Yingying Chen , Jinqiao Wang , Hanqing Lu

Heart diseases constitute a global health burden, and the problem is exacerbated by the error-prone nature of listening to and interpreting heart sounds. This motivates the development of automated classification to screen for abnormal…

Sound · Computer Science 2016-12-07 Yuhao Zhang , Sandeep Ayyar , Long-Huei Chen , Ethan J. Li

Sensor nodes in a wireless sensor network (WSN) for security surveillance applications should preferably be small, energy-efficient, and inexpensive with in-sensor computational abilities. An appropriate data processing scheme in the sensor…

Neural and Evolutionary Computing · Computer Science 2022-05-04 Anand Kumar Mukhopadhyay , Naligala Moses Prabhakar , Divya Lakshmi Duggisetty , Indrajit Chakrabarti , Mrigank Sharad

The prevalent perspectives of scene text recognition are from sequence to sequence (seq2seq) and segmentation. Nevertheless, the former is composed of many components which makes implementation and deployment complicated, while the latter…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Hongxiang Cai , Jun Sun , Yichao Xiong

Many real-world time-series analysis problems are characterised by scarce data. Solutions typically rely on hand-crafted features extracted from the time or frequency domain allied with classification or regression engines which condition…

In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-15 Ee-Leng Tan , Jun Wei Yeow , Santi Peksi , Haowen Li , Ziyi Yang , Woon-Seng Gan

Over the past two decades, CNN architectures have produced compelling models of sound perception and cognition, learning hierarchical organizations of features. Analogous to successes in computer vision, audio feature classification can be…

Sound · Computer Science 2025-05-13 Prateek Verma , Jonathan Berger

This study deals with semantic segmentation of high-resolution (aerial) images where a semantic class label is assigned to each pixel via supervised classification as a basis for automatic map generation. Recently, deep convolutional neural…

Computer Vision and Pattern Recognition · Computer Science 2017-11-22 Pascal Kaiser , Jan Dirk Wegner , Aurelien Lucchi , Martin Jaggi , Thomas Hofmann , Konrad Schindler

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

We propose an efficient end-to-end convolutional neural network architecture, AclNet, for audio classification. When trained with our data augmentation and regularization, we achieved state-of-the-art performance on the ESC-50 corpus with…

Sound · Computer Science 2018-11-19 Jonathan J Huang , Juan Jose Alvarado Leanos