English
Related papers

Related papers: Device-Robust Acoustic Scene Classification via Im…

200 papers

Deep neural networks have become popular in many supervised learning tasks, but they may suffer from overfitting when the training dataset is limited. To mitigate this, many researchers use data augmentation, which is a widely used and…

Machine Learning · Computer Science 2022-05-27 Jianhan Wu , Shijing Si , Jianzong Wang , Jing Xiao

The reverberation time (T60) and the direct-to-reverberant ratio (DRR) are commonly used to characterize room acoustic environments. Both parameters can be measured from an acoustic impulse response (AIR) or using blind estimation methods…

Sound · Computer Science 2019-10-23 Nicholas J. Bryan

Medical audio classification remains challenging due to low signal-to-noise ratios, subtle discriminative features, and substantial intra-class variability, often compounded by class imbalance and limited training data. Synthetic data…

Sound · Computer Science 2026-02-04 David McShannon , Anthony Mella , Nicholas Dietrich

We study the problem of learning robust acoustic models in adverse environments, characterized by a significant mismatch between training and test conditions. This problem is of paramount importance for the deployment of speech recognition…

Sound · Computer Science 2022-06-30 Dino Oglic , Zoran Cvetkovic , Peter Sollich , Steve Renals , Bin Yu

Audio tagging has attracted increasing attention since last decade and has various potential applications in many fields. The objective of audio tagging is to predict the labels of an audio clip. Recently deep learning methods have been…

Sound · Computer Science 2018-08-14 Shengyun Wei , Kele Xu , Dezhi Wang , Feifan Liao , Huaimin Wang , Qiuqiang Kong

Data augmentation is an inexpensive way to increase training data diversity and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that…

Sound · Computer Science 2021-04-20 Turab Iqbal , Karim Helwani , Arvindh Krishnaswamy , Wenwu Wang

The use of deep learning for radio modulation recognition has become prevalent in recent years. This approach automatically extracts high-dimensional features from large datasets, facilitating the accurate classification of modulation…

Machine Learning · Computer Science 2023-11-08 Tao Chen , Shilian Zheng , Kunfeng Qiu , Luxin Zhang , Qi Xuan , Xiaoniu Yang

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Jiarui Hai , Mounya Elhilali

A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a…

Sound · Computer Science 2020-05-21 Ricard Marxer , Hendrik Purwins

Deep learning has recently been applied to automatically classify the modulation categories of received radio signals without manual experience. However, training deep learning models requires massive volume of data. An insufficient…

Signal Processing · Electrical Eng. & Systems 2019-12-11 Liang Huang , Weijian Pan , You Zhang , LiPing Qian , Nan Gao , Yuan Wu

Previous work has shown that it is possible to improve speech recognition by learning acoustic features from paired acoustic-articulatory data, for example by using canonical correlation analysis (CCA) or its deep extensions. One limitation…

Computation and Language · Computer Science 2018-03-21 Qingming Tang , Weiran Wang , Karen Livescu

In this paper, we present an acoustic scene classification framework based on a large-margin factorized convolutional neural network (CNN). We adopt the factorized CNN to learn the patterns in the time-frequency domain by factorizing the 2D…

Sound · Computer Science 2019-10-16 Janghoon Cho , Sungrack Yun , Hyoungwoo Park , Jungyun Eum , Kyuwoong Hwang

Collecting sufficient amount of data that can represent various acoustic environmental attributes is a critical problem for distributed acoustic machine learning. Several audio data augmentation techniques have been introduced to address…

Sound · Computer Science 2021-01-07 Chunheng Jiang , Jae-wook Ahn , Nirmit Desai

Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples' augmentations (e.g., with time or frequency masking). However, they might be…

Sound · Computer Science 2023-04-11 Jian Guan , Feiyang Xiao , Youde Liu , Qiaoxi Zhu , Wenwu Wang

In recent decade, many state-of-the-art algorithms on image classification as well as audio classification have achieved noticeable successes with the development of deep convolutional neural network (CNN). However, most of the works only…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Bold Naranchimeg , Chao Zhang , Takuya Akashi

Deep learning-based methods deliver state-of-the-art performance for solving inverse problems that arise in computational imaging. These methods can be broadly divided into two groups: (1) learn a network to map measurements to the signal…

Image and Video Processing · Electrical Eng. & Systems 2023-10-11 Nebiyou Yismaw , Ulugbek S. Kamilov , M. Salman Asif

Automatic speech emotion recognition (SER) is a challenging task that plays a crucial role in natural human-computer interaction. One of the main challenges in SER is data scarcity, i.e., insufficient amounts of carefully labeled data to…

Sound · Computer Science 2021-08-17 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

The room impulse response (RIR) encodes, among others, information about the distance of an acoustic source from the sensors. Deep neural networks (DNNs) have been shown to be able to extract that information for acoustic distance…

Sound · Computer Science 2024-08-27 Tobias Gburrek , Adrian Meise , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone…

Computation and Language · Computer Science 2015-09-02 Andreas Schwarz , Christian Huemmer , Roland Maas , Walter Kellermann

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Miseul Kim , Soo Jin Park , Kyungguen Byun , Hyeon-Kyeong Shin , Sunkuk Moon , Shuhua Zhang , Erik Visser