English
Related papers

Related papers: Feature Aggregation in Joint Sound Classification …

200 papers

Spectrograms have been widely used in Convolutional Neural Networks based schemes for acoustic scene classification, such as the STFT spectrogram and the MFCC spectrogram, etc. They have different time-frequency characteristics,…

Computer Vision and Pattern Recognition · Computer Science 2018-09-06 Weiping Zheng , Zhenyao Mo , Xiaotao Xing , Gansen Zhao

The ability of deep convolutional neural networks (CNN) to learn discriminative spectro-temporal patterns makes them well suited to environmental sound classification. However, the relative scarcity of labeled data has impeded the…

Sound · Computer Science 2017-04-05 Justin Salamon , Juan Pablo Bello

Recent advancements have highlighted the efficacy of self-supervised learning (SSL) features in various speech-related tasks, providing lightweight and versatile multi-view speech representations. However, our study reveals that while SSL…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-15 Weiqiao Shan , Yuhao Zhang , Yuchen Han , Bei Li , Xiaofeng Zhao , Yuang Li , Min Zhang , Hao Yang , Tong Xiao , Jingbo Zhu

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on…

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while deep learning has…

Sound · Computer Science 2025-09-05 Zhaorui Sun , Yihao Chen , Jialong Wang , Minqiang Xu , Lei Fang , Sian Fang , Lin Liu

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embedding vector each,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Inho Kim , Youngkil Song , Jicheol Park , Won Hwa Kim , Suha Kwak

Following the rapidly growing digital image usage, automatic image categorization has become preeminent research area. It has broaden and adopted many algorithms from time to time, whereby multi-feature (generally, hand-engineered features)…

Computer Vision and Pattern Recognition · Computer Science 2017-05-12 Thangarajah Akilan , Q. M. Jonathan Wu , Wei Jiang

Deep networks have gained immense popularity in Computer Vision and other fields in the past few years due to their remarkable performance on recognition/classification tasks surpassing the state-of-the art. One of the keys to their success…

Machine Learning · Computer Science 2018-06-04 Rudrasis Chakraborty , Chun-Hao Yang , Baba C. Vemuri

Visual recognition requires rich representations that span levels from low to high, scales from small to large, and resolutions from fine to coarse. Even with the depth of features in a convolutional network, a layer in isolation is not…

Computer Vision and Pattern Recognition · Computer Science 2019-01-07 Fisher Yu , Dequan Wang , Evan Shelhamer , Trevor Darrell

Efficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods…

Sound · Computer Science 2023-12-27 Yiyuan Yang , Kaichen Zhou , Niki Trigoni , Andrew Markham

Self-supervised learning (SSL) models offer powerful representations for sound event detection (SED), yet their synergistic potential remains underexplored. This study systematically evaluates state-of-the-art SSL models to guide optimal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Hanfang Cui , Longfei Song , Li Li , Dongxing Xu , Yanhua Long

Message Passing Graph Neural Networks (MPGNNs) have emerged as the preferred method for modeling complex interactions across diverse graph entities. While the theory of such models is well understood, their aggregation module has not…

Machine Learning · Computer Science 2024-10-01 Mitchell Keren Taraday , Almog David , Chaim Baskin

Effective fusion of multi-scale features is crucial for improving speaker verification performance. While most existing methods aggregate multi-scale features in a layer-wise manner via simple operations, such as summation or concatenation.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-04 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Qian Chen , Jiajun Qi

In most state-of-the-art hashing-based visual search systems, local image descriptors of an image are first aggregated as a single feature vector. This feature vector is then subjected to a hashing function that produces a binary hash code.…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Thanh-Toan Do , Dang-Khoa Le Tan , Trung T. Pham , Ngai-Man Cheung

Acoustic scene classification (ASC) has been approached in the last years using deep learning techniques such as convolutional neural networks or recurrent neural networks. Many state-of-the-art solutions are based on image classification…

This study addresses the challenge of integrating complex, high-dimensional deep semantic features with simple, interpretable structural cues for lyrical content classification. We introduce a novel Synergistic Fusion Layer (SFL)…

Machine Learning · Computer Science 2025-11-18 M. A. Gameiro

Different layers of deep convolutional neural networks(CNNs) can encode different-level information. High-layer features always contain more semantic information, and low-layer features contain more detail information. However, low-layer…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Chen Du , Chunheng Wang , Yanna Wang , Cunzhao Shi , Baihua Xiao

We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an…

Multimedia · Computer Science 2017-01-05 Naoya Takahashi , Michael Gygli , Luc Van Gool

In this paper, we propose the use of spatial and harmonic features in combination with long short term memory (LSTM) recurrent neural network (RNN) for automatic sound event detection (SED) task. Real life sound recordings typically have…