English
Related papers

Related papers: Acoustic scene analysis with multi-head attention …

200 papers

Recently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations…

Sound · Computer Science 2022-10-12 Fan Yu , Shiliang Zhang , Pengcheng Guo , Yuhao Liang , Zhihao Du , Yuxiao Lin , Lei Xie

Cloud cover can significantly hinder the use of remote sensing images for Earth observation, prompting urgent advancements in cloud removal technology. Recently, deep learning strategies have shown strong potential in restoring…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Wenli Huang , Ye Deng , Yang Wu , Jinjun Wang

Audio question answering (AQA), acting as a widely used proxy task to explore scene understanding, has got more attention. The AQA is challenging for it requires comprehensive temporal reasoning from different scales' events of an audio…

Sound · Computer Science 2023-05-30 Guangyao Li , Yixin Xu , Di Hu

Environmental sound classification (ESC) is an important and challenging problem. In contrast to speech, sound events have noise-like nature and may be produced by a wide variety of sources. In this paper, we propose to use a novel deep…

Sound · Computer Science 2018-08-28 Zhichao Zhang , Shugong Xu , Shan Cao , Shunqing Zhang

A promising approach for steering auditory attention in complex listening environments relies on Auditory Attention Decoding (AAD), which aim to identify the attended speech stream in a multiple speaker scenario from neural recordings.…

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Olivier Siohan

This technical report describes the SurreyAudioTeam22s submission for DCASE 2022 ASC Task 1, Low-Complexity Acoustic Scene Classification (ASC). The task has two rules, (a) the ASC framework should have maximum 128K parameters, and (b)…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-03 Arshdeep Singh , James A King , Xubo Liu , Wenwu Wang , Mark D. Plumbley

We introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector…

Sound · Computer Science 2017-06-06 Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maass , Radoslaw Mazur , Alfred Mertins

Acoustic scene classification (ASC) suffers from device-induced domain shift, especially when labels are limited. Prior work focuses on curriculum-based training schedules that structure data presentation by ordering or reweighting training…

Sound · Computer Science 2026-02-02 Peihong Zhang , Yuxuan Liu , Rui Sang , Zhixin Li , Yiqiang Cai , Yizhou Tan , Shengchen Li

This paper introduces a model of environmental acoustic scenes which adopts a morphological approach by ab-stracting temporal structures of acoustic scenes. To demonstrate its potential, this model is employed to evaluate the performance of…

Machine Learning · Statistics 2015-02-03 Mathieu Lagrange , Grégoire Lafay , Mathias Rossignol , Emmanouil Benetos , Axel Roebel

Machine sounds exhibit consistent and repetitive patterns in both the frequency and time domains, which vary significantly across scales for different machine types. For instance, rotating machines often show periodic features in short time…

Sound · Computer Science 2025-08-26 Yucong Zhang , Juan Liu , Ming Li

Acoustic scene classification and related tasks have been dominated by Convolutional Neural Networks (CNNs). Top-performing CNNs use mainly audio spectograms as input and borrow their architectural design primarily from computer vision. A…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-09 Khaled Koutini , Hamid Eghbal-zadeh , Gerhard Widmer

Given a question-image input, the Visual Commonsense Reasoning (VCR) model can predict an answer with the corresponding rationale, which requires inference ability from the real world. The VCR task, which calls for exploiting the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Xuejiao Tang , Wenbin Zhang

This research identifies a gap in weakly-labelled multivariate time-series classification (TSC), where state-of-the-art TSC models do not per-form well. Weakly labelled time-series are time-series containing noise and significant…

Machine Learning · Computer Science 2021-09-20 Surayez Rahman , Chang Wei Tan

The novelty of this study consists in a multi-modality approach to scene classification, where image and audio complement each other in a process of deep late fusion. The approach is demonstrated on a difficult classification problem,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Jordan J. Bird , Diego R. Faria , Cristiano Premebida , Anikó Ekárt , George Vogiatzis

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

Attention-based models have been widely used in many areas, such as computer vision and natural language processing. However, relevant applications in time series classification (TSC) have not been explored deeply yet, causing a significant…

Machine Learning · Computer Science 2022-07-18 Bowen Zhao , Huanlai Xing , Xinhan Wang , Fuhong Song , Zhiwen Xiao

Sound event localization and detection (SELD) is a task for the classification of sound events and the identification of direction of arrival (DoA) utilizing multichannel acoustic signals. For effective classification and localization, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-18 Yusun Shul , Dayun Choi , Jung-Woo Choi

Current feature matching methods focus on point-level matching, pursuing better representation learning of individual features, but lacking further understanding of the scene. This results in significant performance degradation when…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Xiaoyong Lu , Yaping Yan , Tong Wei , Songlin Du

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end (E2E) Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great success by…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-22 Ruizhi Li , Xiaofei Wang , Sri Harish Mallidi , Shinji Watanabe , Takaaki Hori , Hynek Hermansky