English
Related papers

Related papers: Enhancing 1-Second 3D SELD Performance with Filter…

200 papers

In this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD). First, we explore two spatial augmentation techniques, namely audio…

Sound · Computer Science 2023-03-08 Qing Wang , Jun Du , Hua-Xin Wu , Jia Pan , Feng Ma , Chin-Hui Lee

As of today, the best accuracy in line segment detection (LSD) is achieved by algorithms based on convolutional neural networks - CNNs. Unfortunately, these methods utilize deep, heavy networks and are slower than traditional model-based…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Lev Teplyakov , Leonid Erlygin , Evgeny Shvets

Recent advances in deep generative models have lead to remarkable progress in synthesizing high quality images. Following their successful application in image processing and representation learning, an important next step is to consider…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Thomas Unterthiner , Sjoerd van Steenkiste , Karol Kurach , Raphael Marinier , Marcin Michalski , Sylvain Gelly

This paper presents a novel method to involve both spatial and temporal features for semantic video segmentation. Current work on convolutional neural networks(CNNs) has shown that CNNs provide advanced spatial features supporting a very…

Computer Vision and Pattern Recognition · Computer Science 2016-09-05 Mohsen Fayyaz , Mohammad Hajizadeh Saffar , Mohammad Sabokrou , Mahmood Fathy , Reinhard Klette , Fay Huang

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-22 Hengyi Hong , Qing Wang , Jun Du , Ruoyu Wei , Mingqi Cai , Xin Fang

For learning-based sound event localization and detection (SELD) methods, different acoustic environments in the training and test sets may result in large performance differences in the validation and evaluation stages. Different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-21 Jinbo Hu , Yin Cao , Ming Wu , Feiran Yang , Ziying Yu , Wenwu Wang , Mark D. Plumbley , Jun Yang

When recognizing a long-range activity, exploring the entire video is exhaustive and computationally expensive, as it can span up to a few minutes. Thus, it is of great importance to sample only the salient parts of the video. We propose…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Noureldien Hussein , Mihir Jain , Babak Ehteshami Bejnordi

We leverage unsupervised learning of depth, egomotion, and camera intrinsics to improve the performance of single-image semantic segmentation, by enforcing 3D-geometric and temporal consistency of segmentation masks across video frames. The…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Ankita Pasad , Ariel Gordon , Tsung-Yi Lin , Anelia Angelova

Recently, window-based attention methods have shown great potential for computer vision tasks, particularly in Single Image Super-Resolution (SISR). However, it may fall short in capturing long-range dependencies and relationships between…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Dinh Phu Tran , Dao Duy Hung , Daeyoung Kim

Sound Event Detection and Localization (SELD) constitutes a complex task that depends on extensive multichannel audio recordings with annotated sound events and their respective locations. In this paper, we introduce a self-supervised…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Orlem Lima dos Santos , Karen Rosero , Roberto de Alencar Lotufo

Dataset distillation aims to synthesize compact yet informative datasets that allow models trained on them to achieve performance comparable to training on the full dataset. While this approach has shown promising results for image data,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Zhenghao Zhao , Haoxuan Wang , Kai Wang , Yuzhang Shang , Yuan Hong , Yan Yan

Sound Event Localization and Detection (SELD) is a problem related to the field of machine listening whose objective is to recognize individual sound events, detect their temporal activity, and estimate their spatial location. Thanks to the…

State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This incurs heavy computational costs that scales quadratically…

Sound · Computer Science 2025-10-30 Christodoulos Benetatos , Yongyi Zang , Randal Leistikow

Triangular, overlapping Mel-scaled filters ("f-banks") are the current standard input for acoustic models that exploit their input's time-frequency geometry, because they provide a psycho-acoustically motivated time-frequency geometry for a…

Machine Learning · Computer Science 2019-01-03 Sean Robertson , Gerald Penn , Yingxue Wang

Joint sound event localization and detection (SELD) is an emerging audio signal processing task adding spatial dimensions to acoustic scene analysis and sound event detection. A popular approach to modeling SELD jointly is using…

Sound · Computer Science 2021-09-28 Parthasaarathy Sudarsanam , Archontis Politis , Konstantinos Drossos

Semantic Segmentation is an important module for autonomous robots such as self-driving cars. The advantage of video segmentation approaches compared to single image segmentation is that temporal image information is considered, and their…

Computer Vision and Pattern Recognition · Computer Science 2019-07-17 Andreas Pfeuffer , Klaus Dietmayer

Despite the steady progress in video analysis led by the adoption of convolutional neural networks (CNNs), the relative improvement has been less drastic as that in 2D static image classification. Three main challenges exist including…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Saining Xie , Chen Sun , Jonathan Huang , Zhuowen Tu , Kevin Murphy

Environment shifts and conflicts present significant challenges for learning-based sound event localization and detection (SELD) methods. SELD systems, when trained in particular acoustic settings, often show restricted generalization…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

Recently, an event-based end-to-end model (SEDT) has been proposed for sound event detection (SED) and achieves competitive performance. However, compared with the frame-based model, it requires more training data with temporal annotations…

Sound · Computer Science 2022-04-07 Zhirong Ye , Xiangdong Wang , Hong Liu , Yueliang Qian , Rui Tao , Long Yan , Kazushige Ouchi

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual displays of an audio…

Computer Vision and Pattern Recognition · Computer Science 2017-06-23 M. Huzaifah