中文
相关论文

相关论文: An evaluation of data augmentation methods for sou…

200 篇论文

In anomalous sound detection, the discriminative method has demonstrated superior performance. This approach constructs a discriminative feature space through the classification of the meta-information labels for normal sounds. This feature…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Takuya Fujimura , Ibuki Kuroyanagi , Tomoki Toda

Some studies have revealed that contexts of scenes (e.g., "home," "office," and "cooking") are advantageous for sound event detection (SED). Mobile devices and sensing technologies give useful information on scenes for SED without the use…

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Jiarui Hai , Mounya Elhilali

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for…

音频与语音处理 · 电气工程与系统科学 2022-10-11 Wim Boes , Hugo Van hamme

Obtaining data to train robust artificial intelligence (AI)-based models for species classification can be challenging, particularly for rare species. Data augmentation can boost classification accuracy by increasing the diversity of…

声音 · 计算机科学 2025-12-16 Anthony Gibbons , Emma King , Ian Donohue , Andrew Parnell

Nowadays voice search for points of interest (POI) is becoming increasingly popular. However, speech recognition for local POI has remained to be a challenge due to multi-dialect and massive POI. This paper improves speech recognition…

音频与语音处理 · 电气工程与系统科学 2021-07-08 Songjun Cao , Yike Zhang , Xiaobing Feng , Long Ma

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…

With the popularity of cellular phones, events are often recorded by multiple devices from different locations and shared on social media. Several different recordings could be found for many events. Such recordings are usually noisy, where…

声音 · 计算机科学 2024-09-02 Shiran Aziz , Yossi Adi , Shmuel Peleg

Despite consistent advancement in powerful deep learning techniques in recent years, large amounts of training data are still necessary for the models to avoid overfitting. Synthetic datasets using generative adversarial networks (GAN) have…

声音 · 计算机科学 2023-04-05 Yunhao Chen , Yunjie Zhu , Zihui Yan , Jianlu Shen , Zhen Ren , Yifan Huang

Recent image-to-audio models have shown impressive performance on object-centric visual scenes. However, their application to satellite imagery remains limited by the complex, wide-area semantic ambiguity of top-down views. While satellite…

多媒体 · 计算机科学 2026-04-17 Kunlin Wu , Yanning Wang , Haofeng Tan , Boyi Chen , Teng Fei , Xianping Ma , Yang Yue , Zan Zhou , Xiaofeng Liu

In this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD). First, we explore two spatial augmentation techniques, namely audio…

声音 · 计算机科学 2023-03-08 Qing Wang , Jun Du , Hua-Xin Wu , Jia Pan , Feng Ma , Chin-Hui Lee

Audio tagging aims at predicting sound events occurred in a recording. Traditional models require enormous laborious annotations, otherwise performance degeneration will be the norm. Therefore, we investigate robust audio tagging models in…

声音 · 计算机科学 2021-10-05 Zhiling Zhang , Zelin Zhou , Haifeng Tang , Guangwei Li , Mengyue Wu , Kenny Q. Zhu

Detection and Classification Acoustic Scene and Events Challenge 2021 Task 4 uses a heterogeneous dataset that includes both recorded and synthetic soundscapes. Until recently only target sound events were considered when synthesizing the…

音频与语音处理 · 电气工程与系统科学 2024-01-02 Francesca Ronchini , Romain Serizel , Nicolas Turpault , Samuele Cornell

Sound event localization and detection (SELD) is a joint task of sound event detection and direction-of-arrival estimation. In DCASE 2022 Task 3, types of data transform from computationally generated spatial recordings to recordings of…

音频与语音处理 · 电气工程与系统科学 2022-09-12 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

While various sensors have been deployed to monitor vehicular flows, sensing pedestrian movement is still nascent. Yet walking is a significant mode of travel in many cities, especially those in Europe, Africa, and Asia. Understanding…

音频与语音处理 · 电气工程与系统科学 2025-08-11 Chaeyeon Han , Pavan Seshadri , Yiwei Ding , Noah Posner , Bon Woo Koo , Animesh Agrawal , Alexander Lerch , Subhrajit Guhathakurta

Data augmentation is an inexpensive way to increase training data diversity and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that…

声音 · 计算机科学 2021-04-20 Turab Iqbal , Karim Helwani , Arvindh Krishnaswamy , Wenwu Wang

We introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector…

声音 · 计算机科学 2017-06-06 Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maass , Radoslaw Mazur , Alfred Mertins

Acoustic source localization has been applied in different fields, such as aeronautics and ocean science, generally using multiple microphones array data to reconstruct the source location. However, the model-based beamforming methods fail…

声音 · 计算机科学 2022-04-01 Guanxing Zhou , Hao Liang , Xinghao Ding , Yue Huang , Xiaotong Tu , Saqlain Abbas

Audio perception is a key to solving a variety of problems ranging from acoustic scene analysis, music meta-data extraction, recommendation, synthesis and analysis. It can potentially also augment computers in doing tasks that humans do…

声音 · 计算机科学 2020-02-12 Prateek Verma , Kenneth Salisbury

Identifying acoustic events from a continuously streaming audio source is of interest for many applications including environmental monitoring for basic research. In this scenario neither different event classes are known nor what…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Matthias Meyer , Jan Beutel , Lothar Thiele