中文
相关论文

相关论文: Executable Boundary Contracts for Sound Event Trac…

200 篇论文

We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations.…

音频与语音处理 · 电气工程与系统科学 2024-12-02 Florian Schmid , Tobias Morocutti , Francesco Foscarin , Jan Schlüter , Paul Primus , Gerhard Widmer

Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Kranti Kumar Parida , Omar Emara , Hazel Doughty , Dima Damen

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…

计算与语言 · 计算机科学 2018-08-08 Yu-Hsuan Wang , Hung-yi Lee , Lin-shan Lee

Sound event detection (SED) is the task of tagging the absence or presence of audio events and their corresponding interval within a given audio clip. While SED can be done using supervised machine learning, where training data is fully…

声音 · 计算机科学 2021-02-08 Heinrich Dinkel , Mengyue Wu , Kai Yu

In many situations, we would like to hear desired sound events (SEs) while being able to ignore interference. Target sound extraction (TSE) tackles this problem by estimating the audio signal of the sounds of target SE classes in a mixture…

音频与语音处理 · 电气工程与系统科学 2022-11-03 Marc Delcroix , Jorge Bennasar Vázquez , Tsubasa Ochiai , Keisuke Kinoshita , Yasunori Ohishi , Shoko Araki

In this technique report, we present a bunch of methods for the task 4 of Detection and Classification of Acoustic Scenes and Events 2017 (DCASE2017) challenge. This task evaluates systems for the large-scale detection of sound events using…

声音 · 计算机科学 2017-11-28 Yong Xu , Qiuqiang Kong , Wenwu Wang , Mark D. Plumbley

In this paper, we describe in detail our systems for DCASE 2020 Task 4. The systems are based on the 1st-place system of DCASE 2019 Task 4, which adopts weakly-supervised framework with an attention-based embedding-level pooling module and…

声音 · 计算机科学 2020-11-03 Yuxin Huang , Liwei Lin , Shuo Ma , Xiangdong Wang , Hong Liu , Yueliang Qian , Min Liu , Kazushige Ouch

Action chunking is widely used in generative visuomotor policies, yet the recurring execution discontinuities at chunk boundaries still lack a mechanistic explanation. This paper treats chunk-boundary artifact as an analyzable mechanism…

机器人学 · 计算机科学 2026-05-22 Rui Wang

This paper is motivated by the automation of neuropsychological tests involving discourse analysis in the retellings of narratives by patients with potential cognitive impairment. In this scenario the task of sentence boundary detection in…

计算与语言 · 计算机科学 2017-08-17 Marcos V. Treviso , Christopher D. Shulby , Sandra M. Aluisio

In this paper we present our work on Task 1 Acoustic Scene Classi- fication and Task 3 Sound Event Detection in Real Life Recordings. Among our experiments we have low-level and high-level features, classifier optimization and other…

This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating…

声音 · 计算机科学 2025-07-25 Quoc Thinh Vo , David Han

Unsupervised spoken term discovery consists of two tasks: finding the acoustic segment boundaries and labeling acoustically similar segments with the same labels. We perform segmentation based on the assumption that the frame feature…

音频与语音处理 · 电气工程与系统科学 2020-07-28 Saurabhchand Bhati , Jesús Villalba , Piotr Żelasko , Najim Dehak

In this paper, an approximation recursive formula of the mean-square error lower bound for the discrete-time nonlinear filtering problem when noises of dynamic systems are temporally correlated is derived based on the Van Trees (posterior)…

动力系统 · 数学 2015-04-28 Zhiguo Wang , Xiaojing Shen

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose…

音频与语音处理 · 电气工程与系统科学 2025-01-13 Han Yin , Jisheng Bai , Yang Xiao , Hui Wang , Siqi Zheng , Yafeng Chen , Rohan Kumar Das , Chong Deng , Jianfeng Chen

Automated Audio Captioning is a cross-modal task, generating natural language descriptions to summarize the audio clips' sound events. However, grounding the actual sound events in the given audio based on its corresponding caption has not…

声音 · 计算机科学 2021-02-24 Xuenan Xu , Heinrich Dinkel , Mengyue Wu , Kai Yu

This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted from the mel-spectral…

音频与语音处理 · 电气工程与系统科学 2025-06-27 Jongyeon Park , Joonhee Lee , Do-Hyeon Lim , Hong Kook Kim , Hyeongcheol Geum , Jeong Eun Lim

Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of…

声音 · 计算机科学 2024-12-30 Yunyi Liu , Craig Jin

Boundary samples are special inputs to artificial neural networks crafted to identify the execution environment used for inference by the resulting output label. The paper presents and evaluates algorithms to generate transparent boundary…

机器学习 · 计算机科学 2021-06-15 Alexander Schlögl , Tobias Kupek , Rainer Böhme

Self-supervised learning has drawn attention through its effectiveness in learning in-domain representations with no ground-truth annotations; in particular, it is shown that properly designed pretext tasks (e.g., contrastive prediction…

计算机视觉与模式识别 · 计算机科学 2022-01-17 Jonghwan Mun , Minchul Shin , Gunsoo Han , Sangho Lee , Seongsu Ha , Joonseok Lee , Eun-Sol Kim

In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a complex task that combines temporal event classification with…

音频与语音处理 · 电气工程与系统科学 2025-09-09 Davide Berghi , Philip J. B. Jackson