中文
相关论文

相关论文: Metric Analysis for Spatial Semantic Segmentation …

200 篇论文

To advance immersive communication, the Detection and Classification of Acoustic Scenes and Events (DCASE) 2025 Challenge recently introduced Task 4 on Spatial Semantic Segmentation of Sound Scenes (S5). An S5 system takes a multi-channel…

音频与语音处理 · 电气工程与系统科学 2026-02-02 Binh Thien Nguyen , Masahiro Yasuda , Daiki Takeuchi , Daisuke Niizumi , Noboru Harada

Immersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services. Aiming at its further realization, the DCASE 2025 Challenge has recently introduced a task for…

音频与语音处理 · 电气工程与系统科学 2025-06-10 Binh Thien Nguyen , Masahiro Yasuda , Daiki Takeuchi , Daisuke Niizumi , Yasunori Ohishi , Noboru Harada

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline -…

声音 · 计算机科学 2025-07-24 Tobias Morocutti , Jonathan Greif , Paul Primus , Florian Schmid , Gerhard Widmer

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal…

音频与语音处理 · 电气工程与系统科学 2022-04-22 Thilo von Neumann , Keisuke Kinoshita , Christoph Boeddeker , Marc Delcroix , Reinhold Haeb-Umbach

This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted from the mel-spectral…

音频与语音处理 · 电气工程与系统科学 2025-06-27 Jongyeon Park , Joonhee Lee , Do-Hyeon Lim , Hong Kook Kim , Hyeongcheol Geum , Jeong Eun Lim

Language-queried audio source separation (LASS) aims to separate an audio source guided by a text query, with the signal-to-distortion ratio (SDR)-based metrics being commonly used to objectively measure the quality of the separated audio.…

声音 · 计算机科学 2025-01-07 Feiyang Xiao , Jian Guan , Qiaoxi Zhu , Xubo Liu , Wenbo Wang , Shuhan Qi , Kejia Zhang , Jianyuan Sun , Wenwu Wang

Source separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on…

声音 · 计算机科学 2019-12-10 Qiuqiang Kong , Yong Xu , Wenwu Wang , Mark D. Plumbley

The spatial semantic segmentation task focuses on separating and classifying sound objects from multichannel signals. To achieve two different goals, conventional methods fine-tune a large classification model cascaded with the separation…

音频与语音处理 · 电气工程与系统科学 2025-09-22 Younghoo Kwon , Jung-Woo Choi

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Simon Dahl Jepsen , Mads Græsbøll Christensen , Jesper Rindom Jensen

Music source separation aims to extract individual sound sources (e.g., vocals, drums, guitar) from a mixed music recording. However, evaluating the quality of separated audio remains challenging, as commonly used metrics like the…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Noah Jaffe , John Ashley Burgoyne

Spatial Semantic Segmentation of Sound Scenes (S5) aims to enhance technologies for sound event detection and separation from multi-channel input signals that mix multiple sound events with spatial information. This is a fundamental basis…

This paper introduces a multi-stage self-directed framework designed to address the spatial semantic segmentation of sound scene (S5) task in the DCASE 2025 Task 4 challenge. This framework integrates models focused on three distinct tasks:…

音频与语音处理 · 电气工程与系统科学 2025-09-18 Younghoo Kwon , Dongheon Lee , Dohwan Kim , Jung-Woo Choi

In speech enhancement and source separation, signal-to-noise ratio is a ubiquitous objective measure of denoising/separation quality. A decade ago, the BSS_eval toolkit was developed to give researchers worldwide a way to evaluate the…

声音 · 计算机科学 2018-11-07 Jonathan Le Roux , Scott Wisdom , Hakan Erdogan , John R. Hershey

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech…

This paper presents an overview of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2026 Challenge Task 4, Spatial Semantic Segmentation of Sound Scenes (S5). The S5 task focuses on the joint detection and separation…

A dictionary learning based audio source classification algorithm is proposed to classify a sample audio signal as one amongst a finite set of different audio sources. Cosine similarity measure is used to select the atoms during dictionary…

声音 · 计算机科学 2015-10-28 K V Vijay Girish , T V Ananthapadmanabha , A G Ramakrishnan

Source separation is the task to separate an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular…

声音 · 计算机科学 2020-02-07 Qiuqiang Kong , Yuxuan Wang , Xuchen Song , Yin Cao , Wenwu Wang , Mark D. Plumbley

In this paper, we propose a two-step training procedure for source separation via a deep neural network. In the first step we learn a transform (and it's inverse) to a latent space where masking-based separation performance using oracles is…

机器学习 · 计算机科学 2021-05-12 Efthymios Tzinis , Shrikant Venkataramani , Zhepei Wang , Cem Subakan , Paris Smaragdis

The acquisition of high-quality labeled synthetic aperture radar (SAR) data is challenging due to the demanding requirement for expert knowledge. Consequently, the presence of unreliable noisy labels is unavoidable, which results in…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yimin Fu , Zhunga Liu , Dongxiu Guo , Longfei Wang

Semantic segmentation metrics for 3D point clouds, such as mean Intersection over Union (mIoU) and Overall Accuracy (OA), present two key limitations in the context of aerial LiDAR data. First, they treat all misclassifications equally…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Alex Salvatierra , José Antonio Sanz , Christian Gutiérrez , Mikel Galar
‹ 上一页 1 2 3 10 下一页 ›