中文
相关论文

相关论文: SoundDet: Polyphonic Moving Sound Event Detection …

200 篇论文

Automatic detection and classification of animal sounds has many applications in biodiversity monitoring and animal behaviour. In the past twenty years, the volume of digitised wildlife sound available has massively increased, and automatic…

Acoustic Scene Classification (ASC) and Sound Event Detection (SED) are two separate tasks in the field of computational sound scene analysis. In this work, we present a new dataset with both sound scene and sound event labels and use this…

音频与语音处理 · 电气工程与系统科学 2019-07-02 Helen L. Bear , Ines Nolasco , Emmanouil Benetos

Time delay estimation or Time-Difference-Of-Arrival estimates is a critical component for multiple localization applications such as multilateration, direction of arrival, and self-calibration. The task is to estimate the time difference…

声音 · 计算机科学 2024-11-21 Erik Tegler , Magnus Oskarsson , Kalle Åström

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains. In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring…

音频与语音处理 · 电气工程与系统科学 2019-07-18 Jee-weon Jung , Hee-Soo Heo , Ju-ho Kim , Hye-jin Shim , Ha-Jin Yu

In the field of acoustic scene analysis, this paper presents a novel approach to find spatio-temporal latent representations from in-the-wild audio data. By using WE-LIVE, an in-house collected dataset that includes audio recordings in…

音频与语音处理 · 电气工程与系统科学 2024-12-11 Claudia Montero-Ramírez , Esther Rituerto-González , Carmen Peláez-Moreno

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by…

音频与语音处理 · 电气工程与系统科学 2025-09-24 Jiarui Hai , Mounya Elhilali

Objective: Ultrasound Shear Wave Elastography (SWE) demonstrates great potential in assessing soft-tissue pathology by mapping tissue stiffness, which is linked to malignancy. Traditional SWE methods have shown promise in estimating tissue…

图像与视频处理 · 电气工程与系统科学 2025-05-27 Ahsan Habib Akash , MD Jahin Alam , Md. Kamrul Hasan

This paper presents a methodology for early detection of audio events from audio streams. Early detection is the ability to infer an ongoing event during its initial stage. The proposed system consists of a novel inference step coupled with…

声音 · 计算机科学 2019-04-09 Huy Phan , Philipp Koch , Ian McLoughlin , Alfred Mertins

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose…

音频与语音处理 · 电气工程与系统科学 2025-01-13 Han Yin , Jisheng Bai , Yang Xiao , Hui Wang , Siqi Zheng , Yafeng Chen , Rohan Kumar Das , Chong Deng , Jianfeng Chen

Polyphonic sound event localization and detection (SELD) aims at detecting types of sound events with corresponding temporal activities and spatial locations. In this paper, a track-wise ensemble event independent network with a novel data…

声音 · 计算机科学 2022-03-22 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

In recent years, anomaly events detection in crowd scenes attracts many researchers' attention, because of its importance to public safety. Existing methods usually exploit visual information to analyze whether any abnormal events have…

计算机视觉与模式识别 · 计算机科学 2021-10-29 Junyu Gao , Maoguo Gong , Xuelong Li

We introduce a U-net model for 360{\deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth…

音频与语音处理 · 电气工程与系统科学 2026-03-30 Belman Jahir Rodriguez , Sergio F. Chevtchenko , Marcelo Herrera Martinez , Yeshwanth Bethi , Saeed Afshar

Sound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localization presents additional challenges, especially when large…

音频与语音处理 · 电气工程与系统科学 2018-11-12 Sandeep Kothinti , Keisuke Imoto , Debmalya Chakrabarty , Gregory Sell , Shinji Watanabe , Mounya Elhilali

This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition on each event. Exploiting the lightweight nature of Temporal Convolution…

Partial deepfake speech detection requires identifying manipulated regions that may occur within short temporal portions of an otherwise bona fide utterance, making the task particularly challenging for conventional utterance-level…

声音 · 计算机科学 2026-04-06 Inbal Rimon , Oren Gal , Haim Permuter

Sound event localization and detection (SELD) consists of two subtasks, which are sound event detection and direction-of-arrival estimation. While sound event detection mainly relies on time-frequency patterns to distinguish different sound…

音频与语音处理 · 电气工程与系统科学 2022-06-07 Thi Ngoc Tho Nguyen , Karn N. Watcharasupat , Ngoc Khanh Nguyen , Douglas L. Jones , Woon-Seng Gan

Change detection in remote sensing imagery is essential for a variety of applications such as urban planning, disaster management, and climate research. However, existing methods for identifying semantically changed areas overlook the…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Maximilian Bernhard , Niklas Strauß , Matthias Schubert

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and…

音频与语音处理 · 电气工程与系统科学 2024-06-24 Ya Jiang , Qing Wang , Jun Du , Maocheng Hu , Pengfei Hu , Zeyan Liu , Shi Cheng , Zhaoxu Nian , Yuxuan Dong , Mingqi Cai , Xin Fang , Chin-Hui Lee

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Yidi Li , Hong Liu , Bing Yang