中文
相关论文

相关论文: SoundDet: Polyphonic Moving Sound Event Detection …

200 篇论文

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to produce a processed…

音频与语音处理 · 电气工程与系统科学 2018-09-18 Francois G. Germain , Qifeng Chen , Vladlen Koltun

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality, existing methods are rarely leverage language model for…

声音 · 计算机科学 2024-08-06 Hualei Wang , Jianguo Mao , Zhifang Guo , Jiarui Wan , Hong Liu , Xiangdong Wang

Automatic event detection from time series signals has wide applications, such as abnormal event detection in video surveillance and event detection in geophysical data. Traditional detection methods detect events primarily by the use of…

机器学习 · 计算机科学 2018-09-26 Yue Wu , Youzuo Lin , Zheng Zhou , David Chas Bolton , Ji Liu , Paul Johnson

Sound event detection (SED) is a task to detect sound events in an audio recording. One challenge of the SED task is that many datasets such as the Detection and Classification of Acoustic Scenes and Events (DCASE) datasets are weakly…

声音 · 计算机科学 2020-08-25 Qiuqiang Kong , Yong Xu , Wenwu Wang , Mark D. Plumbley

Event-based keypoint detection and matching holds significant potential, enabling the integration of event sensors into highly optimized Visual SLAM systems developed for frame cameras over decades of research. Unfortunately, existing…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yannick Burkhardt , Simon Schaefer , Stefan Leutenegger

Dynamic Vision Sensor (DVS) can asynchronously output the events reflecting apparent motion of objects with microsecond resolution, and shows great application potential in monitoring and other fields. However, the output event stream of…

计算机视觉与模式识别 · 计算机科学 2022-03-23 Jinze Chen , Yang Wang , Yang Cao , Feng Wu , Zheng-Jun Zha

Recognizing the sounding objects in scenes is a longstanding objective in embodied AI, with diverse applications in robotics and AR/VR/MR. To that end, Audio-Visual Segmentation (AVS), taking as condition an audio signal to identify the…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Artem Sokolov , Swapnil Bhosale , Xiatian Zhu

Polyphonic sound event detection and localization (SELD) task is challenging because it is difficult to jointly optimize sound event detection (SED) and direction-of-arrival (DOA) estimation in the same network. We propose a general network…

音频与语音处理 · 电气工程与系统科学 2020-11-17 Thi Ngoc Tho Nguyen , Ngoc Khanh Nguyen , Huy Phan , Lam Pham , Kenneth Ooi , Douglas L. Jones , Woon-Seng Gan

Performing sound event detection on real-world recordings often implies dealing with overlapping target sound events and non-target sounds, also referred to as interference or noise. Until now these problems were mainly tackled at the…

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions,…

声音 · 计算机科学 2025-03-18 Chen Liu , Liying Yang , Peike Li , Dadong Wang , Lincheng Li , Xin Yu

In this paper, we propose a stacked convolutional and recurrent neural network (CRNN) with a 3D convolutional neural network (CNN) in the first layer for the multichannel sound event detection (SED) task. The 3D CNN enables the network to…

声音 · 计算机科学 2018-01-30 Sharath Adavanne , Archontis Politis , Tuomas Virtanen

This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing and retrieval of similar audio events using an audio example…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Sagar Dutta , Vipul Arora

Sound event localization and detection (SELD) involves sound event detection (SED) and direction of arrival (DoA) estimation tasks. SED mainly relies on temporal dependencies to distinguish different sound classes, while DoA estimation…

音频与语音处理 · 电气工程与系统科学 2024-03-21 Weiming Huang , Qinghua Huang , Liyan Ma , Chuan Wang

Sound event localization and detection (SELD) is an emerging research topic that aims to unify the tasks of sound event detection and direction-of-arrival estimation. As a result, SELD inherits the challenges of both tasks, such as noise,…

音频与语音处理 · 电气工程与系统科学 2021-10-28 Thi Ngoc Tho Nguyen , Karn N. Watcharasupat , Zhen Jian Lee , Ngoc Khanh Nguyen , Douglas L. Jones , Woon Seng Gan

Polyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requirement of most…

音频与语音处理 · 电气工程与系统科学 2022-06-07 Thi Ngoc Tho Nguyen , Douglas L. Jones , Karn N. Watcharasupat , Huy Phan , Woon-Seng Gan

The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Chengjian Feng , Yujie Zhong , Zequn Jie , Xiangxiang Chu , Haibing Ren , Xiaolin Wei , Weidi Xie , Lin Ma

Event-based vision sensors, such as the Dynamic Vision Sensor (DVS), are ideally suited for real-time motion analysis. The unique properties encompassed in the readings of such sensors provide high temporal resolution, superior sensitivity…

计算机视觉与模式识别 · 计算机科学 2020-01-14 Anton Mitrokhin , Cornelia Fermuller , Chethan Parameshwara , Yiannis Aloimonos

In recent years, many deep learning techniques for single-channel sound source separation have been proposed using recurrent, convolutional and transformer networks. When multiple microphones are available, spatial diversity between…

音频与语音处理 · 电气工程与系统科学 2022-08-23 Ali Aroudi , Stefan Uhlich , Marc Ferras Font

Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been…

声音 · 计算机科学 2024-09-23 Xutong Jin , Chenxi Xu , Ruohan Gao , Jiajun Wu , Guoping Wang , Sheng Li

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

声音 · 计算机科学 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler