English
Related papers

Related papers: Robust Sound Source Tracking Using SRP-PHAT and 3D…

200 papers

Modeling late reverberation in real-time interactive applications is a challenging task when multiple sound sources and listeners are present in the same environment. This is especially problematic when the environment is geometrically…

Sound · Computer Science 2025-10-14 Matteo Scerbo , Sebastian J. Schlecht , Randall Ali , Lauri Savioja , Enzo De Sena

We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into…

Sound · Computer Science 2021-08-24 Yuhang He , Niki Trigoni , Andrew Markham

We present a detail-driven deep neural network for point set upsampling. A high-resolution point set is essential for point-based rendering and surface reconstruction. Inspired by the recent success of neural image super-resolution…

Computer Vision and Pattern Recognition · Computer Science 2019-03-22 Wang Yifan , Shihao Wu , Hui Huang , Daniel Cohen-Or , Olga Sorkine-Hornung

We consider the problem of estimating the directions of arrival (DOAs) of multiple sources from a single snapshot of an antenna array, a task with many practical applications. In such settings, the classical Bartlett beamformer is commonly…

Signal Processing · Electrical Eng. & Systems 2025-09-22 Lioz Berman , Sharon Gannot , Tom Tirer

In traditional sound event localization and detection (SELD) tasks, the focus is typically on sound event detection (SED) and direction-of-arrival (DOA) estimation, but they fall short of providing full spatial information about the sound…

Sound · Computer Science 2025-01-22 Yuxuan Dong , Qing Wang , Hengyi Hong , Ya Jiang , Shi Cheng

Environmental sound classification systems often do not perform robustly across different sound classification tasks and audio signals of varying temporal structures. We introduce a multi-stream convolutional neural network with temporal…

Sound · Computer Science 2019-01-28 Xinyu Li , Venkata Chebiyyam , Katrin Kirchhoff

Recently, progressive learning has shown its capacity to improve speech quality and speech intelligibility when it is combined with deep neural network (DNN) and long short-term memory (LSTM) based monaural speech enhancement algorithms,…

Sound · Computer Science 2020-01-14 Andong Li , Minmin Yuan , Chengshi Zheng , Xiaodong Li

More powerful feature representations derived from deep neural networks benefit visual tracking algorithms widely. However, the lack of exploitation on temporal information prevents tracking algorithms from adapting to appearances changing…

Computer Vision and Pattern Recognition · Computer Science 2019-08-05 Tao Hu , Lichao Huang , Xianming Liu , Han Shen

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Yidi Li , Hong Liu , Hao Tang

Radio map, or pathloss map prediction, is a crucial method for wireless network modeling and management. By leveraging deep learning to construct pathloss patterns from geographical maps, an accurate digital replica of the transmission…

Signal Processing · Electrical Eng. & Systems 2025-01-14 Yuxuan Li , Cheng Zhang , Wen Wang , Yongming Huang

Direction-of-arrival (DoA) is a critical parameter in wireless channel estimation. With the ever-increasing requirement of high data rate and ubiquitous devices in wireless communication systems, effective wideband DoA estimation is…

Signal Processing · Electrical Eng. & Systems 2023-11-21 Xiaorui Ding , Wenbo Xu , Yue Wang

As three-dimensional (3D) data acquisition devices become increasingly prevalent, the demand for 3D point cloud transmission is growing. In this study, we introduce a semantic-aware communication system for robust point cloud classification…

Signal Processing · Electrical Eng. & Systems 2023-06-26 Tianxiao Han , Kaiyi Chi , Qianqian Yang , Zhiguo Shi

Sound source localization (SSL) adds a spatial dimension to auditory perception, allowing a system to pinpoint the origin of speech, machinery noise, warning tones, or other acoustic events, capabilities that facilitate robot navigation,…

Robotics · Computer Science 2025-08-29 Reza Jalayer , Masoud Jalayer , Amirali Baniasadi

Resolving transient atomic configurations in non-crystalline or dynamic environments remains a fundamental bottleneck in the physical sciences. While X-ray absorption spectroscopy (XAS) is a premier probe of local structure, inverting…

Materials Science · Physics 2026-03-31 Suyang Zhong , Boying Huang , Pengwei Xu , Fanjie Xu , Yuhao Zhao , Jun Cheng , Fujie Tang , Weinan E , Zhong-Qun Tian

Radio maps (RMs) serve as a critical foundation for enabling environment-aware wireless communication, as they provide the spatial distribution of wireless channel characteristics. Despite recent progress in RM construction using…

Machine Learning · Computer Science 2025-07-17 Xiucheng Wang , Qiming Zhang , Nan Cheng , Junting Chen , Zezhong Zhang , Zan Li , Shuguang Cui , Xuemin Shen

This work presents a novel technique that performs both orientation and distance localization of a sound source in a three-dimensional (3D) space using only the interaural time difference (ITD) cue, generated by a newly-developed…

Sound · Computer Science 2018-04-11 Deepak Gala , Nathan Lindsay , Liang Sun

As we interact with the world, for example when we communicate with our colleagues in a large open space or meeting room, we continuously analyse the surrounding environment and, in particular, localise and recognise acoustic events. While…

Sound · Computer Science 2019-04-02 Pawel Swietojanski , Ondrej Miksik

In this paper we propose a novel environmental sound classification approach incorporating unsupervised feature learning from codebook via spherical $K$-Means++ algorithm and a new architecture for high-level data augmentation. The audio…

Machine Learning · Computer Science 2019-11-26 Mohammad Esmaeilpour , Patrick Cardinal , Alessandro Lameiras Koerich

DNN-based methods have shown high performance in sound event localization and detection(SELD). While in real spatial sound scenes, reverberation and the imbalanced presence of various sound events increase the complexity of the SELD task.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-18 Siwei Huang , Jianfeng Chen , Jisheng Bai , Yafei Jia , Dongzhe Zhang

In this paper, we investigate the impact of different standard environmental sound representations (spectrograms) on the recognition performance and adversarial attack robustness of a victim residual convolutional neural network. Averaged…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Mohammad Esmaeilpour , Patrick Cardinal , Alessandro Lameiras Koerich
‹ Prev 1 8 9 10 Next ›