English
Related papers

Related papers: Spatial Data Augmentation with Simulated Room Impu…

200 papers

A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks. While there have been several augmentation methods shown to…

Sound · Computer Science 2021-08-09 Gwantae Kim , David K. Han , Hanseok Ko

This paper describes the synthesis of the room acoustics challenge as a part of the generative data augmentation workshop at ICASSP 2025. The challenge defines a unique generative task that is designed to improve the quantity and diversity…

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-18 Yicheng Du , Aditya Arie Nugraha , Kouhei Sekiguchi , Yoshiaki Bando , Mathieu Fontaine , Kazuyoshi Yoshii

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

In time-critical eXtended reality (XR) scenarios where users must rapidly reorient their attention to hazards, alerts, or instructions while engaged in a primary task, spatial audio can provide an immediate directional cue without occupying…

Human-Computer Interaction · Computer Science 2026-05-08 Yoonsang Kim , Swapnil Dey , Arie Kaufman

Referring Image Segmentation (RIS) is an advanced vision-language task that involves identifying and segmenting objects within an image as described by free-form text descriptions. While previous studies focused on aligning visual and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Minhyun Lee , Seungho Lee , Song Park , Dongyoon Han , Byeongho Heo , Hyunjung Shim

This technical report details our systems submitted for Task 3 of the DCASE 2024 Challenge: Audio and Audiovisual Sound Event Localization and Detection (SELD) with Source Distance Estimation (SDE). We address only the audio-only SELD with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Jun Wei Yeow , Ee-Leng Tan , Jisheng Bai , Santi Peksi , Woon-Seng Gan

This paper investigates the use of intelligent reflecting surfaces (IRS) to assist cellular communications and radar sensing operations in a communications and sensing setup. The IRS dynamically allocates reflecting elements to…

Signal Processing · Electrical Eng. & Systems 2025-10-24 Daniyal Munir , Atta Ullah , Danish Mehmood Mughal , Min Young Chung , Hans D. Schotten

In this paper, we consider a challenging secure wireless sensing scenario where a legitimate radar station (LRS) intends to detect a target at unknown location in the presence of an unauthorized radar station (URS). We aim to enhance the…

Information Theory · Computer Science 2023-08-08 Xiaodan Shao , Rui Zhang

Visual surveillance aims to stably detect a foreground object using a continuous image acquired from a fixed camera. Recent deep learning methods based on supervised learning show superior performance compared to classical background…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Jae-Yeul Kim , Jong-Eun Ha

In this paper, we consider a multi-user multiple-input multiple-output (MIMO) system aided by multiple intelligent reflecting surfaces (IRSs) that are deployed to increase the coverage and, possibly, the rank of the channel. We propose an…

Information Theory · Computer Science 2020-12-22 Andrea Abrardo , Davide Dardari , Marco Di Renzo

In this paper, we introduce an intelligent reflecting surface (IRS) to provide a programmable wireless environment for physical layer security. By adjusting the reflecting coefficients, the IRS can change the attenuation and scattering of…

Signal Processing · Electrical Eng. & Systems 2019-06-21 Jie Chen , Ying-Chang Liang , Yiyang Pei , Huayan Guo

In this paper we present a novel algorithm for improved block-online supervised acoustic system identification in adverse noise scenarios by exploiting prior knowledge about the space of Room Impulse Responses (RIRs). The method is based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-06 Thomas Haubner , Andreas Brendel , Walter Kellermann

Soundscape augmentation is an emerging approach for noise mitigation by introducing additional sounds known as "maskers" to increase acoustic comfort. Traditionally, the choice of maskers is often predicated on expert guidance or post-hoc…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-04 Trevor Wong , Karn N. Watcharasupat , Bhan Lam , Kenneth Ooi , Zhen-Ting Ong , Furi Andi Karnapi , Woon-Seng Gan

Intelligent reflecting surface (IRS) has emerged as a promising technology to reconfigure the radio propagation environment by dynamically controlling wireless signal's amplitude and/or phase via a large number of reflecting elements. In…

Signal Processing · Electrical Eng. & Systems 2022-05-26 Xiaodan Shao , Changsheng You , Wenyan Ma , Xiaoming Chen , Rui Zhang

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Mauricio Flores-Vargas , Enda Bates , Rachel McDonnell

Intelligent reflecting surface (IRS) is envisioned to have abundant applications in future wireless networks by smartly reconfiguring the signal propagation for performance enhancement. Specifically, an IRS consists of a large number of…

Information Theory · Computer Science 2018-09-13 Qingqing Wu , Rui Zhang

Recent developments in multi-agent imitation learning have shown promising results for modeling the behavior of human drivers. However, it is challenging to capture emergent traffic behaviors that are observed in real-world datasets. Such…

Psychoacoustic studies have shown that locally-time reversed (LTR) speech, i.e., signal samples time-reversed within a short segment, can be accurately recognised by human listeners. This study addresses the question of how well a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-12 Si-Ioi Ng , Tan Lee