中文
相关论文

相关论文: MB-RIRs: a Synthetic Room Impulse Response Dataset…

200 篇论文

This work presents our end-to-end (E2E) automatic speech recognition (ASR) model targetting at robust speech recognition, called Integraded speech Recognition with enhanced speech Input for Self-supervised learning representation (IRIS).…

声音 · 计算机科学 2022-04-04 Xuankai Chang , Takashi Maekaku , Yuya Fujita , Shinji Watanabe

An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Traditional methods for constructing acoustic models involve expensive and…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Arjun Somayazulu , Sagnik Majumder , Changan Chen , Kristen Grauman

Intelligent Reflecting Surface (IRS) technology is revolutionizing wireless communications by shifting from channel adaptation to a responsive wireless environment. This paper introduces a multi-IRS assisted millimeter wave (mm-wave)…

信号处理 · 电气工程与系统科学 2024-01-04 Alireza Qazavi Khorasgani , Foroogh S. Tabataba , Mehdi Naderi Soorki , Mohammad Sadegh Fazel

This paper proposes a method for estimating the norms of a system in a pure data-driven fashion based on their identified Impulse Response (IR) coefficients. The calculation of norms is briefly reviewed and the main expressions for the…

系统与控制 · 电气工程与系统科学 2021-11-09 L. V. Fiorio , C. L. Remes , L. Campestrini , Y. R. de Novaes

Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are available for…

声音 · 计算机科学 2025-08-05 Yizhu Jin , Zhen Ye , Zeyue Tian , Haohe Liu , Qiuqiang Kong , Yike Guo , Wei Xue

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

In this paper, we propose to utilise diffusion models for data augmentation in speech emotion recognition (SER). In particular, we present an effective approach to utilise improved denoising diffusion probabilistic models (IDDPM) to…

声音 · 计算机科学 2023-05-22 Ibrahim Malik , Siddique Latif , Raja Jurdak , Björn Schuller

We propose a method for generating low-frequency compensated synthetic impulse responses that improve the performance of far-field speech recognition systems trained on artificially augmented datasets. We design linear-phase filters that…

声音 · 计算机科学 2021-09-28 Zhenyu Tang , Hsien-Yu Meng , Dinesh Manocha

Recent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which tend to suffer from over-fitting in low resource scenarios. One solution to tackle this issue is to…

计算与语言 · 计算机科学 2021-07-14 Nick Rossenbach , Mohammad Zeineldeen , Benedikt Hilmes , Ralf Schlüter , Hermann Ney

Glass surfaces create complex interactions of reflected and transmitted light, making single-image reflection removal (SIRR) challenging. Existing datasets suffer from limited physical realism in synthetic data or insufficient scale in real…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yu Guo , Zhiqiang Lao , Xiyun Song , Yubin Zhou , Heather Yu

The spatial impulse response (SIR) method is a well-known approach to calculate transient acoustic fields of arbitrary-shape transducers. It involves the evaluation of a time-dependent surface integral. Although analytic expressions of the…

数值分析 · 数学 2021-11-01 Dimitris Perdios , Florian Martinez , Marcel Arditi , Jean-Philippe Thiran

Driven by large scale datasets and LLM based architectures, automatic speech recognition (ASR) systems have achieved remarkable improvements in accuracy. However, challenges persist for domain-specific terminology, and short utterances…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Jinming Chen , Lu Wang , Zheshu Song , Wei Deng

In this letter, we investigate the signal-to-interference-plus-noise-ratio (SINR) maximization problem in a multi-user massive multiple-input-multiple-output (massive MIMO) system enabled with multiple reconfigurable intelligent surfaces…

信号处理 · 电气工程与系统科学 2023-04-06 Somayeh Aghashahi , Zolfa Zeinalpour-Yazdi , Aliakbar Tadaion , Mahdi Boloursaz Mashhadi , Ahmed Elzanaty

Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To harden models against…

音频与语音处理 · 电气工程与系统科学 2018-07-19 Davis Liang , Zhiheng Huang , Zachary C. Lipton

Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their…

音频与语音处理 · 电气工程与系统科学 2023-12-01 Sina Hafezi , Alastair H. Moore , Pierre H. Guiraud , Patrick A. Naylor , Jacob Donley , Vladimir Tourbabin , Thomas Lunner

There is an emerging need for comparable data for multi-microphone processing, particularly in acoustic sensor networks. However, commonly available databases are often limited in the spatial diversity of the microphones or only allow for…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Daniel Fejgin , Wiebke Middelberg , Simon Doclo

Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-started generation process that enables realistic texture synthesis and captures the…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Wei Zhu , Kai Zhang , Yu Zheng , Lei Luo , Yong Guo , Jian Yang

Background: An early diagnosis together with an accurate disease progression monitoring of multiple sclerosis is an important component of successful disease management. Prior studies have established that multiple sclerosis is correlated…

音频与语音处理 · 电气工程与系统科学 2021-09-28 Emil Svoboda , Tomáš Bořil , Jan Rusz , Tereza Tykalová , Dana Horáková , Charles R. G. Guttman , Krastan B. Blagoev , Hiroto Hatabu , Vlad I. Valtchinov

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Classroom datasets remain limited and not publicly available, and the absence of dedicated classroom noise or Room…

声音 · 计算机科学 2025-10-03 Ahmed Adel Attia , Jing Liu , Carol Espy Wilson

Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous…

音频与语音处理 · 电气工程与系统科学 2025-04-08 Mauricio Flores-Vargas , Enda Bates , Rachel McDonnell