中文
相关论文

相关论文: SS-BRPE: Self-Supervised Blind Room Parameter Esti…

200 篇论文

Smart glasses are increasingly recognized as a key medium for augmented reality, offering a hands-free platform with integrated microphones and non-ear-occluding loudspeakers to seamlessly mix virtual sound sources into the real-world…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Thomas Deppisch , Nils Meyer-Kahlen , Sebastià V. Amengual Garí

In-context learning with attention enables large neural networks to make context-specific predictions by selectively focusing on relevant examples. Here, we adapt this idea to supervised learning procedures such as lasso regression and…

机器学习 · 统计学 2025-12-11 Erin Craig , Robert Tibshirani

This paper introduces a new framework for supervised sound source localization referred to as virtually-supervised learning. An acoustic shoe-box room simulator is used to generate a large number of binaural single-source audio scenes.…

声音 · 计算机科学 2017-03-21 Saurabh Kataria , Clément Gaultier , Antoine Deleforge

Anomalous sound detection for machine condition monitoring has great potential in the development of Industry 4.0. However, these anomalous sounds of machines are usually unavailable in normal conditions. Therefore, the models employed have…

音频与语音处理 · 电气工程与系统科学 2023-02-17 Jisheng Bai , Jianfeng Chen , Mou Wang , Muhammad Saad Ayub , Qingli Yan

Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognition without labeled…

计算与语言 · 计算机科学 2025-01-14 Haoyu Wang , Wei-Qiang Zhang , Hongbin Suo , Yulong Wan

As a type of multi-dimensional sequential data, the spatial and temporal dependencies of electroencephalogram (EEG) signals should be further investigated. Thus, in this paper, we propose a novel spatial-temporal progressive attention model…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Yang Li , Wei Liu , Tianzhi Feng , Fu Li , Chennan Wu , Boxun Fu , Zhifu Zhao , Xiaotian Wang , Guangming Shi

Simulation-based inference (SBI) is a statistical inference approach for estimating latent parameters of a physical system when the likelihood is intractable but simulations are available. In practice, SBI is often hindered by model…

机器学习 · 计算机科学 2025-10-22 Ortal Senouf , Antoine Wehenkel , Cédric Vincent-Cuaz , Emmanuel Abbé , Pascal Frossard

Any audio recording encapsulates the unique fingerprint of the associated acoustic environment, namely the background noise and reverberation. Considering the scenario of a room equipped with a fixed smart speaker device with one or more…

音频与语音处理 · 电气工程与系统科学 2022-12-05 Francesco Nespoli , Daniel Barreda , Patrick A. Naylor

Recently, self-supervised pre-training has gained success in automatic speech recognition (ASR). However, considering the difference between speech accents in real scenarios, how to identify accents and use accent features to improve ASR is…

音频与语音处理 · 电气工程与系统科学 2021-09-16 Keqi Deng , Songjun Cao , Long Ma

Speaker-attributed automatic speech recognition (SA-ASR) improves the accuracy and applicability of multi-speaker ASR systems in real-world scenarios by assigning speaker labels to transcribed texts. However, SA-ASR poses unique challenges…

音频与语音处理 · 电气工程与系统科学 2023-09-29 Xiang Lyu , Yuhang Cao , Qing Wang , Jingjing Yin , Yuguang Yang , Pengpeng Zou , Yanni Hu , Heng Lu

Indoor localization systems in care facilities enable optimization of staff allocation, workload management, and quality of care delivery. Traditional machine learning approaches to Bluetooth Low Energy (BLE)-based localization treat each…

机器学习 · 计算机科学 2026-03-24 Minh Triet Pham , Quynh Chi Dang , Le Nhat Tan

Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural…

声音 · 计算机科学 2024-12-31 Sangwook Park , David K. Han , Mounya Elhilali

Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech…

声音 · 计算机科学 2024-07-02 Hok-Shing Lau , Mark Huntly , Nathon Morgan , Adesua Iyenoma , Biao Zeng , Tim Bashford

In this paper we present a novel algorithm for improved block-online supervised acoustic system identification in adverse noise scenarios by exploiting prior knowledge about the space of Room Impulse Responses (RIRs). The method is based on…

音频与语音处理 · 电气工程与系统科学 2020-07-06 Thomas Haubner , Andreas Brendel , Walter Kellermann

Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge. Data-driven supervised approaches, including ones based on…

Stress is a complex issue with wide-ranging physical and psychological impacts on human daily performance. Specifically, acute stress detection is becoming a valuable application in contextual human understanding. Two common approaches to…

机器学习 · 计算机科学 2022-03-21 Van-Tu Ninh , Manh-Duy Nguyen , Sinéad Smyth , Minh-Triet Tran , Graham Healy , Binh T. Nguyen , Cathal Gurrin

Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase,…

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

信息检索 · 计算机科学 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer

Audio-visual speech enhancement (AVSE) has been found to be particularly useful at low signal-to-noise (SNR) ratios due to the immunity of the visual features to acoustic noise. However, a significant gap exists in AVSE methods tailored to…

音频与语音处理 · 电气工程与系统科学 2025-10-21 Danielle Yaffe , Ferdinand Campe , Prachi Sharma , Dorothea Kolossa , Boaz Rafaely

Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units based on e.g. byte-pair encoding (BPE). The mapping from…

音频与语音处理 · 电气工程与系统科学 2021-04-16 Mohammad Zeineldeen , Albert Zeyer , Wei Zhou , Thomas Ng , Ralf Schlüter , Hermann Ney