中文
相关论文

相关论文: Autoregressive Guidance of Deep Spatially Selectiv…

200 篇论文

Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Jakob Kienegger , Alina Mannanova , Huajian Fang , Timo Gerkmann

Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in…

音频与语音处理 · 电气工程与系统科学 2026-01-22 Jakob Kienegger , Timo Gerkmann

Recent speaker extraction methods using deep non-linear spatial filtering perform exceptionally well when the target direction is known and stationary. However, spatially dynamic scenarios are considerably more challenging due to…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Jakob Kienegger , Timo Gerkmann

Audio-visual speaker tracking has drawn increasing attention over the past few years due to its academic values and wide applications. Audio and visual modalities can provide complementary information for localization and tracking. With…

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

声音 · 计算机科学 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

In a scenario with multiple persons talking simultaneously, the spatial characteristics of the signals are the most distinct feature for extracting the target signal. In this work, we develop a deep joint spatial-spectral non-linear filter…

音频与语音处理 · 电气工程与系统科学 2023-04-05 Kristina Tesch , Timo Gerkmann

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are…

音频与语音处理 · 电气工程与系统科学 2025-06-26 Taous Iatariene , Can Cui , Alexandre Guérin , Romain Serizel

This paper introduces an alternative approach to sampling from autoregressive models. Autoregressive models are typically sampled sequentially, according to the transition dynamics defined by the model. Instead, we propose a sampling…

机器学习 · 计算机科学 2021-12-20 Vivek Jayaram , John Thickstun

Autoregressive models are widely used for tasks such as image and audio generation. The sampling process of these models, however, does not allow interruptions and cannot adapt to real-time computational resources. This challenge impedes…

机器学习 · 计算机科学 2021-02-24 Yilun Xu , Yang Song , Sahaj Garg , Linyuan Gong , Rui Shu , Aditya Grover , Stefano Ermon

Data fusion plays an important role in many technical applications that require efficient processing of multimodal sensory observations. A prominent example is audiovisual signal processing, which has gained increasing attention in…

计算机视觉与模式识别 · 计算机科学 2019-03-15 Christopher Schymura , Dorothea Kolossa

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight visual frontend based…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Zexu Pan , Wupeng Wang , Shengkui Zhao , Chong Zhang , Kun Zhou , Yukun Ma , Bin Ma

Extended object tracking methods based on random matrices, founded on Bayesian filters, have been able to achieve efficient recursive processes while jointly estimating the kinematic states and extension of the targets. Existing random…

信号处理 · 电气工程与系统科学 2025-05-27 Zhixing Wang , Le Zheng , Shi Yan , Ruud J. G. van Sloun , Nir Shlezinger , Yonina C. Eldar

This paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations. In contrast to other speech representation learning methods that aim to remove noise or speaker variabilities, ours is…

计算与语言 · 计算机科学 2019-06-20 Yu-An Chung , Wei-Ning Hsu , Hao Tang , James Glass

For compressive sensing of dynamic sparse signals, we develop an iterative pursuit algorithm. A dynamic sparse signal process is characterized by varying sparsity patterns over time/space. For such signals, the developed algorithm is able…

统计理论 · 数学 2012-10-15 Dave Zachariah , Saikat Chatterjee , Magnus Jansson

Modeling nonstationary processes is of paramount importance to many scientific disciplines including environmental science, ecology, and finance, among others. Consequently, flexible methodology that provides accurate estimation across a…

统计方法学 · 统计学 2014-08-13 Wen-Hsi Yang , Scott H. Holan , Christopher K. Wikle

Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model that maps a reference…

音频与语音处理 · 电气工程与系统科学 2022-02-17 Mohamed Elminshawi , Wolfgang Mack , Emanuël A. P. Habets

Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal enhancement. However, the textbook solutions for optimal…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Reinhold Haeb-Umbach , Tomohiro Nakatani , Marc Delcroix , Christoph Boeddeker , Tsubasa Ochiai

Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning-based tracking with beamforming…

This paper investigates continuous representations of steering vectors over frequency and microphone/source positions for augmented listening (e.g., spatial filtering and binaural rendering), enabling user-parameterized control of the…

音频与语音处理 · 电气工程与系统科学 2026-04-17 Diego Di Carlo , Shoichi Koyama , Nugraha Aditya Arie , Fontaine Mathieu , Bando Yoshiaki , Yoshii Kazuyoshi
‹ 上一页 1 2 3 10 下一页 ›