中文
相关论文

相关论文: Can all variations within the unified mask-based b…

200 篇论文

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal…

声音 · 计算机科学 2021-04-06 Yong Xu , Zhuohuang Zhang , Meng Yu , Shi-Xiong Zhang , Dong Yu

Multichannel linear filters, such as the Multichannel Wiener Filter (MWF) and the Generalized Eigenvalue (GEV) beamformer are popular signal processing techniques which can improve speech recognition performance. In this paper, we present…

声音 · 计算机科学 2017-11-16 Ziteng Wang , Emmanuel Vincent , Romain Serizel , Yonghong Yan

We consider a multiuser multiple-input single-output interference channel where the receivers are characterized by both quality-of-service (QoS) and radio-frequency (RF) energy harvesting (EH) constraints. We consider the power splitting…

信息论 · 计算机科学 2014-07-11 Stelios Timotheou , Ioannis Krikidis , Gan Zheng , Björn Ottersten

Neural network approaches to single-channel speech enhancement have received much recent attention. In particular, mask-based architectures have achieved significant performance improvements over conventional methods. This paper proposes a…

音频与语音处理 · 电气工程与系统科学 2023-09-22 Bengt J. Borgstrom , Michael S. Brandstein

We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture model of the…

声音 · 计算机科学 2019-04-09 Lukas Drude , Jahn Heymann , Reinhold Haeb-Umbach

Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted speech waveform.…

音频与语音处理 · 电气工程与系统科学 2023-03-10 Kai Liu , Ziqing Du , Xucheng Wan , Huan Zhou

Coordinated beamforming (Co-BF) is a key multi-access-point coordination (MAPC) technique for dense Wi-Fi deployments, but its performance can be hindered by the large channel state information (CSI) feedback required through channel…

网络与互联网体系结构 · 计算机科学 2026-04-16 Ibrahim Aboushehada , Boris Bellalta , Giovanni Geraci , Lorenzo Galati Giordano

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no…

音频与语音处理 · 电气工程与系统科学 2021-08-05 Tomohiro Nakatani , Rintaro Ikeshita , Keisuke Kinoshita , Hiroshi Sawada , Shoko Araki

Multi-frequency interferometry (MFI) is well known as an accurate phase-based measurement scheme. The paper reveals the inherent relationship of the unambiguous measurement range (UMR), the outlier probability, the MSE performance with the…

信息论 · 计算机科学 2012-10-09 Li Wei , Wangdong Qi

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

音频与语音处理 · 电气工程与系统科学 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

Recently, frequency domain all-neural beamforming methods have achieved remarkable progress for multichannel speech separation. In parallel, the integration of time domain network structure and beamforming also gains significant attention.…

声音 · 计算机科学 2022-12-27 Rongzhi Gu , Shi-Xiong Zhang , Yuexian Zou , Dong Yu

Federated learning (FL) has emerged as an appealing machine learning approach to deal with massive raw data generated at multiple mobile devices, {which needs to aggregate the training model parameter of every mobile device at one base…

机器学习 · 计算机科学 2023-08-21 Xuming An , Rongfei Fan , Shiyuan Zuo , Han Hu , Hai Jiang , Ning Zhang

Target sound extraction (TSE) is the task of extracting a target sound specified by a query from an audio mixture. Much prior research has focused on the problem setting under the Fully Matched Query (FMQ) condition, where the query…

音频与语音处理 · 电气工程与系统科学 2025-09-11 Ryo Sato , Chiho Haruta , Nobuhiko Hiruma , Keisuke Imoto

The state-of-art methods for acoustic beamforming in multi-channel ASR are based on a neural mask estimator that predicts the presence of speech and noise. These models are trained using a paired corpus of clean and noisy recordings…

音频与语音处理 · 电气工程与系统科学 2019-12-02 Rohit Kumar , Anirudh Sreeram , Anurenjan Purushothaman , Sriram Ganapathy

Complexity reduction of optimal linear receiver is considered in a scenario where both the number of single-antenna user equipments (UEs) $K$ and base station (BS) antennas $N$ are large. Two-stage beamforming (TSB) greatly alleviates the…

信号处理 · 电气工程与系统科学 2019-12-03 Hossein Asgharimoghaddam , Antti Tölli

Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE achieve strong results.…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Aviv Navon , Aviv Shamsian , Yael Segal-Feldman , Neta Glazer , Gil Hetz , Joseph Keshet

This paper investigates the performance of optimal single stream beamforming schemes in multiple-input multiple-output (MIMO) dual-hop amplify-and-forward (AF) systems. Assuming channel state information is not available at the source and…

信息论 · 计算机科学 2016-11-17 Caijun Zhong , Tharm Ratnarajah , Shi Jin , Kai Kit Wong

We present Masked Frequency Modeling (MFM), a unified frequency-domain-based approach for self-supervised pre-training of visual models. Instead of randomly inserting mask tokens to the input embeddings in the spatial domain, in this paper,…

计算机视觉与模式识别 · 计算机科学 2023-04-26 Jiahao Xie , Wei Li , Xiaohang Zhan , Ziwei Liu , Yew Soon Ong , Chen Change Loy

Acoustic beamformers have been widely used to enhance audio signals. The best current methods are DNN-powered variants of the generalized eigenvalue beamformer, and DNN-based filterestimation methods that directly compute beamforming…

声音 · 计算机科学 2020-03-03 Yuichiro Koyama , Bhiksha Raj

Multi-speaker speech recognition has been one of the keychallenges in conversation transcription as it breaks the singleactive speaker assumption employed by most state-of-the-artspeech recognition systems. Speech separation is consideredas…

音频与语音处理 · 电气工程与系统科学 2020-09-08 Jian Wu , Zhuo Chen , Jinyu Li , Takuya Yoshioka , Zhili Tan , Ed Lin , Yi Luo , Lei Xie