中文
相关论文

相关论文: RTF-steered binaural MVDR beamforming incorporatin…

200 篇论文

In room acoustic environments, the Relative Transfer Functions (RTFs) are controlled by few underlying modes of variability. Accordingly, they are confined to a low-dimensional manifold. In this letter, we investigate a RTF inverse…

声音 · 计算机科学 2017-10-26 Ziteng Wang , Emmanuel Vincent , Yonghong Yan

Telepresence aims to create an immersive but virtual experience of the audio and visual scene at the far end for users at the near end. In this contribution, we propose an array-based binaural rendering system that converts the array…

音频与语音处理 · 电气工程与系统科学 2023-03-07 Yicheng Hsu , Chenghumg Ma , Mingsian R. Bai

This work proposes a multichannel speech separation method with narrow-band Conformer (named NBC). The network is trained to learn to automatically exploit narrow-band speech separation information, such as spatial vector clustering of…

声音 · 计算机科学 2022-07-04 Changsheng Quan , Xiaofei Li

This paper deals with joint source and relay beamforming (BF) design for an amplify-and-forward (AF) multi-antenna multirelay network. Considering that the channel state information (CSI) from relays to destination is imperfect, we aim to…

信号处理 · 电气工程与系统科学 2020-03-05 Hongying Tang , Wen Chen , Jun Li , Haibin Wan

In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and ease of learning. However, because the ultimate goal is to…

声音 · 计算机科学 2025-08-28 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

We investigate an alternative solution method to the joint signal-beamformer optimization problem considered by Setlur and Rangaswamy[1]. First, we directly demonstrate that the problem, which minimizes the received noise, interference, and…

系统与控制 · 计算机科学 2018-02-14 Sean M. O'Rourke , Pawan Setlur , Muralidhar Rangaswamy , A. Lee Swindlehurst

Dereverberation of recorded speech signals is one of the most pertinent problems in speech processing. In the present work, the objective is to understand and implement dereverberation techniques that aim at enhancing the magnitude…

音频与语音处理 · 电气工程与系统科学 2025-12-30 Dhruv Nigam

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

音频与语音处理 · 电气工程与系统科学 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

Measuring personal head-related transfer functions (HRTFs) is essential in binaural audio. Personal HRTFs are not only required for binaural rendering and for loudspeaker-based binaural reproduction using crosstalk cancellation, but they…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Tobias Kabzinski , Peter Jax

Binaural rendering of ambisonic signals is of broad interest to virtual reality and immersive media. Conventional methods often require manually measured Head-Related Transfer Functions (HRTFs). To address this issue, we collect a paired…

声音 · 计算机科学 2022-11-07 Yin Zhu , Qiuqiang Kong , Junjie Shi , Shilei Liu , Xuzhou Ye , Ju-chiang Wang , Junping Zhang

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast…

音频与语音处理 · 电气工程与系统科学 2021-01-20 Eunwoo Song , Ryuichi Yamamoto , Min-Jae Hwang , Jin-Seob Kim , Ohsung Kwon , Jae-Min Kim

To make a good balance between performance, cost, and power consumption, a hybrid intelligent reflecting surface (IRS)-aided directional modulation (DM) network is investigated in this paper, where the hybrid IRS consists of passive and…

信息论 · 计算机科学 2023-02-15 Rongen Dong , Hangjia He , Feng Shu , Riqing Chen , Jiangzhou Wang

To address practical challenges in establishing and maintaining robust wireless connectivity such as multi-path effects, low latency, size reduction, and high data rate, the digital beamformer is performed by the hybrid antenna array at the…

网络与互联网体系结构 · 计算机科学 2022-11-04 Somayeh Komeylian , Christopher Paolini

Neural rendering of implicit surfaces performs well in 3D vision applications. However, it requires dense input views as supervision. When only sparse input images are available, output quality drops significantly due to the shape-radiance…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Haoyu Wu , Alexandros Graikos , Dimitris Samaras

In this paper we present a new robust sound source localization and tracking method using an array of eight microphones (US patent pending) . The method uses a steered beamformer based on the reliability-weighted phase transform (RWPHAT)…

机器人学 · 计算机科学 2016-04-07 Jean-Marc Valin , François Michaud , Jean Rouat

Attention-based models have made tremendous progress on end-to-end automatic speech recognition(ASR) recently. However, the conventional transformer-based approaches usually generate the sequence results token by token from left to right,…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Xi Chen , Songyang Zhang , Dandan Song , Peng Ouyang , Shouyi Yin

A new spatial IIR beamformer based direction-of-arrival (DoA) estimation method is proposed in this paper. We propose a retransmission based spatial feedback method for an array of transmit and receive antennas that improves the performance…

信号处理 · 电气工程与系统科学 2023-09-07 Parth Mehta , Kumar Appaiah , Rajbabu Velmurugan

Mobile robots in real-life settings would benefit from being able to localize sound sources. Such a capability can nicely complement vision to help localize a person or an interesting event in the environment, and also to provide enhanced…

机器人学 · 计算机科学 2016-03-01 Jean-Marc Valin , François Michaud , Brahim Hadjou , Jean Rouat

We propose a new robust distributed linearly constrained beamformer which utilizes a set of linear equality constraints to reduce the cross power spectral density matrix to a block-diagonal form. The proposed beamformer has a convenient…

信号处理 · 电气工程与系统科学 2019-05-28 Andreas I. Koutrouvelis , Thomas W. Sherson , Richard Heusdens , Richard C. Hendriks

Acoustic beamforming models typically assume wide-sense stationarity of speech signals within short time frames. However, voiced speech is better modeled as a cyclostationary (CS) process, a random process whose mean and autocorrelation are…

音频与语音处理 · 电气工程与系统科学 2025-12-16 Giovanni Bologni , Richard Heusdens , Richard C. Hendriks