中文
相关论文

相关论文: BeamTransformer: Microphone Array-based Overlappin…

200 篇论文

Beamforming techniques are utilized in millimeter wave (mmWave) communication to address the inherent path loss limitation, thereby establishing and maintaining reliable connections. However, adopting standard defined beamforming approach…

网络与互联网体系结构 · 计算机科学 2025-09-16 Muhammad Baqer Mollah , Honggang Wang , Hua Fang

In this paper we present a Transformer-Transducer model architecture and a training technique to unify streaming and non-streaming speech recognition models into one model. The model is composed of a stack of transformer layers for audio…

声音 · 计算机科学 2020-10-08 Anshuman Tripathi , Jaeyoung Kim , Qian Zhang , Han Lu , Hasim Sak

Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have…

声音 · 计算机科学 2018-05-16 Hiroshi Seki , Takaaki Hori , Shinji Watanabe , Jonathan Le Roux , John R. Hershey

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

音频与语音处理 · 电气工程与系统科学 2026-02-18 Ilai Zaidel , Sharon Gannot

The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder…

声音 · 计算机科学 2017-03-16 Tsubasa Ochiai , Shinji Watanabe , Takaaki Hori , John R. Hershey

Automatic speech recognition (ASR) technologies have been significantly advanced in the past few decades. However, recognition of overlapped speech remains a highly challenging task to date. To this end, multi-channel microphone array data…

音频与语音处理 · 电气工程与系统科学 2021-08-31 Jianwei Yu , Shi-Xiong Zhang , Bo Wu , Shansong Liu , Shoukang Hu , Mengzhe Geng , Xunying Liu , Helen Meng , Dong Yu

The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…

音频与语音处理 · 电气工程与系统科学 2025-04-03 Tianqin Zheng , Jilu Jin , Hanchen Pei , Gongping Huang , Jingdong Chen , Jacob Benesty

Spatial aliasing affects spaced microphone arrays, causing directional ambiguity above certain frequencies, degrading spatial and spectral accuracy of beamformers. Given the limitations of conventional signal processing and the scarcity of…

音频与语音处理 · 电气工程与系统科学 2025-10-21 Mateusz Guzik , Giulio Cengarle , Daniel Arteaga

Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are…

音频与语音处理 · 电气工程与系统科学 2025-06-26 Taous Iatariene , Can Cui , Alexandre Guérin , Romain Serizel

We address the problem of effectively handling overlapping speech in a diarization system. First, we detail a neural Long Short-Term Memory-based architecture for overlap detection. Secondly, detected overlap regions are exploited in…

音频与语音处理 · 电气工程与系统科学 2019-10-28 Latané Bullock , Hervé Bredin , Leibny Paola Garcia-Perera

Millimeter wave (mmWave) communication, utilizing beamforming techniques to address the inherent path loss limitation, is considered as one of the key technologies to support ever increasing high throughput and low latency demands of…

网络与互联网体系结构 · 计算机科学 2026-02-17 Muhammad Baqer Mollah , Honggang Wang , Mohammad Ataul Karim , Hua Fang

This paper introduces a parallel and asynchronous Transformer framework designed for efficient and accurate multilingual lip synchronization in real-time video conferencing systems. The proposed architecture integrates translation, speech…

多媒体 · 计算机科学 2025-12-23 Eren Caglar , Amirkia Rafiei Oskooei , Mehmet Kutanoglu , Mustafa Keles , Mehmet S. Aktas

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions to the real-world. Distributed sensor arrays that consider several devices with a few microphones is a viable…

声音 · 计算机科学 2020-03-17 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows…

音频与语音处理 · 电气工程与系统科学 2025-03-25 Luan Vinícius Fiorio , Bruno Defraene , Johan David , Alex Young , Frans Widdershoven , Wim van Houtum , Ronald M. Aarts

The spatial covariance matrix has been considered to be significant for beamformers. Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and…

声音 · 计算机科学 2021-09-03 Andong Li , Wenzhe Liu , Chengshi Zheng , Xiaodong Li

We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It consists of subsystems for signal synchronization,…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Tobias Gburrek , Christoph Boeddeker , Thilo von Neumann , Tobias Cord-Landwehr , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

This paper investigates deep learning techniques to predict transmit beamforming based on only historical channel data without current channel information in the multiuser multiple-input-single-output downlink. This will significantly…

信息论 · 计算机科学 2023-02-03 Juping Zhang , Gan Zheng , Yangyishi Zhang , Ioannis Krikidis , Kai-Kit Wong

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and…

音频与语音处理 · 电气工程与系统科学 2018-10-17 Zhong Meng , Shinji Watanabe , John R. Hershey , Hakan Erdogan

Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and…

声音 · 计算机科学 2018-04-19 Yi Luo , Zhuo Chen , Nima Mesgarani

In this paper, we propose a speech enhancement method us ing dual-path Multi-Channel Linear Prediction (MCLP) filters and multi-norm beamforming. Specifically, the MCLP part in the proposed method is designed with dual-path filters in both…

音频与语音处理 · 电气工程与系统科学 2025-07-25 Chengyuan Qin , Wenmeng Xiong , Jing Zhou , Maoshen Jia , Changchun Bao