中文
相关论文

相关论文: Rethinking the Separation Layers in Speech Separat…

200 篇论文

In recent years, continuous-aperture multiple-input multiple-output (CAP-MIMO) is reinvestigated to achieve improved communication performance with limited antenna apertures. Unlike the classical MIMO composed of discrete antennas, CAP-MIMO…

信息论 · 计算机科学 2023-01-03 Zijian Zhang , Linglong Dai

The electromagnetic (EM) features of reconfigurable intelligent surfaces (RISs) fundamentally determine their operating principles and performance. Motivated by these considerations, we study a single-input single-output (SISO) system in…

信息论 · 计算机科学 2024-04-15 Nemanja Stefan Perović , Le-Nam Tran , Marco Di Renzo , Mark F. Flanagan

Single Input-Multiple Output (SIMO) systems are key enablers of high data rates in the next generation wireless communications. However in SIMO systems, channel estimation and equalization are challenging particularly in the presence of…

信息论 · 计算机科学 2025-12-09 K Sai Praneeth , P Aswathylakshmi , Radhakrishna Ganti

In this paper, we present three new signal designs for Enhanced Spatial Modulation (ESM), which was recently introduced by the present authors. The basic idea of ESM is to convey information bits not only by the index(es) of the active…

网络与互联网体系结构 · 计算机科学 2016-05-31 Chien-Chun Cheng , Hikmet Sari , Serdar Sezginer , Yu T. Su

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by…

音频与语音处理 · 电气工程与系统科学 2021-11-22 Tom O'Malley , Arun Narayanan , Quan Wang , Alex Park , James Walker , Nathan Howard

Massive multiple-input multiple-output (MIMO) is expected to be a vital component in future 5G systems. As such, there is a need for new modeling in order to investigate the performance of massive MIMO not only at the physical layer, but…

信号处理 · 电气工程与系统科学 2019-05-13 Emma Fitzgerald , Michał Pióro , Fredrik Tufvesson

Multiplexing services as a key communication technique to effectively combine multiple signals into one signal and transmit over a shared medium. Multiplexing can increase the channel capacity by requiring more resources on the transmission…

信息论 · 计算机科学 2017-12-27 Chunlin Ji , Ruopeng Liu

While the performance of offline neural speech separation systems has been greatly advanced by the recent development of novel neural network architectures, there is typically an inevitable performance gap between the systems and their…

声音 · 计算机科学 2023-02-22 Kai Li , Yi Luo

We propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES contains two methods, speaker separation using speaker…

声音 · 计算机科学 2020-10-22 Peidong Wang , Zhuo Chen , DeLiang Wang , Jinyu Li , Yifan Gong

General audio foundation models have recently achieved remarkable progress, enabling strong performance across diverse tasks. However, state-of-the-art models remain extremely large, often with hundreds of millions of parameters, leading to…

人工智能 · 计算机科学 2026-04-29 Mohammed Ali El Adlouni , Aurian Quelennec , Pierre Chouteau , Geoffroy Peeters , Slim Essid

The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by…

We first conceive a novel transmission protocol for a multi-relay multiple-input--multiple-output orthogonal frequency-division multiple-access (MIMO-OFDMA) cellular network based on joint transmit and receive beamforming. We then address…

信息论 · 计算机科学 2015-03-04 Kent Tsz Kan Cheung , Shaoshi Yang , Lajos Hanzo

This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides single-modal…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

Language models trained on large-scale corpora can generate remarkably fluent results in open-domain dialogue. However, for the persona-based dialogue generation task, consistency and coherence are also key factors, which are great…

人工智能 · 计算机科学 2023-05-23 Junkai Zhou , Liang Pang , Huawei Shen , Xueqi Cheng

The goal of speech separation is to extract multiple speech sources from a single microphone recording. Recently, with the advancement of deep learning and availability of large datasets, speech separation has been formulated as a…

音频与语音处理 · 电气工程与系统科学 2021-11-17 Midia Yousefi , John H. L. Hansen

A massive single-input multiple-output (SIMO) system with a single transmit antenna and a large number of receive antennas in intersymbol interference (ISI) channels is considered. Contrast to existing energy detection (ED)-based…

信号处理 · 电气工程与系统科学 2019-10-02 Huiqiang Xie , Weiyang Xu , Wei Xiang , Ke Shao , Shengbo Xu

In multiple-input multiple-output (MIMO), multiple radio frequency (RF) chains are usually required to simultaneously transmit multiple data streams. As a special MIMO technology, spatial modulation (SM) activates one transmit antenna with…

信息论 · 计算机科学 2020-09-03 Qiang Li , Miaowen Wen , Marco Di Renzo

Continuous speech separation for meeting pre-processing has recently become a focused research topic. Compared to the data in utterance-level speech separation, the meeting-style audio stream lasts longer, has an uncertain number of…

音频与语音处理 · 电气工程与系统科学 2022-02-11 Chenda Li , Lei Yang , Weiqin Wang , Yanmin Qian

Diffusion models, emerging as powerful deep generative tools, excel in various applications. They operate through a two-steps process: introducing noise into training samples and then employing a model to convert random noise into new…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Huijie Zhang , Yifu Lu , Ismail Alkhouri , Saiprasad Ravishankar , Dogyoon Song , Qing Qu

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li