English
Related papers

Related papers: Can all variations within the unified mask-based b…

200 papers

Simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-aided cell-free massive multiple-input multiple-output (CF-mMIMO) systems are investigated under spatially correlated fading channels using realistic…

Information Theory · Computer Science 2025-01-06 Zeping Sui , Hien Quoc Ngo , Michail Matthaiou , Lajos Hanzo

Optimal BeamFormers (BFs) that maximize the Weighted Sum Rate (WSR) for a Multiple-Input Multiple-Output (MIMO) interference broadcast channel (IBC) remains an important research area. Under practical scenarios, the problem is compounded by…

Signal Processing · Electrical Eng. & Systems 2017-10-27 Kalyana Gopala , Dirk Slock

A major bottleneck of standard auto-regressive large language models is that their inference process is inherently sequential, resulting in very long and costly inference times. To circumvent this, practitioners proposed a class of language…

Machine Learning · Computer Science 2025-11-11 Sitan Chen , Kevin Cong , Jerry Li

In this paper, we consider fast wireless data aggregation via over-the-air computation (AirComp) in Internet of Things (IoT) networks, where an access point (AP) with multiple antennas aim to recover the arithmetic mean of sensory data from…

Signal Processing · Electrical Eng. & Systems 2021-05-12 Wenzhi Fang , Yinan Zou , Hongbin Zhu , Yuanming Shi , Yong Zhou

Recent neural network strategies for source separation attempt to model audio signals by processing their waveforms directly. Mean squared error (MSE) that measures the Euclidean distance between waveforms of denoised speech and the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-06-05 Shrikant Venkataramani , Ryley Higa , Paris Smaragdis

Explainable boosting machines (EBMs) are popular "glass-box" models that learn a set of univariate functions using boosting trees. These achieve explainability through visualizations of each feature's effect. However, unlike linear model…

Machine Learning · Statistics 2026-03-31 Haimo Fang , Kevin Tan , Jonathan Pipping-Gamon , Giles Hooker

This paper proposes the use of two task-aware warping factors in mask-based speech enhancement (SE). One controls the balance between speech-maintenance and noise-removal in training phases, while the other controls SE power applied to…

Sound · Computer Science 2021-08-30 Qiongqiong Wang , Kong Aik Lee , Takafumi Koshinaka , Koji Okabe , Hitoshi Yamamoto

Acoustic beamformers have been widely used to enhance audio signals. Currently, the best methods are the deep neural network (DNN)-powered variants of the generalized eigenvalue and minimum-variance distortionless response beamformers and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Yuichiro Koyama , Bhiksha Raj

Joint optimization of multi-channel front-end and automatic speech recognition (ASR) has attracted much interest. While promising results have been reported for various tasks, past studies on its meeting transcription application were…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-30 Xiaofei Wang , Naoyuki Kanda , Yashesh Gaur , Zhuo Chen , Zhong Meng , Takuya Yoshioka

This paper introduces an innovative quantum-inspired method for beamforming (BF) optimization in multiple-input multiple-output (MIMO) arrays. The method leverages the simulated bifurcation (SB) algorithm to address the complex…

As a fundamental task in computer vision, semantic segmentation is widely applied in fields such as autonomous driving, remote sensing image analysis, and medical image processing. In recent years, Transformer-based segmentation methods…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Tai An , Weiqiang Huang , Da Xu , Qingyuan He , Jiacheng Hu , Yujia Lou

The evolution of fifth generation (5G) wireless communication networks has led to an increased need for wireless resource management solutions that provide higher data rates, wide coverage, low latency, and power efficiency. Yet, many of…

Information Theory · Computer Science 2024-06-13 Cemil Vahapoglu , Timothy J. O'Shea , Tamoghna Roy , Sennur Ulukus

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-17 Zhong Meng , Shinji Watanabe , John R. Hershey , Hakan Erdogan

In this paper, we develop a unified dynamic intelligent reflecting surface (IRS) beamforming framework to boost the sum computation rate of an IRS-aided mobile edge computing (MEC) system, where each device follows a binary offloading…

Information Theory · Computer Science 2022-06-01 Guangji Chen , Qingqing Wu

We study the properties of beamformers in their ability to either maintain or estimate the true signal power of the signal of interest (SOI). Our focus is particularly on the Capon beamformer and the minimum mean squared error (MMSE)…

Signal Processing · Electrical Eng. & Systems 2025-06-23 Esa Ollila , Xavier Mestre , Elias Raninen

We study physical layer multicasting in multicell networks where each base station, equipped with multiple antennas, transmits a common message using a single beamformer to multiple users in the same cell. We investigate two coordinated…

Information Theory · Computer Science 2012-10-23 Zhengzheng Xiang , Meixia Tao , Xiaodong Wang

Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Thomas Drugman , Yannis Stylianou

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters…

Sound · Computer Science 2024-01-08 Shulin He , Jinjiang liu , Hao Li , Yang Yang , Fei Chen , Xueliang Zhang

Over-the-air computation (AirComp) based federated learning (FL) is capable of achieving fast model aggregation by exploiting the waveform superposition property of multiple access channels. However, the model aggregation performance is…

Information Theory · Computer Science 2022-04-01 Zhibin Wang , Jiahang Qiu , Yong Zhou , Yuanming Shi , Liqun Fu , Wei Chen , Khaled B. Lataief

A generalization of the maximum noise fraction (MNF) transform is proposed. Powers of each band are included as new bands before the MNF transform is performed. The generalized MNF (GMNF) is shown to perform better than the MNF on a time…

Data Analysis, Statistics and Probability · Physics 2009-11-06 Christopher Gordon