English
Related papers

Related papers: Can all variations within the unified mask-based b…

200 papers

The literature is abundant with methodologies focusing on using transformer architectures due to their prominence in wireless signal processing and their capability to capture long-range dependencies via attention mechanisms. In particular,…

Information Theory · Computer Science 2025-04-17 Cemil Vahapoglu , Timothy J. O'Shea , Wan Liu , Tamoghna Roy , Sennur Ulukus

Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder architectures with one-step inference. Inspired by test-time…

Sound · Computer Science 2026-03-12 Zhenghai You , Ying Shi , Lantian Li , Dong Wang

Recent masked diffusion models (MDMs) have shown competitive performance compared to autoregressive models (ARMs) for language modeling. While most literature has focused on performance enhancing sampling procedures, efficient sampling from…

Machine Learning · Computer Science 2025-06-02 Heli Ben-Hamu , Itai Gat , Daniel Severo , Niklas Nolte , Brian Karrer

We present the first neural target speech extraction (TSE) system that uses human feedback for iterative refinement. Our approach allows users to mark specific segments of the TSE output, generating an edit mask. The refinement system then…

Sound · Computer Science 2025-08-06 Malek Itani , Ashton Graves , Sefik Emre Eskimez , Shyamnath Gollakota

We investigate feature selection problem for generic machine learning models. We introduce a novel framework that selects features considering the outcomes of the model. Our framework introduces a novel feature masking approach to eliminate…

Machine Learning · Computer Science 2024-12-10 Mehmet E. Lorasdagi , Mehmet Y. Turali , Suleyman S. Kozat

The events of recent years have highlighted the importance of telemedicine solutions which could potentially allow remote treatment and diagnosis. Relatedly, Computational Paralinguistics, a unique subfield of Speech Processing, aims to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-31 Tamás Grósz , Mittul Singh , Sudarsana Reddy Kadiri , Hemant Kathania , Mikko Kurimo

In this paper, we present a Small Energy Masking (SEM) algorithm, which masks inputs having values below a certain threshold. More specifically, a time-frequency bin is masked if the filterbank energy in this bin is less than a certain…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-18 Chanwoo Kim , Kwangyoun Kim , Sathish Reddy Indurthi

We present an analytical, simulation, and experimental-based study of beamforming Multiple Input Single Output (MISO) systems. We analyze the performance of beamforming MISO systems taking into account implementation complexity and effects…

Information Theory · Computer Science 2016-11-15 Melissa Duarte , Ashutosh Sabharwal , Chris Dick , Raghu Rao

In this paper, we propose two mask-based beamforming methods using a deep neural network (DNN) trained by multichannel loss functions. Beamforming technique using time-frequency (TF)-masks estimated by a DNN have been applied to many…

Sound · Computer Science 2019-07-12 Yoshiki Masuyama , Masahito Togami , Tatsuya Komatsu

Audio-Visual Target Speaker Extraction (AV-TSE) aims to mimic the human ability to enhance auditory perception using visual cues. Although numerous models have been proposed recently, most of them estimate target signals by primarily…

Sound · Computer Science 2025-04-02 Wenxuan Wu , Xueyuan Chen , Shuai Wang , Jiadong Wang , Lingwei Meng , Xixin Wu , Helen Meng , Haizhou Li

Motivated by massive deployment of low data rate Internet of things (IoT) and ehealth devices with requirement for highly reliable communications, this paper proposes receive beamforming techniques for the uplink of a single-input…

Information Theory · Computer Science 2017-05-17 Majid Bavand , Steven D. Blostein

We propose a framework for extracting the bone surface from B-mode images employing the eigenspace minimum variance (ESMV) beamformer and a ridge detection method. We show that an ESMV beamformer with a rank-1 signal subspace can preserve…

Medical Physics · Physics 2016-09-07 Saeed Mehdizadeh , Sebastien Muller , Gabriel Kiss , Tonni F. Johansen , Sverre Holm

This work revisits the joint beamforming (BF) and antenna selection (AS) problem, as well as its robust beamforming (RBF) version under imperfect channel state information (CSI). Such problems arise due to various reasons, e.g., the costly…

Signal Processing · Electrical Eng. & Systems 2023-04-12 Sagar Shrestha , Xiao Fu , Mingyi Hong

Far-field speech recognition is a challenging task that conventionally uses signal processing beamforming to attack noise and interference problem. But the performance has been found usually limited due to heavy reliance on environmental…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Dongdi Zhao , Jianbo Ma , Lu Lu , Jinke Li , Xuan Ji , Lei Zhu , Fuming Fang , Ming Liu , Feijun Jiang

Linear hybrid beamformer designs are conceived for the decentralized estimation of a vector parameter in a millimeter wave (mmWave) multiple-input multiple-output (MIMO) Internet of Things network (IoTNe). The proposed designs incorporate…

This paper considers base station (BS) cooperation in the form of coordinated beamforming, focusing on min-max fairness in the power usage subject to target SINR constraints. We show that the optimal beamforming strategies have an…

Information Theory · Computer Science 2012-02-06 Randa Zakhour , Stephen V. Hanly

Multi-objective embedding-based retrieval (EBR) has become increasingly critical due to the growing complexity of user behaviors and commercial objectives. While traditional approaches often suffer from data sparsity and limited information…

Information Retrieval · Computer Science 2025-04-18 Hao Deng , Haibo Xing , Kanefumi Matsuyama , Moyu Zhang , Jinxin Hu , Hong Wen , Yu Zhang , Xiaoyi Zeng , Jing Zhang

The concept of reconfigurable fluid antennas (FA) is a potential and promising solution to enhance the spectral efficiency of wireless communication networks. Despite their many advantages, FA-enabled communications have limitations as they…

Information Theory · Computer Science 2022-12-19 Christodoulos Skouroumounis , Ioannis Krikidis

Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may…

Sound · Computer Science 2025-08-12 Shu Wu , Anbin Qi , Yanzhang Xie , Xiang Xie

Downlink beamforming is a key technology for cellular networks. However, computing the transmit beamformer that maximizes the weighted sum rate subject to a power constraint is an NP-hard problem. As a result, iterative algorithms that…

Signal Processing · Electrical Eng. & Systems 2020-06-16 Lissy Pellaco , Mats Bengtsson , Joakim Jaldén