English
Related papers

Related papers: Can all variations within the unified mask-based b…

200 papers

In dynamic acoustic environments with time-varying interferers, effective beamforming requires identifying stationary regions over time. The Capon beamformer, a whitened matched filter constrained to maintain unity gain in the desired…

Signal Processing · Electrical Eng. & Systems 2026-05-26 Manan Mittal , Ryan M. Corey , Diego Cuji , John R. Buck , Andrew C. Singer

The choice of an optimal time-frequency resolution is usually a difficult but important step in tasks involving speech signal classification, e.g., speech anti-spoofing. The variations of the performance with different choices of…

Sound · Computer Science 2021-10-12 Wei Liu , Meng Sun , Xiongwei Zhang , Hugo Van hamme , Thomas Fang Zheng

This study presents a novel method for source extraction, referred to as the similarity-and-independence-aware beamformer (SIBF). The SIBF extracts the target signal using a rough magnitude spectrogram as the reference signal. The advantage…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-22 Atsuo Hiroe

This work proposes a mixed learning-based and optimization-based approach to the weighted-sum-rates beamforming problem in a multiple-input multiple-output (MIMO) wireless network. The conventional methods, i.e., the fractional programming…

Information Theory · Computer Science 2026-01-07 Jianhang Zhu , Tsung-Hui Chang , Liyao Xiang , Kaiming Shen

Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Lei Ke , Martin Danelljan , Xia Li , Yu-Wing Tai , Chi-Keung Tang , Fisher Yu

In end-to-end multi-channel speech enhancement, the traditional approach of designating one microphone signal as the reference for processing may not always yield optimal results. The limitation is particularly in scenarios with large…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Wang Dai , Xiaofei Li , Archontis Politis , Tuomas Virtanen

Foundation Models (FMs) have revolutionized machine learning with their adaptability and high performance across tasks; yet, their integration into Federated Learning (FL) is challenging due to substantial communication overhead from their…

Machine Learning · Computer Science 2023-11-30 Vasileios Tsouvalas , Yuki Asano , Aaqib Saeed

Ultrasound (US) imaging exhibits substantial heterogeneity across anatomical structures and acquisition protocols, posing significant challenges to the development of generalizable analysis models. Most existing methods are task-specific,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Bo Deng , Yitong Tang , Jiake Li , Yuxin Huang , Li Wang , Yu Zhang , Yufei Zhan , Hua Lu , Xiaoshen Zhang , Jieyun Bai

Time-frequency domain dual-path models have demonstrated strong performance and are widely used in source separation. Because their computational cost grows with the number of frequency bins, these models often use the band-split (BS)…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-10 Kohei Saijo , Yoshiaki Bando

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common application scenario is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-14 Zhuohuang Zhang , Takuya Yoshioka , Naoyuki Kanda , Zhuo Chen , Xiaofei Wang , Dongmei Wang , Sefik Emre Eskimez

Masked Autoencoders (MAEs) learn rich semantic representations in audio classification through an efficient self-supervised reconstruction task. However, general-purpose models fail to generalize well when applied directly to fine-grained…

Machine Learning · Computer Science 2025-08-20 Lukas Rauch , René Heinrich , Ilyass Moummad , Alexis Joly , Bernhard Sick , Christoph Scholz

In this paper, we develop a functional weighted minimum mean-squared error (WMMSE) algorithm for downlink beamforming in multiuser continuous aperture array (CAPA) systems where both the base station (BS) and users are equipped with CAPAs.…

Signal Processing · Electrical Eng. & Systems 2025-11-19 Shiyong Chen , Shengqian Han , Jia Guo

Massive multi-user (MU) multiple-input multiple-output (MIMO) promises significant gains in spectral efficiency compared to traditional, small-scale MIMO technology. Linear equalization algorithms, such as zero forcing (ZF) or minimum…

Information Theory · Computer Science 2018-11-12 Charles Jeon , Kaipeng Li , Joseph R. Cavallaro , Christoph Studer

Although there have been extensive studies on transmit beamforming in multi-input single-output (MISO) multicell networks, achieving optimal sum-rate with limited channel state information (CSI) is still a challenge even with a single user…

Networking and Internet Architecture · Computer Science 2019-04-11 Youjin Kim , Hyun Jong Yang

We consider the problem of optimal distributed beamforming in a sensor network where the sensors observe a dynamic parameter in noise and coherently amplify and forward their observations to a fusion center (FC). The FC uses a Kalman filter…

Information Theory · Computer Science 2013-04-02 Feng Jiang , Jie Chen , A. Lee Swindlehurst

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-24 Zhepei Wang , Ritwik Giri , Devansh Shah , Jean-Marc Valin , Michael M. Goodwin , Paris Smaragdis

The main challenges in designing downlink coordinated multicast beamforming in massive multiple-input multiple output (MIMO) cellular networks are the complex computational solutions and significant fronthaul overhead for centralized…

Signal Processing · Electrical Eng. & Systems 2025-02-20 Shiqi Yin , Min Dong

We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy…

Machine Learning · Computer Science 2026-03-24 Sitong Chen , Shen Nie , Jiacheng Sun , Zijin Feng , Zhenguo Li , Ji-Rong Wen , Chongxuan Li

Formulations of the turbo equalization approach to iterative equalization and decoding vary greatly when channel knowledge is either partially or completely unknown. Maximum aposteriori probability (MAP) and minimum mean square error (MMSE)…

Systems and Control · Computer Science 2012-03-20 Nargiz Kalantarova , Kyeongyeon Kim , Suleyman S. Kozat , Andrew C. Singer

Face recognition under ideal conditions is now considered a well-solved problem with advances in deep learning. Recognizing faces under occlusion, however, still remains a challenge. Existing techniques often fail to recognize faces with…

Computer Vision and Pattern Recognition · Computer Science 2022-02-16 Shaozhe Hao , Chaofeng Chen , Zhenfang Chen , Kwan-Yee K. Wong